Skip to main content
Glama

Server Details

Connect your health, fitness, nutrition, sleep, and wearable data to your AI assistant.

Ownership verified
Status
Healthy
Uptime
73.8% over 54 days
OAuth
Works in Glama
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
turnnoblindeye/wellness-project-mcp
GitHub Stars
2
Server Listing
Wellness Project MCP

TDQS

A3.9/5.0

Scored across 81 tools

Disambiguation3/5

Many tools share domains (multiple meal, workout, sleep, and wearable tools, plus over a dozen visualization tools), and boundaries are clear mainly because descriptions include explicit routing heuristics. An agent could still confuse list_body_metrics vs show_body_weight vs show_body_composition, or list_wearable_data vs show_health_overview vs show_recovery, so overlap is more than one or two tools.

Naming Consistency5/5

All tools use lower snake_case with a predictable verb-first pattern (list_, log_, update_, delete_, get_, show_, manage_, set_). The only minor deviation is the compound verb in add_or_update_personal_context, but overall naming is highly consistent and readable.

Tool Count1/5

81 tools is far beyond any reasonable scope for a single server and exceeds the 50+ threshold for extreme mismatch. Even though the server spans many wellness domains, the surface is too large to be well-scoped and imposes unnecessary selection burden.

Completeness4/5

Core CRUD coverage is present for most domains (meals, workouts, cycle, injuries, lab results, supplements, recovery, wellbeing, goals, locations). Missing operations like update_run, delete_sleep, delete_wearable, and delete_location are workaroundable via delete-and-recreate or device-level fixes, so gaps are minor relative to the huge surface.

Available Tools

81 tools
add_or_update_personal_contextAInspect

Add a new Personal Context memory, or update an existing one by id. A memory is a durable circumstance or preference that should carry across future unrelated conversations (e.g. "travels most weeks", "gym has no squat rack", "prefers short home workouts", "wants blunt feedback"). Use list_personal_context first to check whether an existing memory already covers the subject, and pass its id with operation update rather than creating a duplicate.

Health history does not belong here. Injuries, lab results, meals, workouts, sleep and body metrics each have their own dedicated tools that store them as structured data the app can chart and reason over; writing any of them as a memory duplicates that record and degrades it to loose text.

This writes immediately with no separate approval step. Free accounts are capped at 3 memories and Pro accounts at 50; updating an existing memory by id is always allowed even at the cap. There is no delete or bulk-write capability here.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesThe memory text: a compact, durable fact or preference with no conversational filler.
memory_idNoRequired for update. The numeric id from list_personal_context. Omit for add.
operationYesadd creates a new memory. update replaces the content of an existing one, identified by memory_id.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that writes happen immediately with no separate approval step, that free accounts are capped at 3 memories and Pro at 50, that updates by id are allowed even at the cap, and that there is no delete/bulk capability. This meaningfully extends what the annotations alone communicate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than typical, but each paragraph earns its place: purpose, do-not-duplicate workflow, exclusions for structured health data, and write/cap behavior. The core action is front-loaded before broader guardrails.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 3-parameter tool with full schema coverage, an output schema, and meaningful annotations, the description closes the important gaps: when to update vs add, what content belongs, what caps apply, and why duplicate structured records should be avoided. No critical dimension appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all three parameters at 100%, but the description adds real semantics: it clarifies that content must be a compact durable fact/preference, illustrates valid content with examples, and explains the relationship between memory_id and operation update in the list-first workflow.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Add a new Personal Context memory, or update an existing one by id.' It also distinguishes itself from nearby tools in the sibling list by explaining that memories are durable circumstances/preferences and that structured health data belongs in other dedicated tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit, actionable guidance: use list_personal_context first, pass the existing id with operation update instead of duplicating, and avoid storing metrics/trackable health data here because dedicated tools exist for those. This is strong when-versus-alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_rest_dayA
DestructiveIdempotent
Inspect

Remove a previously declared Rest Day, Sick Day, or Travel Day status.

INFER — do not ask:

  • date: parse the user's reference; default to today.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoThe date to un-mark. Format: YYYY-MM-DD. Default: today.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the agent knows this mutates state and is safe to retry. The description adds only that the target is a 'previously declared' status of one of three types; it says nothing about permissions, error behavior on a non-existent day, or reversibility beyond what annotations already cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the mutation front-loaded and the inference rule clearly separated; no filler. Slightly verbose in enumerating the three status types, but nothing wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no prose, the single parameter is fully documented in the schema, and annotations cover the destructive/idempotent profile. For a one-parameter undo tool, this is nearly complete; only edge behavior on missing entries is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter's description already includes 'Format: YYYY-MM-DD. Default: today.' The description's 'default to today' and inference instruction duplicate rather than extend the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Remove/cancel) and a precise resource (previously declared Rest Day, Sick Day, or Travel Day status), which is clearer than the tool name alone and distinguishes it from log_rest_day/list_rest_days. It does not, however, explicitly name the sibling it complements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'previously declared' implies this is for undoing an existing marking, and the INFER block gives concrete invocation guidance for the date. There is no explicit when-not-to-use or named alternative (e.g., log_rest_day to re-mark), so the guidance remains implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_goalInspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Create a new Wellness Project goal or standard target. Call this directly when the user wants to establish a goal; there is no schema-discovery or list_goals prerequisite. Available goal types and inputs come from the canonical Goals definitions.

This tool loads the user's current goals itself before writing, so do not call list_goals first: it reports what a standard target changed from, and refuses to stack a second goal on top of one that already covers the same thing, naming that goal's ID to use with update_goal. Infer the goal_type and canonical inputs from the request. Formal goal fields are described on the generated inputs.

WEIGHT INPUTS: For body_weight_target_lb, use target_value_lb. For body_comp metric lean_mass_lb, use start_value_lb and target_value_lb. These aliases are unit-aware; in-app calls use the preferred unit, while external calls use pounds unless input_weight_unit is set. Legacy generic fields for these weight goals remain canonical pounds. Other standard targets use target_value; body-fat values remain percentages.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoConcise title, inferred from the goal inputs.
weeksNoFor weight_loss. Duration in weeks when target_date is not supplied. For body_comp. Duration in weeks when target_date is not supplied.
metricNoFor consistency. What consistency behavior to track. Required on create. Set at create, not editable later. For body_comp. Body composition metric. Required on create. Set at create, not editable later. For nutrition. Legacy nutrition metric.
new_nameNoFor n1_experiment. Name for a new supplement when not using supplement_id.
goal_typeYesGoal type to create.
new_brandNoFor n1_experiment. Optional brand for a new supplement.
race_dateNoFor race. Race date in YYYY-MM-DD format. Required on create.
start_dateNoYYYY-MM-DD. Default: today.
start_valueNoFor body_comp. Starting value when the goal begins. Required on create. Set at create, not editable later.
target_dateNoFor weight_loss. Target date in YYYY-MM-DD format. Use this when the user names a deadline. For body_comp. Target date in YYYY-MM-DD format. Use this when the user names a deadline.
week_windowNoFor consistency. How the week this goal is measured against is bounded: rolling = the last 7 days, sunday/monday = a calendar week that resets on that day. Default: rolling.
start_1rm_lbNoFor strength. Estimated 1RM when the goal starts. Required on create. Set at create, not editable later. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_hoursNoFor consistency. Nightly sleep target in hours. Required when metric is sleep_duration.
target_valueNoFor body_comp. Target body composition value. Required on create. For nutrition. Legacy nutrition target value. For standard targets, this is the numeric target value.
exercise_nameNoFor strength. Exercise name. Required on create. Set at create, not editable later.
new_dose_unitNoFor n1_experiment. Dose unit for a new supplement.
supplement_idNoFor n1_experiment. Existing supplement ID. Use either supplement_id or the new-supplement fields.
target_1rm_lbNoFor strength. Target 1RM. Required on create. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
start_value_lbNoStarting lean mass in body_comp. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
new_dose_amountNoFor n1_experiment. Dose amount for a new supplement.
start_weight_lbNoFor weight_loss. Starting body weight. Required on create. Set at create, not editable later. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_per_weekNoFor consistency. Target occurrences per week. Required on create.
target_time_secNoFor race. Target finish time in seconds. Required on create.
target_value_lbNoWeight goal target. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_weight_lbNoFor weight_loss. Target body weight. Required on create. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
input_weight_unitNoSet to kg when the user gave kg for the _lb fields in this object. Omit when they are already lb.
intervention_daysNoFor n1_experiment. Intervention duration in days.
baseline_directionNoFor n1_experiment. Use already logged previous 14 days or collect the next 14 days.
target_distance_miNoFor race. Target race distance. Required on create. In mi, or km with input_distance_unit set. See UNIT INPUTS.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.
acknowledged_warningsNoWarning keys the user explicitly acknowledged after a guarded create attempt. Omit otherwise.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.
delete_cycleA
DestructiveIdempotent
Inspect

Delete a period record. Only delete if the user explicitly asks to remove a specific record. Do not delete to "fix" a record — use update_cycle instead.

SELECTOR — pass id if known, or date (the period's start date, or any date that falls within it) to resolve it. Exactly one required. If date matches more than one record, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRecord ID. Alternative to date.
dateNoAlternative to id: a date (YYYY-MM-DD) that identifies the period — its start date, or any day within it. Resolves only when exactly one record matches.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the destructive nature is covered. The description adds valuable behavioral detail: exactly one of id or date is required, date resolution matches only the period start date or any date within it, and an ambiguous match causes an error that returns candidate IDs to retry with. This goes beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by usage constraints and selector details. Every sentence earns its place, with no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter destructive tool with annotations and an output schema, the description is complete. It covers what the tool does, when to use it, when not to use it, how to select the target, and how ambiguity is handled. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. However, the description adds important semantics: 'Exactly one required' clarifies the optional-looking schema, and it explains how date resolves and what happens on ambiguous matches. This meaningfully aids correct invocation beyond the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: delete a period record. It explicitly distinguishes itself from update_cycle, telling the agent it must not use this tool to 'fix' a record. This makes it easy to tell apart from the sibling update_cycle tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: only when the user explicitly asks to remove a specific record. It also gives a clear exclusion and alternative: do not delete to fix a record; use update_cycle instead. The selector instructions further clarify how to invoke it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_injuryA
DestructiveIdempotent
Inspect

Permanently delete an injury entry. Also removes all severity history for that injury.

SELECTOR — pass id if known, or injury (a body part or injury type substring, case-insensitive, e.g. "shoulder") optionally narrowed by date (an injury active on that day). Exactly one of id or injury required. If injury matches more than one entry, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoInjury ID to delete. Alternative to injury.
dateNoOptional, narrows the injury selector to one active on this date (YYYY-MM-DD). Ignored when id is given.
injuryNoAlternative to id: body part or injury type substring, case-insensitive (e.g. "shoulder"). Optionally narrow with date.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this destructive, but the description adds meaningful behavioral context: deletion is permanent, severity history is also removed, and ambiguous injury matches cause an error returning candidate IDs. This goes well beyond the structured hints and clearly sets expectations for side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: effect first, side effect second, then selector rules. Each sentence earns its place and no information is repeated from the schema or annotations. The use of an em-dash separator and explicit rule statements makes it easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations cover safety, the description supplies everything else needed: selection strategy, disambiguation behavior, requiredness despite no schema-required params, and destructive consequences. The tool is fully usable based on this description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is met. The description adds value by explaining the mutual exclusivity between id and injury, the exact-one-required contract, the role of date as a narrowing filter, and the error behavior on ambiguous matches — all beyond the schema's individual property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Permanently delete an injury entry.' The scope is precise, and naming the cascading removal of severity history distinguishes this from update_injury and log_injury. An agent can tell exactly what this tool does and what side effects come with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear operational guidance: how to choose between id and injury, when to use date to narrow, and that exactly one selector is required. It does not explicitly discuss alternatives like update_injury for non-destructive changes, but the selector guidance is detailed enough for practical use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_lab_resultA
DestructiveIdempotent
Inspect

Permanently delete one or more lab results. Use when the user explicitly asks to remove or delete a logged lab result. Never guess a selector.

SELECTOR, pass exactly one of id, date, date+marker, or draw_id:

  • id: deletes a single marker's row.

  • draw_id: deletes every result sharing that draw_id at once, unambiguous by construction. Irreversible.

  • date (optionally narrowed by panel_name): deletes every result from that draw at once, but ONLY when exactly one draw exists on that date — see AMBIGUITY below. Irreversible and can remove many rows in one call. Confirm with the user before a date-scoped delete, especially one not narrowed by panel_name or marker.

  • date + marker (a marker-name substring, case-insensitive, optionally narrowed by panel_name): resolves to and deletes one marker's row, same as id. Errors with candidate IDs if more than one marker on that date matches.

AMBIGUITY: a bare date (optionally + panel_name) selector is rejected, with nothing deleted, if it would match more than one physical draw — an explicit draw_id from one source plus legacy rows with none, two distinct draw_ids, or two differently-named legacy sources on the same day. The error names every draw found; retry with draw_id, marker, or a narrower panel_name.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoLab result ID. Deletes a single marker row. Alternative to date/draw_id.
dateNoCollection date of the draw to delete. Format: YYYY-MM-DD. Deletes every result from that draw (see AMBIGUITY above), or (with marker) one row. Alternative to id/draw_id.
markerNoOptional with date: marker-name substring, case-insensitive (e.g. "LDL"), narrowing the date selector to delete a single marker row instead of the whole draw. Ignored when id or draw_id is given.
draw_idNoOpaque label of your choosing grouping a set of results from one visit, normalized server-side. Deletes every result sharing that label, unambiguous by construction. Alternative to id/date. Ignored when id is given.
panel_nameNoOptional, narrows a date (or date+marker) selector to one panel within that draw (e.g. "Lipid Panel"). Ignored when id is given.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true and idempotentHint=true, but the description adds crucial behavioral context: deletion is irreversible, date-scoped deletes can remove many rows, and ambiguous dates are rejected with nothing deleted and candidate IDs returned. This fully discloses what gets destroyed and how failures behave.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every block earns its place given the destructive and ambiguous nature of the operation. The core action and use condition are front-loaded, selector rules are organized, and the AMBIGUITY section handles the trickiest behavior clearly. No filler or repetition beyond what improves safe usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with five optional parameters and complex selector semantics, the description covers the key decision space: which selector to use, how ambiguity is resolved, what is irreversible, and when to confirm with the user. The output schema exists, so return-value details do not need to be repeated here. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, yet the description adds substantial meaning beyond the schema: pass exactly one selector, marker matching is a case-insensitive substring, draw_id deletes all rows sharing that label, and panel_name narrows date selectors. It also explains resolution and rejection behavior that the schema alone cannot convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Permanently delete one or more lab results.' It clearly distinguishes this destructive action from the sibling update_lab_result and the many list/log tools, and even states the triggering user intent. The selector breakdown reinforces exactly what the tool operates on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: when the user asks to remove or delete a logged lab result. It also provides strong exclusions and guardrails: 'Never guess a selector,' warns about ambiguity, and instructs confirming with the user before date-scoped deletes. This goes well beyond implied usage and actively prevents misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_mealA
DestructiveIdempotent
Inspect

Permanently delete a meal entry. Use when the user explicitly asks to remove or delete a logged meal.

FIND THE MEAL: pass id if already known. Otherwise pass date (YYYY-MM-DD, defaults to today) and, only if more than one meal was logged that day, name (a substring of the food description, case-insensitive) to narrow it down. This action is irreversible — a match that isn't exactly one meal returns an error explaining why, with nothing deleted; retry with id or a narrower name, never guess.

HYDRATION: any assistant hydration events linked to this meal are deleted by the database in the same food-row delete. Do not issue a separate hydration delete.

CAFFEINE: caffeine sidecars linked to this meal are deleted by the database in the same food-row delete. Do not call a separate caffeine delete tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoMeal ID, if already known. Alternative to date + name — see FIND THE MEAL above.
dateNoDate the meal was logged. Format: YYYY-MM-DD. Used with name to find the meal when id is omitted; defaults to today if id and date are both omitted.
nameNoSubstring of the food description (case-insensitive) to disambiguate multiple meals on the same date. Only used when id is omitted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already flag the operation as destructive, but the description adds meaningful behavioral detail: the action is irreversible, ambiguous matches return an error with nothing deleted, retries should use id or a narrower name, and linked hydration/caffeine events are removed automatically. This goes well beyond the annotation metadata and directly shapes safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but every section earns its place. The intent is front-loaded in the first sentence, and the FIND THE MEAL / HYDRATION / CAFFEINE sections are clearly labeled and logically separated, making the content easy for an agent to parse and apply.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, ambiguity-prone delete operation, this description covers the entire invocation context: how to identify the target, what happens on failure, retry guidance, and automatic cascade behavior for linked records. An output schema exists to cover return values, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already has a strong 100% coverage with per-parameter descriptions, the tool description enriches those semantics with the decision tree: id takes priority, date defaults to today, and name is only a disambiguator used when multiple meals were logged that day. It even warns that an ambiguous match will error and the agent should never guess, which is genuinely useful parameter-level guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Permanently delete a meal entry,' naming the exact verb and resource. It also states the precise user-intent trigger ('when the user explicitly asks to remove or delete a logged meal'), which clearly differentiates it from update_meal and the other meal-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use condition and then expands into a structured disambiguation protocol: use id if known, otherwise date, and add name only when multiple meals exist on that date. It also provides strong exclusions, telling the agent not to issue separate hydration or caffeine delete calls because the database handles those cascades.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_recovery_sessionA
DestructiveIdempotent
Inspect

Permanently delete a recovery session log entry. This action is irreversible. If the user's intent is ambiguous, ask which session to remove.

SELECTOR — pass id if known, or session_date (+ optional session_category to narrow) to resolve it. Exactly one of id or session_date required. If it matches more than one session, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRecovery session ID. Alternative to session_date.
session_dateNoAlternative to id: the date (YYYY-MM-DD) the session was logged on. Optionally narrow with session_category.
session_categoryNoOptional, narrows session_date to one category when more than one session shares that date. Ignored when id is given.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal destructiveness and idempotency, and the description adds irreversibility, ambiguity-handling guidance, and the error behavior with candidate IDs. The description and annotations align, and the additional context meaningfully exceeds what annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the most important facts (permanence, irreversibility, ambiguity handling) followed by a compact selector block. Every sentence adds operational value and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's destructive nature and the non-obvious selector logic, the description fully equips an agent to call it correctly: it covers user clarification, parameter selection, ambiguity resolution, and error recovery. The output schema and annotations cover remaining details, so nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds crucial cross-parameter semantics: the selector relationship between id and session_date, the optional narrowing role of session_category, that session_category is ignored when id is given, and the exactly-one-required constraint. This goes well beyond the baseline schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Permanently delete') and a specific resource ('a recovery session log entry'), clearly distinguishing it from sibling delete/update/log tools. The title annotation adds a parallel label, and the resource is unambiguous against the large set of sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs the agent on when to ask the user for clarification, how to resolve the target session using either id or session_date, and what to do if multiple sessions match. It also clearly states that exactly one of id or session_date is required, which is not reflected in the schema's required parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_runA
DestructiveIdempotent
Inspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Delete a run. Use when the user wants to remove a run entry. ASK for confirmation if the user's intent is ambiguous.

SELECTOR — pass id if known, or date (+ optional distance_mi to narrow, matched approximately within 0.25 mi) to resolve it. Exactly one of id or date required. If it matches more than one run, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRun UUID. Alternative to date.
dateNoAlternative to id: the date (YYYY-MM-DD) the run was logged on. Optionally narrow with distance_mi.
distance_miNoOptional, narrows date to a run within 0.25 mi of this value when more than one run shares that date. Ignored when id is given. In mi, or km with input_distance_unit set. See UNIT INPUTS.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already supply destructiveHint and readOnlyHint; the description adds deletion-safety behavior such as the id-or-date selector, approximate 0.25 mi matching, an error-with-candidate-IDs failure mode, and a confirmation rule. This is valuable context beyond the annotations and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The useful selector guidance is buried after four lines of UNIT INPUTS instructions for parameters that do not exist on delete_run. This is not front-loaded and wastes tokens; the whole message could be reduced to two focused sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema and annotations, the description covers the full calling flow: selector, one-of requirement, confirmation when ambiguous, and behavior on multiple matches. The irrelevant unit preamble subtracts from comprehensibility, but no critical deletion-handling detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage the baseline is 3; the description adds the essential one-of constraint ('Exactly one of id or date required') and explains id-first selection and distance narrowing. The unrelated UNIT INPUTS block is confusing noise, but the selector paragraph genuinely augments the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Core statement 'Delete a run' and 'Use when the user wants to remove a run entry' name a clear verb and resource, distinguishing delete_run from sibling delete_* tools by target. But the opening UNIT INPUTS block references parameters like value/alternate_unit/canonical_unit/precedence that are not in the schema, which muddies the first impression.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it when the user wants to remove a run entry and tells the agent to ASK for confirmation if intent is ambiguous. It does not name specific alternative tools or exclusion cases, but the run-targeted scope is enough to route the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_wellbeingA
DestructiveIdempotent
Inspect

Permanently delete a wellbeing entry.

SELECTOR — pass id if known, or date (any day within the entry's period) to resolve it. Exactly one of id or date required. If date matches more than one entry, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoWellbeing entry ID to delete. Alternative to date.
dateNoAlternative to id: a date (YYYY-MM-DD) that falls within the entry's period. Resolves only when exactly one entry matches.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reinforces the destructive nature ('Permanently delete') and adds transparency about the idempotent behavior (error with candidate IDs on ambiguous date). It goes beyond the annotations by detailing the exact failure mode and retry guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet complete, comprising two sentences and a selector note. Every word contributes to usage clarity, with no redundant or vague statements.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete operation, the description covers all necessary context: target, selector, exclusivity, and error handling. The lack of output schema details is acceptable since a delete confirmation is standard and does not hinder usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (id and date) are fully described in the schema and the description, including the semantics of date ('any day within the entry's period') and the exclusivity requirement. High schema coverage combined with clear prose makes parameter usage unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Permanently delete') and the target ('a wellbeing entry'), with no ambiguity. It effectively distinguishes this from sibling tools like update or log by specifying deletion and the selector mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains when to use the tool by providing a selector rule: pass id or date, exactly one required. It also describes the error condition when the date matches multiple entries, giving clear guidance for resolution.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workoutA
DestructiveIdempotent
Inspect

Permanently delete a workout session and all its exercises and sets. Use when the user wants to remove a logged workout entirely.

FIND THE SESSION: call directly, no preliminary list for an ID. Pass session_id if already known. Otherwise pass session_date (YYYY-MM-DD, defaults to today) and, only if more than one session was logged that day, name (a substring of the workout's focus/type, e.g. "Push" or "Leg Day", case-insensitive) to narrow it down. This action is irreversible and removes the session, all supersets, and all sets — a match that isn't exactly one session returns an error explaining why, with nothing deleted; retry with session_id or a narrower name, never guess.

SAVED WORKOUTS: pass saved_workout_id to delete a reusable Saved Workout instead of completed workout history.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSubstring of the workout's focus/type (e.g. "Push", "Leg Day"), case-insensitive, to disambiguate multiple sessions on the same date. Only used when session_id is omitted.
session_idNoPositive session ID returned by a workout tool. If unknown, omit it and use session_date + name; never guess.
session_dateNoDate the session was logged. Format: YYYY-MM-DD. Used with name to find the session when session_id is omitted; defaults to today if both are omitted.
saved_workout_idNoSaved Workout ID to delete. When present, deletes the reusable prescription and does not touch completed workout history.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark destructiveHint=true and idempotentHint=true, but the description goes well beyond: it states the action is irreversible, removes the session, all supersets, and all sets, and explains that a non-exact match returns an error with nothing deleted. It also clarifies that saved_workout_id does not touch completed workout history, providing full behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections: the core purpose, 'FIND THE SESSION' instructions, and 'SAVED WORKOUTS' alternative. It front-loads the primary use case, then provides necessary lookup steps, and every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with multiple targeting methods, the description covers irreversibility, error behavior, parameter interactions, and the alternative mode for saved workouts. An output schema exists, so return values are presumably documented elsewhere. No critical information is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter already has a description. The tool description adds crucial usage logic: session_id is preferred and never to be guessed, session_date defaults to today, name is a case-insensitive substring for disambiguation, and saved_workout_id switches the operation to delete a reusable prescription. This clarifies parameter interactions beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool permanently deletes a workout session and all its exercises and sets, which is a specific verb+resource. It also distinguishes between deleting a logged session and a reusable Saved Workout, making it distinct from sibling delete_* tools like delete_cycle or delete_meal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Use when the user wants to remove a logged workout entirely' and provides a detailed decision path: pass session_id if known, otherwise session_date and name for disambiguation, and saved_workout_id for the reusable prescription. It also warns against guessing and directs retries with a narrower name or session_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_app_guide_sectionA
Read-onlyIdempotent
Inspect

Look up customer-facing Wellness Project product knowledge. Use it for app navigation, visible features, integrations, metric meanings, logging methods, permissions/sync questions, or troubleshooting.

Wellness Project-specific answers are closed-book: answer them from this tool's result, never from memory. The result states its own grounding rules. For how-to questions, select the relevant feature topic; a matching published video link is included when one exists.

For recognized exercise names or whether an exercise is in the app's exercise library, use list_exercises rather than guessing or expecting the app guide to enumerate exercises.

For cost, price, Free vs Pro, Founding Member, upgrading, or the 3-analysis limit, use topic=pricing. Reproduce exactly one returned pricing message verbatim, include the supplied subscription link, and add no other pricing detail. The app attaches its own upgrade button to that message when one applies; never type out a button, chip, link markup, or call-to-action of your own.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYesChoose the narrowest relevant area. pages_dashboard = Dashboard, Me, Settings, AI Assistants. pages_training = Fitness, workouts, running, cycling, heart, recovery. pages_nutrition = nutrition, hydration, caffeine, sleep, body, wellbeing, labs. personas = AI specialists. logging = ways to log/correct data. photos = meal photos, labels, barcodes. wearables = integrations, connection requirements, known provider limits. metrics = user-facing metric meanings. goals = goals/targets. challenges = friend challenges. privacy = account controls/policy links. sync_details = permissions, history, missing-data behavior. troubleshooting = unexpected behavior/recovery. pricing = approved pricing copy.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/no-open-world, and the description adds substantial behavior beyond them: the result carries its own grounding rules, published video links may be included, pricing answers must be reproduced verbatim with exactly one message and the supplied link, and the app attaches its own upgrade button so the agent must never emit its own CTA markup. That is unusually rich operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four paragraphs, but each is front-loaded on its rule (purpose, closed-book, exercise routing, pricing) and every sentence carries a distinct constraint. Dense rather than padded, though the length is at the upper edge of what an agent can absorb at selection time.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation here, and the description covers the remaining gaps: which topics to pick, how to handle pricing output, the grounding requirement, and the sibling alternatives. Nothing needed to invoke or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the enum already defines every topic value in detail, so the baseline is 3. The description adds the routing principle ('select the relevant feature topic', 'narrowest relevant area') and the pricing-topic constraint, but mostly reinforces what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Look up customer-facing Wellness Project product knowledge,' and immediately enumerates the domains it covers (navigation, features, integrations, metrics, logging, permissions/sync, troubleshooting). It also explicitly separates itself from a sibling by naming list_exercises as the tool for exercise-library questions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing rules: use list_exercises for exercise names/library membership, use topic=pricing for all cost/upgrade/limit questions, and select the relevant feature topic for how-to questions. It also states the closed-book constraint (answer from this result, never from memory), which is exactly the kind of when/when-not guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_exercise_historyA
Read-onlyIdempotent
Inspect

Look up everything the user has done for ONE exercise: all-time PR plus recent performance, across many sessions.

USE FOR:

  • PR lookups — "what's my bench PR?", "have I ever squatted 315?". Returns the est. 1RM PR and the exact set it came from (date, weight, reps, RPE, banded vs unbanded, superset siblings, notes), plus rep-range bests (1RM/3RM/5RM/10RM). Banded and unbanded PRs are shown side-by-side when both exists.

  • Recent-activity questions — "how has my squat been lately?", "when did I last deadlift?". Returns the most recent N sessions containing the exercise, formatted like get_workout.

  • Trend questions — "am I getting stronger on incline DB press?". Includes a one-line delta of current best vs ~30-90 days ago.

NOT for a full session (every exercise in one workout — use get_workout) or a date-window list regardless of exercise (use list_workouts).

INFER — do not ask: exercise_name (take the user's words; resolves to canonical, or says so if never logged), recent_limit (default 10 sessions), since_date (optional — narrows only the Recent block; the PR is always all-time).

ParametersJSON Schema
NameRequiredDescriptionDefault
since_dateNoOptional YYYY-MM-DD lower bound for the Recent block. Does not affect the PR section, which is always all-time.
recent_limitNoHow many recent sessions containing this exercise to surface. Optional — default 10, capped at 50.
exercise_nameYesExercise to look up. Required. Free-text — the tool resolves to canonical (e.g. "bench" → "Bench Press").

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, non-destructive, closed-world. The description adds genuine behavioral detail beyond them: that the PR is always all-time regardless of since_date, that banded and unbanded PRs appear side-by-side, rep-range bests, and that since_date narrows only the Recent block. Some return-content detail overlaps the output schema, keeping this just short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Strongly front-loaded and organized under USE FOR / NOT for / INFER headers, with the scoping constraint stated first. Slightly long overall, but no sentence is clearly disposable, so a 4 rather than a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with rich annotations, a full output schema, and three fully documented parameters, nothing an agent needs to invoke it correctly is missing — selection criteria, argument inference, and the since_date/PR interaction are all covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns extra: it clarifies exercise_name resolution (free-text resolved to canonical, or reports if never logged) and reinforces that since_date affects only the Recent block while the PR stays all-time. This adds meaning the schema states only in passing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a precise verb and resource — look up everything for ONE exercise, all-time PR plus recent performance across sessions. It explicitly contrasts against the two most confusable siblings (get_workout for a full session, list_workouts for a date-window list regardless of exercise), so an agent can route to it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The USE FOR block enumerates three distinct intents with example user phrasings (PR lookups, recent-activity, trend), and the NOT-for section names the exact alternative tools and the condition that selects them. The INFER block further tells the agent which arguments to derive rather than ask for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrient_contributorsA
Read-onlyIdempotent
Inspect

Top logged foods and supplements contributing to ONE nutrient over a day or short range, highest amount first. Use for "what gave me most of my sodium today" or "what's driving my potassium this week".

INFER -- do not ask:

  • date: default to today (resolved in the user's own timezone)

  • days: default to 1 (just date); set higher for a range ending on date

  • limit: default to 10

nutrient is required and accepts common aliases (e.g. "b12", "carbs", "fibre").

FORMAT: default 'compact' replies with id (this nutrient's number in get_nutrient_summary's nutrient dictionary, whose unit then applies) instead of key/unit, and drops the per-contributor unit field it would otherwise repeat on every row. A nutrient outside that dictionary has no id, so key/unit are used regardless of format. 'verbose' always uses key/unit and keeps unit on every contributor, as before.

DATA COMPLETENESS: totals are sums over the logged items that carry a value for that nutrient; coverage tells how many did. Say "based on foods with X data" when coverage is partial; never call a low intake a deficiency; intake is not a lab result; supplement vs food split is reported separately. A daily supplement counts on every day it's active by default -- the user isn't expected to check it off -- unless a day was explicitly marked not taken, so its share of a total can include days with no check-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoEnd date (or the only date if days=1). Format: YYYY-MM-DD. Optional -- omit for today.
daysNoHow many days, ending on date. Optional -- default 1.
limitNoMax contributors to return. Optional -- default 10.
formatNocompact (default) packs values by nutrient id, see FORMAT above. verbose spells out each nutrient's name and unit.
nutrientYesNutrient name or alias, e.g. "vitamin_d", "calcium", "b12". Required.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safe read-only/idempotent profile, and the description adds substantial behavior beyond them: default resolution rules, the compact-vs-verbose output contract, coverage semantics ('based on foods with X data'), and the non-obvious rule that an active daily supplement counts on every day unless marked not taken. These are exactly the traits an agent cannot infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and is clearly sectioned (INFER / FORMAT / DATA COMPLETENESS), so scanning is easy. It is on the long side and the FORMAT block partially restates the schema's format field, but the detail is load-bearing rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be re-explained, and the description nonetheless covers defaults, alias handling, output-format tradeoffs, and the statistical caveats (coverage, not a lab result, supplement-vs-food split). Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: nutrient accepts common aliases (b12, carbs, fibre), format semantics explain how output changes per value, and default/omission behavior for date, days, and limit is spelled out. This exceeds what the schema alone conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: top logged foods/supplements contributing to ONE nutrient over a day or range, ordered highest first. This clearly differentiates it from aggregation siblings like get_nutrient_summary and get_nutrient_history, and the example phrasings ('what gave me most of my sodium today') remove any ambiguity about what it returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage triggers via quoted example questions, and the defaults section tells the agent how to proceed without asking. However it never explicitly names alternatives (get_nutrient_summary, get_nutrient_history) or states when NOT to use this tool, so the routing guidance is context-only rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrient_historyA
Read-onlyIdempotent
Inspect

Day-by-day history for ONE nutrient over a range, with its target, how many days met it, and how many days went over the upper limit. Use for "show my vitamin D over the last month" or "am I usually over on sodium". Only dates with logged data are returned.

INFER -- do not ask:

  • days: default to 30

  • end_date: default to today (resolved in the user's own timezone)

nutrient is required and accepts common aliases (e.g. "b12", "carbs", "fibre").

FORMAT: default 'compact' replies with id (this nutrient's number in get_nutrient_summary's nutrient dictionary, whose unit then applies) instead of key/unit, and drops the target's redundant unit field. A nutrient outside that dictionary has no id, so key/unit are used regardless of format. 'verbose' always uses key/unit, as before.

DATA COMPLETENESS: totals are sums over the logged items that carry a value for that nutrient; coverage tells how many did. Say "based on foods with X data" when coverage is partial; never call a low intake a deficiency; intake is not a lab result; supplement vs food split is reported separately. A daily supplement counts on every day it's active by default -- the user isn't expected to check it off -- unless a day was explicitly marked not taken, so its share of a total can include days with no check-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoHow many days of history, ending on end_date. Optional -- default 30.
formatNocompact (default) packs values by nutrient id, see FORMAT above. verbose spells out each nutrient's name and unit.
end_dateNoLast date of the window. Format: YYYY-MM-DD. Optional -- omit for today.
nutrientYesNutrient name or alias, e.g. "vitamin_d", "calcium", "b12". Required.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial behavior beyond them: only dates with logged data are returned, the compact/verbose output distinction, and the non-obvious supplement rule (a daily supplement counts on every active day unless explicitly marked not taken). It also warns against misinterpreting intake as a lab result or a deficiency, which materially affects how the agent should report results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, and the INFER/FORMAT/DATA COMPLETENESS sections are clearly delimited. It is long, but the length is largely justified by the tool's complexity (two output modes, alias handling, supplement counting); some of the FORMAT prose could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with an output schema, four parameters, and non-obvious data semantics, the description covers everything an agent needs: what is returned, which dates are included, default inference, alias support, and the caveats required to report intake honestly. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it confirms nutrient accepts common aliases ('b12', 'carbs', 'fibre'), restates the days/end_date defaults in user-facing terms, and explains what 'format' actually changes in the payload. This goes beyond the schema's field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Day-by-day history for ONE nutrient over a range, with its target, how many days met it, and how many days went over the upper limit.' It also implicitly distinguishes itself from the aggregate sibling get_nutrient_summary (which it references for the nutrient dictionary) and from get_nutrient_contributors. An agent can tell exactly what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage examples ('show my vitamin D over the last month', 'am I usually over on sodium') and explicit inference rules for days/end_date defaults, which is strong context. It does not, however, explicitly state when to prefer get_nutrient_summary or get_nutrient_contributors over this tool, so the routing is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_nutrient_summaryA
Read-onlyIdempotent
Inspect

Summarize logged micronutrient (and macro) intake for one day or a short range, with each nutrient's target and coverage. Use for "how's my vitamin D / sodium / iron this week", or a general "how are my micronutrients looking" (default: every tracked nutrient except amino acids, for today).

INFER -- do not ask:

  • date: default to today (resolved in the user's own timezone)

  • days: default to 1 (just date); for a range, set days to the number of days ending on date (e.g. "this week" -> days=7)

  • nutrients: default to every nutrient with logged data this period, excluding amino acids (those are rarely tracked and would crowd out the rest). Ask for specific nutrients by name (aliases like "b12", "carbs", "fibre" resolve) only to see amino acids or a nutrient with no data yet.

Returns up to 60 nutrients. days=1 reports that day's totals; days>1 reports the range average per nutrient plus how many days were complete.

FORMAT: default 'compact' packs each panel (known/food/supplement/target_value/target_upper_limit/target_percent) as one "id:value;id:value" string, ids ascending, values to at most 3 significant digits, units are the dictionary's below -- never repeat a nutrient's name or unit back to the user from these strings, resolve the id first. coverage/items_total/items_with_value/estimated_items/target_source/target_kind stay small id-keyed records. A nutrient outside the dictionary (amino acids, Cronometer extras) has no id, so its value is dropped and its name listed once under omitted; call again with format:'verbose' to see it. 'verbose' returns one full object per nutrient instead, as before.

NUTRIENT ID DICTIONARY (id=name, unit per the id; also used by get_nutrient_history/get_nutrient_contributors's id field): 1=fiber_g 2=sugar_g 3=saturated_fat_g 4=monounsaturated_fat_g 5=polyunsaturated_fat_g 6=trans_fat_g 7=cholesterol_mg 8=sodium_mg 9=potassium_mg 10=calcium_mg 11=iron_mg 12=magnesium_mg 13=phosphorus_mg 14=zinc_mg 15=copper_mg 16=manganese_mg 17=selenium_ug 18=chloride_mg 19=chromium_ug 20=iodine_ug 21=molybdenum_ug 22=vitamin_a_ug 23=vitamin_c_mg 24=vitamin_d_ug 25=vitamin_e_mg 26=vitamin_k_ug 27=thiamin_b1_mg 28=riboflavin_b2_mg 29=niacin_b3_mg 30=pantothenic_acid_b5_mg 31=vitamin_b6_mg 32=biotin_b7_ug 33=folate_b9_ug 34=folic_acid_ug 35=vitamin_b12_ug 36=choline_mg 37=omega3_g 38=omega6_g 39=caffeine_mg 40=water_g 41=starch_g 42=added_sugar_g 43=total_unsaturated_fat_g 44=fluoride_mg

DATA COMPLETENESS: totals are sums over the logged items that carry a value for that nutrient; coverage tells how many did. Say "based on foods with X data" when coverage is partial; never call a low intake a deficiency; intake is not a lab result; supplement vs food split is reported separately. A daily supplement counts on every day it's active by default -- the user isn't expected to check it off -- unless a day was explicitly marked not taken, so its share of a total can include days with no check-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoEnd date (or the only date if days=1). Format: YYYY-MM-DD. Optional -- omit for today.
daysNoHow many days, ending on date. Optional -- default 1.
formatNocompact (default) packs values by nutrient id, see FORMAT above. verbose spells out each nutrient's name and unit.
nutrientsNoNutrient names or aliases to restrict to (e.g. ["vitamin_d","calcium"]). Optional -- default is every nutrient with data except amino acids.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantial extra behavior: coverage semantics ('totals are sums over items that carry a value'), the instruction never to call low intake a deficiency, food-vs-supplement split reporting, and the subtle default that an active supplement counts on every day unless explicitly marked not taken. These are genuine behavioral facts not derivable from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded (purpose first, then INFER/FORMAT/DICTIONARY/DATA sections) and everything is labeled, but the 44-entry ID dictionary dominates the length and is payload reference data that inflates the definition. Defensible because compact output is unreadable without it, but it is a lot of text for a tool definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, yet the description still documents the two format modes and what each returns. Combined with the completeness caveats and default rules, an agent has everything needed to call this correctly and interpret the results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), and the description goes further by spelling out resolution rules: timezone-local 'today', 'this week' -> days=7, alias handling ('b12', 'carbs', 'fibre'), and that days>1 returns a range average. Much of the date/days/format meaning still overlaps the schema text, so it is strong rather than exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Summarize logged micronutrient (and macro) intake for one day or a short range, with each nutrient's target and coverage.' It also distinguishes itself from the related history/contributor tools by referencing them in the ID dictionary context, so an agent can tell what this tool uniquely returns (a period summary with targets) versus a history or contributor breakdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering phrases ('how's my vitamin D / sodium / iron this week', 'how are my micronutrients looking') and resolves the common inference cases (date=today, days=1, nutrients=all-but-amino-acids) with explicit 'do not ask' instructions. It does not, however, explicitly route against the sibling get_nutrient_history/get_nutrient_contributors for when a summary is preferable to a trend or breakdown, leaving one routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workoutA
Read-onlyIdempotent
Inspect

Retrieve full detail of a workout session: exercises, sets, reps, weights, superset groupings, heart points, notes, and NSI scoring at every grain. Use for detailed questions about a past workout, reviewing training before recommendations, confirming what was logged, or comparing a session to population strength standards.

NSI: session NSI/rating in the header; per-exercise NSI (max set NSI), rating, est. 1RM, and the population_1rm_lb/population_reps benchmark it was measured against; per-set NSI and est. 1RM to see which set drove the exercise score.

EQUIPMENT: shown per exercise when every set shares a tag, else per set; missing means untagged. A wrong or missing tag on a dumbbell exercise silently halves or doubles its NSI score — fix it via update_workout's set_updates or add_exercises equipment field.

session_id: reuse an ID returned by a workout tool; unknown: list_workouts. Never guess.

SAVED WORKOUTS: pass saved_workout_id to read a reusable Saved Workout prescription. Do not combine it with session_id/session_date.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID returned by a workout tool.
saved_workout_idNoSaved Workout ID from list_workouts(saved_workouts=true). When present, returns the reusable prescription instead of a completed workout session.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation read-only, idempotent, and non-destructive. The description goes well beyond this by explaining NSI scoring behavior at session, exercise, and set grains, disclosing that missing/wrong equipment tags can silently alter NSI scores, and explaining how saved workout prescriptions differ from completed sessions. This is rich, honest behavioral context that an agent would otherwise have no way to infer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but well structured with labeled sections (NSI, EQUIPMENT, SAVED WORKOUTS) and front-loads the core purpose in the first sentence. Some detail, such as the equipment tag consequence, could arguably live in update_workout, but it is relevant here because it explains how to interpret the retrieved data. It is organized enough to remain navigable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, rich output schema, and clear annotations, the description covers everything an agent needs: how to retrieve session details, how to interpret NSI and equipment fields, how to handle saved workouts, and how to recover from unknown IDs. There are no obvious gaps that would prevent correct invocation or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful guidance beyond the schema: session IDs should be reused from workout tool results, unknown IDs should be resolved via list_workouts, saved_workout_id returns a prescription rather than a logged session, and the two IDs should not be combined. This materially helps an agent choose and populate the parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific verb and resource: 'Retrieve full detail of a workout session,' and enumerates the included data fields. It differentiates itself from list_workouts and update_workout by naming them, though it does not explicitly distinguish itself from the sibling show_workout, leaving some ambiguity about how the two compare.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: for detailed questions about a past workout, reviewing training, confirming logs, or comparing to population standards. It also gives clear routing guidance for unknown session IDs ('list_workouts. Never guess') and for saved workouts, including a warning not to combine saved_workout_id with session_id. It does not mention show_workout as an alternative, but the provided context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_blog_postsA
Read-onlyIdempotent
Inspect

Search the public Crew Blog at /blog for advisor-authored daily posts. Only call when the user explicitly asks about the blog or what an advisor has written; don't volunteer posts in normal conversation.

Returns each matching post's slug, title, summary, advisor name, and date. Link a post inline as /blog/.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional — number of posts to return. Default 10, max 30.
queryNoOptional — substring filter applied to title and summary (case-insensitive).
advisor_slugNoOptional — filter to one advisor (e.g. "nutritionist" for Casey Mills).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds valuable behavioral context beyond that: it is a search over public content, returns specific fields, and prescribes how results should be linked inline. This is more than the annotations alone provide, though it does not cover edge cases like empty results or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well organized: purpose first, usage constraint second, return format and link handling last. Every sentence adds useful information, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple, optional-parameter read tool with a rich schema, output schema, and full annotations. The description covers when to use it, what it returns, and how to format links, so an agent has everything needed to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented with types and descriptions. The tool description adds no new parameter semantics beyond the general search framing, which matches the baseline for fully covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: searching the public Crew Blog at /blog for advisor-authored posts. It clearly identifies what the tool returns and the /blog/<slug> link format, making its purpose unambiguous and distinguishable from the many list/show siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: only when the user explicitly asks about the blog or what an advisor has written. It also provides a clear when-not-to-use instruction: don't volunteer posts in normal conversation. No sibling tool covers this same domain, so naming an alternative is unnecessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_body_metricsA
Read-onlyIdempotent
Inspect

List body composition entries within a date range. Use when the user asks about their weight history, body fat trend, or any body metrics over time.

Maximum range: 31 days per call. For longer periods, make multiple calls with sequential date ranges.

INFER — do not ask:

  • start_date: default to 30 days ago

  • end_date: default to today

BMI in the output is derived from the user's canonical height and that day's resolved weight -- do not recompute it yourself.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 30 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint false), the description adds behavioral context: it states BMI in the output is derived from the user's canonical height and that day's resolved weight, and instructs the agent not to recompute it. It also explains default parameter behavior (infer 30 days ago to today). These details help the agent understand side effects and derived data, exceeding annotation-only transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It covers purpose, usage conditions, range constraints, and behavioral notes without unnecessary verbosity. Each sentence serves a distinct purpose: identifying the resource, stating when to use, noting the 31-day limit, and explaining parameter inference. No fluff or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description need not explain return values. It fully covers input parameters, defaults, usage context, range limits, and derived data behavior. The description provides all necessary information for an agent to invoke the tool correctly in various scenarios, making it contextually complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema descriptions for start_date and end_date specify format (YYYY-MM-DD) and defaults ('30 days ago', 'today'). The description reinforces these defaults and adds the 'INFER — do not ask' guidance, making the parameter semantics fully clear. Schema coverage is 100%, and no enums exist, so the description effectively complements the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists body composition entries within a date range. It specifies the resource ('body composition entries') and scope ('date range'), and the context signals show sibling tools include similar list tools, enabling an agent to distinguish this from list_workouts, list_meals, etc. The verb 'list' and explicit resource make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool ('when the user asks about their weight history, body fat trend, or any body metrics over time'). It also provides guidance for handling longer periods by making multiple calls with sequential date ranges, and instructs the agent to infer default parameters rather than asking the user. This gives clear, actionable usage instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_caffeineA
Read-onlyIdempotent
Inspect

List caffeine doses by local date with daily totals. Use for caffeine history or today's total. Supports up to 366 days per call; daily totals cover every dose in the range, while the dose list shows the 200 most recent and says how many it omitted.

INFER: from/to default to the last 14 local calendar days through today.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoLast local date, YYYY-MM-DD. Default: today.
fromNoFirst local date, YYYY-MM-DD. Default: 13 days before today.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, the description discloses important behavioral details: the 366-day limit, daily totals covering all doses, the dose list capping at 200 most recent, and that omitted doses are reported. It also explains default date behavior via the INFER note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core purpose, and uses a clear INFER section for defaults. No sentence is wasted; the behavioral limits and use cases are presented efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and comprehensive annotations, the description covers all necessary operational context: purpose, defaults, limits, and truncation behavior. An agent has enough information to invoke the tool correctly without further inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter already has a clear description. The tool description adds value by explaining the default range behavior ('from/to default to the last 14 local calendar days through today') and by framing parameters in terms of local calendar dates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List caffeine doses by local date with daily totals.' It also clarifies its intended use cases ('caffeine history or today's total'), making it distinct from other list tools like list_sleep or list_meals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use for caffeine history or today's total,' which gives clear usage context. It does not mention alternatives or when not to use it, but the use cases are specific enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_cycleA
Read-onlyIdempotent
Inspect

List the user's logged period records within a date range. update_cycle and delete_cycle can resolve a record on their own from date, so this is no longer required before either — use it when the user asks about their cycle history, or get an id after an ambiguous update_cycle/delete_cycle match.

INFER — do not ask:

  • start_date: default to 6 months ago

  • end_date: default to today

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 6 months ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context beyond annotations: it clarifies date-range scoping, explains the relationship to update/delete operations, and instructs the agent to infer defaults rather than ask the user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose. The sibling differentiation, usage condition, and default-inference instruction each earn their place without redundant filler. It is structured for quick scanning by an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with two optional parameters, full schema coverage, a safe read-only annotation profile, and an output schema, the description provides everything needed to invoke it correctly. It also covers the relevant edge case of ambiguous update/delete matches, making it contextually complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both start_date and end_date already documented including format and defaults. The description mostly repeats these defaults in the INFER block. It adds the 'do not ask' operational instruction, but that is not new parameter semantic meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List the user's logged period records within a date range.' It also explicitly differentiates this tool from update_cycle and delete_cycle, so an agent can tell exactly what list_cycle is for relative to its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: use it when the user asks about cycle history, or to get an id after an ambiguous update/delete match. It also explicitly says the tool is no longer required before update_cycle/delete_cycle because those resolve records by date, which prevents unnecessary calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_exercisesA
Read-onlyIdempotent
Inspect

Returns all canonical exercise names from the exercise library, grouped by muscle group. lookup: before logging, adding or renaming exercises when canonical names are not already known. Metadata/set-value edits: no catalog lookup. Canonical names ensure proper exercise tracking and NSI score calculation.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the grouping behavior and the downstream consequence ('Canonical names ensure proper exercise tracking and NSI score calculation'), which is useful context. It does not describe pagination or ordering, but for a zero-parameter catalog read this is a reasonable level of disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all informative and front-loaded with the core behavior. The second sentence is a compact usage rule, and the third explains the rationale. Minor redundancy exists between 'canonical exercise names' and the final sentence, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only, idempotent catalog tool with an output schema, the description covers the essential decision points: what it returns, when to call it, and when not to. It does not detail the output schema structure, but the output schema itself carries that burden. The only minor gap is not naming a specific sibling alternative, though the usage rule makes the boundary clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics to rely on. The description compensates by clarifying what the returned data represents (canonical exercise names grouped by muscle group) and why that matters. There is no parameter ambiguity to resolve.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns'), a specific resource ('canonical exercise names from the exercise library'), and a grouping dimension ('by muscle group'). It clearly distinguishes this from the many list_* siblings by focusing on the exercise catalog rather than user logs or metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it: 'before logging, adding or renaming exercises when canonical names are not already known.' It also gives an exclusion: 'Metadata/set-value edits: no catalog lookup.' This is strong routing guidance relative to the large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_goalsA
Read-onlyIdempotent
Inspect

Analyze or show the user's current and past goals. Returns active/paused formal goals, completed/historical formal goals, and current standard targets.

Use this when the user asks what goals they have or asks to review/analyze their goals. Pass include_capabilities: true ONLY when the user asks what kinds of goals Wellness Project supports; it appends the full catalog of goal types and their inputs, which is large. Do not call this tool merely to obtain an ID before create_goal or update_goal; those write tools resolve current goals themselves.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_capabilitiesNoAppend the supported goal types and their inputs. Default false. Only set this when the user is asking what kinds of goals the app supports.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds meaningful behavioral context beyond that: it explains that include_capabilities=true appends a 'full catalog' that is 'large', informing the agent of an output-size tradeoff. It also clarifies that write tools resolve current goals themselves, preventing unnecessary calls. Minor gaps remain (e.g., pagination or exact response shape), but the output schema exists to cover structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each earning its place: the first states purpose and output, the second gives usage triggers, and the third covers negative usage and the parameter caveat. It is front-loaded with the core purpose and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter read-only tool with a full output schema and comprehensive annotations, the description covers everything an agent needs: when to use, what it returns, when to pass the parameter, and when not to use it. No critical context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by tying include_capabilities to a specific user intent ('what kinds of goals Wellness Project supports') and explicitly warning that the appended catalog is 'large'. This helps the agent decide when to set the flag, which is more actionable than the schema's generic phrasing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Analyze or show the user's current and past goals'), names the resource ('user's goals'), and enumerates the exact return categories: active/paused formal goals, completed/historical formal goals, and current standard targets. It also implicitly distinguishes itself from sibling write tools by noting it should not be used to fetch an ID for create_goal or update_goal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage conditions are provided: use when the user asks about their goals or wants a review/analysis. The description also gives a clear when-not-to-use rule ('Do not call this tool merely to obtain an ID before create_goal or update_goal') and a precise condition for setting include_capabilities. This fully routes an agent to the correct tool and parameter choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_hydrationA
Read-onlyIdempotent
Inspect

Review hydration events and stored effective hydration totals. Use when the user explicitly asks about hydration history or fluid intake. Hydration tracking must already be enabled in Settings. Maximum range 31 days. Defaults to the last 7 days.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd date in YYYY-MM-DD format. Default today.
start_dateNoStart date in YYYY-MM-DD format. Default 6 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description only needs to add extra context. It adds meaningful behavioral details: maximum range of 31 days, defaulting to the last 7 days, and the Settings prerequisite. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences cover purpose, usage trigger, prerequisite, range, and defaults with no filler. The most important scoping information is front-loaded in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only two-parameter list tool with an output schema, the description is complete: it states what is returned, when to use it, the prerequisite, the allowed range, and the default behavior. Nothing needed for correct invocation or selection is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes both parameters and their defaults. The description adds value by stating the maximum range (31 days) and the default window (last 7 days), which helps an agent validate input without opening the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Review') with a clear resource ('hydration events and stored effective hydration totals'), and the hydration domain distinguishes it from the many list_* siblings. It states exactly what the tool returns without relying on the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use when the user explicitly asks about hydration history or fluid intake' and adds a prerequisite ('Hydration tracking must already be enabled in Settings'). It does not name an alternative tool, but no direct hydration-list sibling exists, so the guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_injuriesA
Read-onlyIdempotent
Inspect

List injuries from the injury log. update_injury and delete_injury can resolve an injury on their own from injury (+ optional date), so this is no longer required before either — use it to review the injury log, answer questions about injury history or rehab progress, or get an id after an ambiguous update_injury/delete_injury match. Defaults to active and monitoring injuries.

INFER — do not ask:

  • status: default to showing Active and Monitoring; use 'all' to include Resolved; use 'Resolved' for history only.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoFilter by status. Default: Active + Monitoring.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description does not need to restate safety. It adds valuable behavioral context beyond the schema: the default filter combines Active and Monitoring, the meaning of 'all' versus 'Resolved', and an explicit instruction to infer rather than ask.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear usage section and a compact INFER block. The only minor flaw is slight redundancy between the sentence 'Defaults to active and monitoring injuries' and the INFER line restating the same default, but the overall size is appropriate and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter, an existing output schema, and strong annotations, the description covers everything an agent needs: what it returns, when to use it, how to handle the status filter, and how it relates to sibling mutation tools. No important operational gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the status filter, so the baseline is 3. The description adds extra semantic value by specifying the default behavior, the exact meaning of each enum option, and an INFER directive that helps the agent invoke the tool without unnecessary clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List injuries from the injury log') and clearly distinguishes the tool's purpose from update_injury and delete_injury by explaining that those can resolve injuries without a prior lookup. It also enumerates concrete use cases: reviewing the log, answering history/rehab questions, and retrieving an id after ambiguous matches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool (review, answer history/rehab questions, get an id after ambiguous matches) and when it is not required (before update_injury/delete_injury). It also gives direct guidance on the status filter values, making the decision boundary clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_lab_markersA
Read-onlyIdempotent
Inspect

Returns all LOINC-coded markers in the reference library: canonical name, LOINC code, panel, typical unit, and common aliases. Call this BEFORE log_lab_results to match user-provided marker names to canonical entries — same pattern as list_exercises for workouts. Prevents name drift and ensures trending works across lab visits.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds useful behavioral context: it returns canonical library data, includes aliases for matching, and supports canonicalization. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The primary return value is front-loaded, the output fields are enumerated compactly, and the usage directive follows naturally. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no parameters, has an output schema, and is fully annotated as read-only and idempotent. The description adds the only missing context: what the data represents, what fields are returned, and how it should be used in the logging workflow. Nothing essential is left out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so there is no parameter burden for the description to carry. The description instead clarifies the semantic meaning of the returned canonical marker set, which is more valuable here. Baseline 4 is appropriate for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns'), a specific resource ('LOINC-coded markers in the reference library'), and enumerates the returned fields. It clearly distinguishes itself from list_lab_results, which is about logged lab results rather than the reference marker library.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent to call this BEFORE log_lab_results, gives the matching purpose, references the analogous list_exercises pattern, and explains why it matters ('prevents name drift and ensures trending works'). This is strong when-to-use guidance with a concrete alternative pattern.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_lab_resultsA
Read-onlyIdempotent
Inspect

List lab/biomarker results within a date range, including each result's ID. update_lab_result and delete_lab_result can resolve a result on their own from date (+ optional marker or panel_name), so this is no longer required before either — use it to review lab history, answer questions about blood work trends or specific marker values over time, or get an id after an ambiguous update_lab_result/delete_lab_result match. Optionally filter by panel or marker name.

INFER — do not ask:

  • start_date: default all time (labs are sparse)

  • end_date: default to today

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
panel_nameNoOptional — filter to a partial panel name (e.g. "Lipid Panel").
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: all time.
marker_nameNoOptional — filter to a partial marker name (e.g. "LDL Cholesterol").

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuine behavior beyond that: results include their ID, and update/delete can resolve results independently so this isn't a mandatory prerequisite. It doesn't describe volume/pagination behavior, but the output schema covers the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then the prerequisite-relaxation note, then use cases, then an INFER block for defaults. Dense but each clause carries routing or default guidance; the em-dash aside about update/delete is slightly long but earns its place by preventing unnecessary calls.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. The description covers purpose, when-to-use, when-not-to-use, alternatives, and the default-date inference policy — everything an agent needs to invoke it correctly without asking the user.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters (start_date, end_date, panel_name, marker_name) are already documented with formats and defaults. The description reinforces the date defaults via the INFER block but adds no syntax or matching semantics beyond what the schema already states, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (lab/biomarker results) plus scope (within a date range) and even what each record carries (the ID). It distinguishes itself from siblings like update_lab_result/delete_lab_result and list_lab_markers, so an agent can pick it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternatives (update_lab_result/delete_lab_result can resolve a result on their own) and states when this tool is NOT required before them. It then enumerates positive use cases: reviewing lab history, answering trend/marker questions, and disambiguating an id after an ambiguous match.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_locationsA
Read-onlyIdempotent
Inspect

List the user's saved training locations (home gym, hotel, gym) with the equipment at each, heaviest dumbbell and barbell load included. Use when the user asks where or with what they can train, before editing a location, or when planning needs to know what equipment is available. Planning rule: listed equipment is known available, not proof that unlisted incidental supports are absent; respect known load ceilings and explicit user limits, and let the user's stated location/equipment override assumptions. Returns each location id for manage_location. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so 'Read-only' is partly redundant. However, the description adds genuine behavioral context the annotations cannot convey: how to interpret the returned equipment list (known-available vs. exhaustively complete), that load ceilings and explicit user limits must be respected, and that user-stated location/equipment overrides assumptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then the planning rule, then the sibling handoff. Dense but every clause carries information; the long middle sentence on availability semantics is the one place a reader must slow down, keeping it just shy of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value documentation is not required, yet the description still flags that location ids are emitted for manage_location. Together with the usage triggers and the availability-interpretation rule, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter detail and instead spends its words on output interpretation, which is the right allocation for a no-arg list tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('List the user's saved training locations') and immediately enumerates the content shape (equipment per location, heaviest dumbbell and barbell load). This distinguishes it from the many other list_* siblings (list_workouts, list_meals, etc.) without requiring the schema to be opened.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names three trigger conditions ('when the user asks where or with what they can train, before editing a location, or when planning needs equipment availability') and points to the sibling that consumes its output ('Returns each location id for manage_location'). It also states an interpretation rule: listed equipment is known-available, not proof that unlisted incidental supports are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_mealsA
Read-onlyIdempotent
Inspect

List all meals logged for a date or date range, including each meal's ID, date, type, food description, and macros. update_meal and delete_meal can resolve a meal on their own from date (+ optional name substring), so this is no longer required before either — use it to answer "what did I eat today/this week/yesterday?", review what has been logged, or get an id after an ambiguous update_meal/delete_meal match.

Maximum range: 31 days per call. For longer periods, make multiple calls with sequential date ranges.

INFER — do not ask:

  • date: default to today

  • end_date: if the user asks about a week or range, set end_date to cover the full period (e.g. "this week" → date=Monday, end_date=today; "last 7 days" → date=7 days ago, end_date=today). For a single day, omit end_date.

ROUTING: Exact rows/IDs/ranges: list_meals. One-day visual diary: show_meal_diary. Multi-day macro trend: show_week_macros.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoStart date (or the single date if no range). Format: YYYY-MM-DD. Optional — omit for today (resolved in the user's own timezone).
end_dateNoEnd date for a range query. Format: YYYY-MM-DD. Optional — omit for a single-day lookup.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, idempotent, and non-destructive. The description adds meaningful behavioral constraints beyond annotations: the 31-day maximum range per call and the parameter inference rules (default today, end_date logic for week/range queries). This goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections and bullets, but it contains some redundancy (e.g., the note about update_meal/delete_meal resolving appears twice in slightly different forms). It is still efficient overall and not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description doesn't need to enumerate all return fields, but it still mentions key outputs (meal ID, date, type, food description, macros). It covers purpose, parameters, usage, routing, and constraints, providing everything an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (date and end_date) are fully described with format (YYYY-MM-DD), optionality, and specific semantics (start vs. single date, end of range). The description also explains default behavior and inference rules, making both parameters unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists meals logged for a date or range, with a specific verb ('List'), resource ('meals'), and scope (date/range). It also distinguishes it from siblings like show_meal_diary and show_week_macros via the ROUTING section.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: answering 'what did I eat today/this week/yesterday?', reviewing logged meals, and retrieving an ID after an ambiguous update/delete match. It also explains when it is not required (update_meal/delete_meal can resolve on their own) and how to handle longer periods with multiple calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_personal_contextA
Read-onlyIdempotent
Inspect

List the user's active Personal Context memories: durable circumstances and preferences remembered across conversations (e.g. travels most weeks, gym has no squat rack, trains early mornings, wants blunt feedback). Use when the user asks what has been remembered about them, or before proposing a new memory to check whether an existing one already covers the subject. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context by specifying that only 'active' memories are returned and by describing the type of content stored, which helps the agent set expectations beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: it states the core action and resource first, gives illustrative examples, then provides usage context. Every sentence adds value and there is no redundancy beyond the harmless 'Read-only' confirmation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only tool with an output schema and annotations covering safety, the description fully covers purpose, scope, content type, and when to invoke it. Nothing essential is missing for an agent to select and call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is trivially complete at 100% coverage. With no parameters to document, the description's focus on the returned content and usage context is appropriate. The baseline for zero-parameter tools is 4, and the description earns it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('List'), a specific resource ('the user's active Personal Context memories'), and includes concrete examples. It clearly distinguishes this tool from the sibling add_or_update_personal_context and from other list tools by focusing on durable remembered preferences and circumstances.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use: when the user asks what has been remembered about them, or before proposing a new memory to check for existing coverage. This gives clear practical guidance and implicitly routes the agent to add_or_update_personal_context when no existing memory covers the subject.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recovery_sessionsA
Read-onlyIdempotent
Inspect

List logged recovery sessions (completions and skips) within a date range, each with its ID. list_recovery_strategies only returns the recurring strategies (the schedule), never the individual logged entries against them — this is the only way to see those.

update_recovery_session and delete_recovery_session can resolve a session on their own from session_date (+ optional session_category), so this is no longer required before either — use it to answer "what recovery sessions have I logged", audit/spot-check past entries (e.g. a sauna and a cold plunge logged separately on the same day that should have been one contrast_therapy entry), or get an id after an ambiguous update_recovery_session/delete_recovery_session match.

INFER — do not ask:

  • start_date / end_date: default to the last 30 days. Widen the range yourself for an older lookup instead of asking the user for exact dates.

  • category: omit to return every category.

Maximum range: 90 days per call. To audit or correct a longer history, make multiple sequential calls walking backwards (days 0-90, then 90-180, then 180-270...) until you have covered the period the user means. Don't stop after one call and report that as the whole history.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional — max sessions to return. Default 50, max 200.
categoryNoOptional — filter to one category. Omit to return every category.
end_dateNoEnd of date range. Format: YYYY-MM-DD. Optional — defaults to today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Optional — defaults to 30 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint, destructiveHint, and idempotentHint, and the description does not contradict them. It adds meaningful behavioral context such as the 90-day maximum range, the need for sequential calls to cover longer periods, and the 'INFER — do not ask' directive, which helps the agent behave correctly. It does not explicitly mention output format or error behavior, but with an output schema present that is less critical.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-organized: it starts with the primary purpose, then contrasts with a sibling, and finally gives usage and inference rules. While it could be trimmed, the added detail on multi-call traversal and inference rules is necessary for correct usage, so the structure is appropriate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple optional parameters, inference rules, range limitations, and a sibling that is easily confused), the description covers all necessary context: it identifies when to use it, how to handle date ranges, how to distinguish from list_recovery_strategies, and how to handle long histories. No critical information is missing for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides descriptions for all four parameters, and the description enriches them further: it explains default values for start_date/end_date, that omitting category returns all categories, and the limit's default/maximum. This goes beyond the schema and gives the agent full understanding of each parameter's meaning and usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List', the resource 'logged recovery sessions', and the scope 'within a date range, each with its ID'. It also explicitly contrasts with the sibling tool list_recovery_strategies to remove ambiguity, so an agent immediately knows what this tool does and how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage scenarios ('answer what recovery sessions have I logged', audit/spot-check, get an id after ambiguous match) and gives concrete inference rules for parameters (default date range, category omission). It also tells the agent when to make multiple calls for longer histories, making the usage guidance highly actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recovery_strategiesA
Read-onlyIdempotent
Inspect

List the user's recovery and mindfulness strategies. Use when the user asks about their recovery practices, mindfulness routines, or you need strategy IDs before logging a session.

INFER — do not ask:

  • filter: default to 'active'; use 'all' for history; use 'historical' for ended strategies only.

Returns each strategy's id, name, category, schedule, start_date, and end_date.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoWhich strategies to return. Default: 'active'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/idempotent/non-destructive behavior. The description builds on that by explaining the INFER behavior ('do not ask'), the default filter, and what fields are returned. This adds useful operational context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose first, use cases second, then a clearly marked INFER block with parameter guidance, and finally a one-line return convention. Every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only list tool with annotations and an output schema, the description covers usage, filter semantics, inference behavior, and return fields. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds meaning not present in the schema: 'use 'all' for history; use 'historical' for ended strategies only.' This disambiguates the enum values and gives the agent a rule for inferring the parameter without asking.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description begins with a specific verb and resource: 'List the user's recovery and mindfulness strategies.' It also gives concrete use cases ('user asks about recovery practices, mindfulness routines, or you need strategy IDs before logging a session'), which distinguishes it from related siblings like manage_recovery_strategy or list_recovery_sessions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool: when the user asks about recovery practices, mindfulness routines, or needs strategy IDs before logging a session. It does not explicitly name alternatives or say when not to use it, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rest_daysA
Read-onlyIdempotent
Inspect

List user-selected Rest Day, Sick Day, and Travel Day context within a range. Use this to understand why a date may have no workout or why recovery, sleep, or nutrition data may look unusual. Travel Day is context-only and can coexist with training.

INFER — do not ask:

  • start_date: default to 30 days ago

  • end_date: default to today

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 30 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: Travel Day is context-only and can coexist with training, meaning a rest-day entry does not imply absence of training, and it explicitly instructs the agent to infer default dates rather than prompt the user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by usage rationale and then a short inference directive. Every sentence carries information, though the INFER block restates defaults that the schema already declares, which is mildly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values do not need explaining, and annotations plus the output schema cover the safety and shape of results. The description supplies the interpretive context an agent needs (what a rest-day entry means, how it interacts with training data). It could be slightly stronger by noting the distinction from log_rest_day/cancel_rest_day, but it is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (start_date, end_date) are already fully documented with format and defaults. The description's 'default to 30 days ago / today' merely duplicates the schema defaults, and the INFER instruction is a call-behavior hint rather than added parameter semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and a precise resource (user-selected Rest Day, Sick Day, and Travel Day context) with an explicit scope (within a date range). The semantic content — 'context-only, can coexist with training' — lets an agent distinguish it from data-bearing siblings like list_wellbeing or list_recovery_sessions without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to reach for this tool: to explain a date with no workout or anomalous recovery/sleep/nutrition data. That is a real usage signal rather than restated purpose. It stops short of naming an alternative or a when-not condition (e.g., use list_personal_context for other context types, or log_rest_day to add one).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsA
Read-onlyIdempotent
Inspect

List runs within a date range, or return full saved detail for one run when run_id is provided. Range reads stay compact and return a Runner State summary first. A run_id read returns that run's stored splits and structured segments plus the detailed metrics Elias needs for interval analysis.

Maximum range: 92 days. Defaults to 7 days ago through today. For a longer span, make several calls covering consecutive windows.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoSpecific run UUID. When provided, returns the full saved run including splits and structured segments and ignores the date range.
end_dateNoEnd of date range. Format: YYYY-MM-DD. Optional. Defaults to today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Optional. Defaults to 7 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the operation read-only, idempotent, and non-destructive, so the description is free to add behavioral value. It does: range reads 'stay compact and return a Runner State summary first,' run_id reads return 'stored splits and structured segments plus detailed metrics,' and the 92-day maximum with multi-call guidance is a meaningful constraint beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short paragraphs front-load the primary distinction and then add behavioral and constraint details. Every sentence earns its place: the mode split, the output differences, the defaults, and the multiple-call instruction. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With three optional parameters, 100% schema coverage, an output schema, and full annotation coverage, the description fills the remaining gaps: output shape per mode, maximum range, defaults, and how to handle spans longer than 92 days. Nothing an agent needs to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all three parameters with formats and defaults, so baseline is 3. The description adds the 92-day maximum range and clarifies the distinction between compact range output and detailed run_id output, reinforcing parameter behavior beyond the schema without repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List runs within a date range, or return full saved detail for one run when run_id is provided.' This clearly distinguishes the two modes of operation and their triggers, so an agent knows exactly what the tool does and when each mode applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: date-range reads return a compact summary, run_id reads return split and segment detail, and long spans require multiple calls covering consecutive windows. It does not explicitly name alternatives like show_runs, so sibling differentiation is left to inference, but the usage conditions themselves are concrete and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sleepA
Read-onlyIdempotent
Inspect

List sleep log entries within a date range. Each entry includes total duration, sleep score, stage breakdown, and the canonical bedtime and wake_time (full ISO 8601 timestamps with preserved timezone offset) for the user's primary overnight sleep session (excluding daytime naps).

Maximum range: 31 days per call. For longer periods, make multiple calls with sequential date ranges.

INFER — do not ask:

  • start_date: default to 14 days ago

  • end_date: default to today

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 14 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already provide readOnlyHint and destructiveHint, and the description adds transparency about the returned data fields (duration, score, stage breakdown) and the nature of the entries. There is no contradiction; the description supplements the annotation adequately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense, covering purpose, scope, parameters, defaults, limits, and exclusions in a clear, structured format without superfluous wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All necessary details for correct usage are present: the data returned, the date range constraints, the default behavior, and the exclusion of naps. Given the tool's simplicity, the description is fully complete and leaves no ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are explained with format and default values in the description, adding guidance on the 31-day limit and the inference of defaults, which goes beyond the schema's basic descriptions. This is particularly valuable given the schema only lists field names and descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'List' clearly specifies the action, the resource 'sleep log entries' is unambiguous, and the scope (date range) is defined. It also implicitly differentiates from logging a new sleep entry via the sibling tool 'log_sleep'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the maximum range of 31 days and instructs to make multiple calls for longer periods. It also clarifies the default values for start_date and end_date, and notes that daytime naps are excluded, providing clear when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_supplementsA
Read-onlyIdempotent
Inspect

List the user's medications and supplements. manage_supplement can resolve an item on its own from supplement_name, so this is no longer required before it — use it when the user asks what medications or supplements they're taking, asks to review their stack, or to get an id after an ambiguous manage_supplement match.

INFER — do not ask:

  • filter: default to 'active' (current items); use 'all' if the user asks about history or a specific past period; use 'historical' for ended items only.

  • category: omit to return both medications and supplements; set to 'medication' or 'supplement' to filter by type.

Returns each item's id, category, name, brand, dose, schedule, start_date, and end_date.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNoWhich items to return. Default: 'active'.
categoryNoFilter by category. Omit for both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered by structured data. The description adds useful behavioral context: default filtering behavior, how to request history, and the full list of returned fields. This goes beyond annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with the core purpose, followed by concise parameter inference guidance and a clear list of return fields. Every sentence adds value, and the use of labeled sections ('INFER — do not ask') makes it easy for an agent to parse. Nothing feels redundant or wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with two optional enum parameters and an output schema, the description is complete. It covers when to use the tool, what each parameter means, how to handle ambiguous matches, and what fields are returned. There is no missing information an agent would need to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the input schema covers both parameters 100%, the description adds significant extra meaning. It defines the inference rules, explains the default 'active' filter, clarifies when to use 'all' vs 'historical', and specifies that omitting category returns both types. This is exactly the kind of parameter guidance that helps an agent choose values correctly without asking the user.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the user's medications and supplements.' It clearly distinguishes itself from the sibling manage_supplement by explaining that list_supplements is for reviewing what the user takes, not for managing individual items. This lets an agent immediately know what the tool does and how it differs from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: when the user asks what medications/supplements they're taking, asks to review their stack, or needs an id after an ambiguous manage_supplement match. It also explains that manage_supplement can resolve items on its own, so list_supplements is not a required prerequisite. This is strong, actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_wearable_dataA
Read-onlyIdempotent
Inspect

List daily wearable data (steps, RHR, HRV, Zone Minutes / AZM including the vigorous-intensity breakdown, VO2max, calories eaten / dietary energy, calories burned / active energy / total energy expenditure, maintenance calories / TDEE, stress, and physiological vital signs reported by connected sources or manual overrides) within a date range. Use when the user asks about their step count, heart rate, HRV trends, vigorous minutes, calories eaten / dietary energy, calories burned, active or total energy expenditure, maintenance calories, TDEE, cardio fitness, blood glucose, vital signs, or any wearable metrics over time.

Zone Minutes (a.k.a. Active Zone Minutes) are shown as a daily total plus, when the per-zone breakdown is available, a moderate (1 pt/min) vs vigorous (2 pts/min — Cardio + Peak zones) split. The zone boundaries are personalized to the user's own resting and maximum heart rate, so they are not a fixed BPM.

Calories burned per day shows active + resting where both are known; when a device reports active energy with no resting figure, a resting estimate is derived from the user's profile BMR and labelled an estimate, never shown as measured. A trailing Maintenance (TDEE) line reports current maintenance calories and which method produced it (formula estimate vs. logged weight trend), or names the missing profile fields when TDEE can't be computed.

Maximum range: 31 days per call. For longer periods, make multiple calls with sequential date ranges.

INFER — do not ask:

  • start_date: default to 14 days ago

  • end_date: default to today

ROUTING: Exact daily values: list_wearable_data. Broad visual: show_health_overview. RHR/HRV: show_recovery. Steps: show_week_steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 14 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it read-only and idempotent, and the description adds meaningful behavioral context: the 31-day maximum, sequential date-range calls, Zone Minutes scoring details, and the explicit rule that derived resting calories are 'labelled an estimate, never shown as measured.' It also explains how TDEE method or missing profile fields are surfaced. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly structured: opening scope, expansion of ambiguous metric semantics, operational limit, inference defaults, and routing rules. Each section adds information an agent needs, and the key verb/resource is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations cover safety, the description covers everything needed to call correctly: metric coverage, date-range limits, batching, estimation caveats, default inference, and sibling routing. No critical behavioral or usage element is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already documents both parameters with defaults and formats, the description adds the actionable instruction 'INFER — do not ask', the default values, and the operational constraint that maximum range is 31 days per call with sequential calls for longer periods. This goes beyond the schema's structured parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List daily wearable data' within a date range, and enumerates the exact metrics covered. The ROUTING section further distinguishes it from show_health_overview, show_recovery, and show_week_steps, so an agent can tell siblings apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit routing ('Exact daily values: list_wearable_data. Broad visual: show_health_overview. RHR/HRV: show_recovery. Steps: show_week_steps') and a 31-day call batching rule. However, the earlier 'Use when the user asks about ... HRV trends' is not fully reconciled with the ROUTING line that sends RHR/HRV to show_recovery, creating a minor ambiguity for that query type.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_wellbeingA
Read-onlyIdempotent
Inspect

List wellbeing log entries within a date range. update_wellbeing and delete_wellbeing can resolve an entry on their own from date, so this is no longer required before either — use it to answer questions about mood, energy, stress, or soreness trends over time, or get an id after an ambiguous update_wellbeing/delete_wellbeing match.

Maximum range: 31 days per call. For longer periods, make multiple calls with sequential date ranges.

INFER — do not ask:

  • start_date: default to 14 days ago

  • end_date: default to today

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Default: today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Default: 14 days ago.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark it as read-only, idempotent, and non-destructive; the description adds further transparency by explaining that the tool is not a prerequisite for update/delete operations and that it returns ids useful for resolving ambiguous matches. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence contributes: the main purpose, the relationship to update/delete, the 31-day limit, and the parameter defaults. No redundant or filler content; the structure flows logically from what the tool does to how to use it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two optional parameters and the presence of an output schema (mentioned in context), the description provides sufficient context for an agent to call the tool correctly. It covers parameter defaults, usage scenarios, and operational limits, so no missing information would cause incorrect invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes both parameters with types and defaults, and the description reinforces their behavior ('INFER — do not ask' with default values), clarifying that the agent should infer these rather than prompt the user. This adds semantic value beyond the schema's basic metadata.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action ('List wellbeing log entries') and the resource, with explicit scope ('within a date range'). It distinguishes itself from sibling tools by noting that update_wellbeing and delete_wellbeing can resolve entries on their own, making list_wellbeing specific to trend queries and id retrieval after ambiguous matches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: for trends over time or getting an id after an ambiguous match, and explicitly states it is not required before update/delete. Also gives operational constraints (maximum 31 days per call) and default parameter values, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workoutsA
Read-onlyIdempotent
Inspect

List workout sessions in a date range plus user-marked Rest Day, Sick Day, or Travel Day context. Day context is returned even when that date has no workout. Treat context labels as explanatory context, not workouts; Travel Day can coexist with training.

Use before get_workout to find a session ID, or to answer training-history questions. Each workout row includes ID, date, focus type, location, and session-level NSI with rating.

DATE RANGE: one call covers the requested span, internally reading bounded 90-day slices. Explicit ranges can be up to 3660 days. Do not make sequential list_workouts calls to cover one user-requested range.

ALL HISTORY: when the user asks for "all", "ever", or their full workout history, set all_history=true. The tool finds the earliest stored workout (respecting activity_type when supplied). If history exceeds 3660 days, it covers the most recent 3660 days and explicitly reports the older stored start date instead of throwing.

OUTPUT BOUND: at most 250 workout rows are emitted. When more match, the response reports how many were omitted (or an honest lower bound for an exercise filter), the earliest returned/scanned date, and tells you to narrow with activity_type, exercise_name, or dates.

FILTERS: activity_type matches the stored workout focus type case-insensitively. exercise_name keeps only sessions containing that exercise in workout sets. Use activity_type for imported modalities such as Stair Stepper; use exercise_name for strength-history questions. Both filters may be combined.

INFER: default start_date to 7 days ago, end_date to today. User-specified dates or years are authoritative. Widen freely for the period the user actually requested.

SAVED WORKOUTS: set saved_workouts=true to list the user's reusable Saved Workouts library instead of completed workout history. Saved workout IDs are separate from workout session IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_dateNoEnd of date range. Format: YYYY-MM-DD. Optional — defaults to today.
start_dateNoStart of date range. Format: YYYY-MM-DD. Optional — defaults to 7 days ago.
all_historyNoSet true for all stored workout history instead of the recent default. Finds the earliest stored workout automatically; activity_type narrows that earliest-date lookup when present.
activity_typeNoStored workout focus type to match case-insensitively, for example Stair Stepper. Defaults to no activity-type filter.
exercise_nameNoExercise name that must appear in the workout sets, matched case-insensitively. Defaults to no exercise filter.
saved_workoutsNoWhen true, list reusable Saved Workouts instead of completed workout sessions. Defaults to false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior, and the description adds substantial operational context beyond them. It discloses internal 90-day slicing, a 3660-day cap, output bound of 250 rows, what happens when results are omitted, and how all_history handles older stored start dates. These details are exactly the kind of behavioral context an agent needs to invoke the tool correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with capitalized section headers that front-load the core purpose and then handle date ranges, all-history behavior, output bounds, filters, inference, and saved workouts. Most sentences earn their place by preventing incorrect invocations, though some return-field detail overlaps with the existing output schema. It is appropriately sized for a complex, multi-mode listing tool, but not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six optional parameters, a rich schema, an output schema, and many siblings, the description is complete enough for correct invocation. It covers filtering, defaults, pagination-like bounds, all-history edge cases, saved workouts, and sibling routing. Nothing critical to selecting or calling the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds meaningful parameter semantics beyond the schema. It explains when to set all_history, how activity_type and exercise_name differ in practice, that both filters may be combined, and that saved_workouts switches to a separate Saved Workouts library with distinct IDs. It also gives date-inference rules and treats user-specified dates as authoritative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: list workout sessions in a date range, plus Rest Day/Sick Day/Travel Day context. It clearly distinguishes the tool from siblings by explaining that it returns session IDs for use before get_workout and that it handles training-history questions. The added warning that context labels are not workouts further sharpens the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: before get_workout to find a session ID, for training-history questions, for all-history queries, and for saved-workout library listing. It also provides when-not guidance by saying not to make sequential list_workouts calls to cover one range and by distinguishing activity_type from exercise_name filter use cases. Alternative sibling tools are not exhaustively named, but the key routing decisions are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_body_metrics
Destructive
Inspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Log or update body composition metrics for a given date. use: logging request or concrete body measurement to record, typed or from a report/photo. Question/hypothetical alone: no write.

PROACTIVE DATA COLLECTION: If the user hasn't shared their data yet, ask them to copy-paste the output from their scale app or upload a photo of the display — this lets you parse all fields at once instead of asking one by one.

INFER — do not ask:

  • date: default to today; infer from context ("this morning", "yesterday")

  • derived fields (lean_mass_lb, fat_mass_lb): calculate from weight and body fat % if possible — lean = weight × (1 - bf%/100), fat = weight × bf%/100

The *_pct fields are percentages 0-100. The *_in fields are manual tape measurements, not bioimpedance scale readings.

BMI is not a field here. It is derived automatically at read time from the user's canonical height and their resolved weight for the day -- never ask the user for BMI, and never try to log it.

Every field except date is optional; log any subset. One row per day. Calling this tool twice on the same date updates the existing entry (upsert).

DELETE: action='delete' with date and delete_fields clears only those measurements from the user's manual entry for that day, keeping the rest; the entry is removed once nothing is left. A weigh-in is weight_lb. Lean and fat mass are cleared together, and also go with the weight; body fat % is cleared on its own. The reply lists exactly what was cleared. Other value fields are ignored. It cannot remove readings a connected scale or app imported: those come back on the next sync, so tell the user to delete them in that app. ASK for confirmation if the user's intent is ambiguous.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesDate of the measurement. Format: YYYY-MM-DD. Default to today.
notesNoAny context worth noting (e.g. "post-workout", "morning fasted").
actionNoDefault 'log'. See DELETE above.
hips_inNoHip circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
chest_inNoChest circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
waist_inNoWaist circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
calf_l_inNoLeft calf circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
calf_r_inNoRight calf circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
weight_lbNoBody weight, between 50 and 700. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
bicep_l_inNoLeft bicep circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
bicep_r_inNoRight bicep circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
thigh_l_inNoLeft thigh circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
thigh_r_inNoRight thigh circumference. In in, or cm with input_length_unit set. See UNIT INPUTS.
fat_mass_lbNoFat mass. Calculate from weight and body fat % if not explicitly stated. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
protein_pctNoProtein percentage (0-100).
body_fat_pctNoBody fat percentage (0-100).
bone_mass_lbNoBone mass. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
lean_mass_lbNoLean (non-fat) mass. Calculate from weight and body fat % if not explicitly stated. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
delete_fieldsNoMeasurements to delete. See DELETE above.
hydration_pctNoBody water/hydration percentage.
muscle_mass_lbNoSkeletal muscle mass. Skeletal muscle tissue specifically, NOT lean_mass_lb, which is total non-fat mass including water, organs and bone. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
input_length_unitNoSet to cm when the user gave cm for the _in fields in this object. Omit when they are already in.
input_weight_unitNoSet to kg when the user gave kg for the _lb fields in this object. Omit when they are already lb.
skeletal_muscle_pctNoSkeletal muscle percentage (0-100).
visceral_fat_ratingNoVisceral fat rating (scale varies by device, typically 1-59).

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.
log_cycleAInspect

Log a period to the user's cycle log. Handles all cases:

  • Starting a period today: "my period started today"

  • Backfilling a past period: "my period started May 3rd and ended May 8th"

  • Resuming a period ended today: "actually I'm still on my period" — detects that today's period was marked ended and reopens it

  • Logging just a start with no end yet: "I just got my period"

Before logging, check that cycle tracking is enabled (cycle_prefs.tracking_enabled AND consented_at, both required -- a user can have consented in the past and later turned tracking off). If not, tell the user to turn it on from the dashboard first.

INFER — do not ask:

  • started_on: default to today for current-period statements

  • ended_on: omit unless the user says it ended; infer from context ("5-day period starting May 3" → ended_on May 7)

Do NOT use this tool to log future dates.

ParametersJSON Schema
NameRequiredDescriptionDefault
ended_onNoPeriod end date. Format: YYYY-MM-DD. Omit if period is still active.
started_onYesPeriod start date. Format: YYYY-MM-DD. Required.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses meaningful behavior: it can reopen a period that was marked ended, it infers started_on and ended_on instead of asking, and it requires a consent/tracking check before logging. No statement contradicts the provided annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear top-line definition, a bulleted list of cases, a prerequisite warning, and an inference section. Every sentence carries information an agent needs; nothing is filler or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers edge cases, prerequisites, inference rules, and forbidden usage. An output schema exists, so the description does not need to explain return values. For a tool that logs cycle data with inference and consent requirements, this is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the schema already documents both parameters at 100% coverage, the description adds crucial semantics: started_on defaults to today for current-period statements, ended_on should be omitted unless explicitly stated, and both dates can be inferred from context. This goes well beyond the schema's basic format descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource, 'Log a period to the user's cycle log', and then enumerates the distinct real-world utterances it handles. This makes the tool's scope unmistakable and clearly separates it from siblings like list_cycle, update_cycle, and delete_cycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit applicability through concrete examples, a prerequisite check for cycle tracking, and an explicit exclusion: 'Do NOT use this tool to log future dates.' It does not explicitly route users to update_cycle for modifications, but for a logging action the guidance is strong and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_injuryAInspect

Log a new injury or aggravation. Use when the user mentions getting hurt, feeling pain, straining something, or describes an injury. Injuries are date ranges — they start on a date and are ongoing until an end_date is set.

INFER — do not ask:

  • start_date: default to today

  • severity: estimate from description (minor twinge=2-3, moderate pain=5-6, severe/acute=8-9)

  • status: default to 'Active' for new injuries

  • affected_movements: infer from body part and injury type (e.g. shoulder strain → pressing, overhead)

  • side: infer from description if mentioned (e.g. "right shoulder" → Right)

ASK only if body_part is entirely unclear.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideNoSide of body. Optional — infer from description.
notesNoAdditional context about how injury occurred, symptoms, etc. Optional.
statusNoStatus. Default: 'Active'. Auto-set to 'Resolved' if end_date is provided.
end_dateNoDate injury resolved. Format: YYYY-MM-DD. Default: null (ongoing). Set when injury is fully resolved.
severityYesSeverity 1-10 (1=minor, 10=severe). Required — estimate from description.
body_partYesBody part affected (e.g. 'Shoulder', 'Lower Back', 'Knee'). Required.
start_dateNoDate injury started. Format: YYYY-MM-DD. Default: today.
injury_typeYesType of injury (e.g. 'Strain', 'Tendonitis', 'Sprain', 'Disc', 'Soreness', 'Acute'). Required — infer from description.
affected_movementsNoMovements affected (e.g. ['Pressing', 'Overhead', 'Bench Press']). Infer from body part.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: injuries are date ranges that stay open until end_date is set, many fields should be inferred rather than asked, and only body_part warrants a clarifying question. The annotations correctly indicate a write operation with readOnlyHint=false, so there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: purpose, trigger conditions, domain model, then a scannable INFER list. Every sentence earns its place, and there is no redundancy with the schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter create tool, the description covers when to call, what to infer, what to ask, and the date-range semantics. Since an output schema exists and schema coverage is 100%, no critical information is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful extra value with numeric severity anchors (minor twinge=2-3, moderate pain=5-6, severe/acute=8-9), an example for affected_movements, and a clear policy to ask only when body_part is unclear. This is above baseline but not maximal because the schema already carries much of the parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Log a new injury or aggravation.' It also lists concrete user signals like 'getting hurt, feeling pain, straining something' that trigger this tool. The word 'new' distinguishes it from sibling tools like update_injury and delete_injury.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: when the user mentions getting hurt, feeling pain, straining, or describes an injury. It doesn't explicitly state when not to use it, such as 'if the injury already exists, use update_injury instead,' but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_lab_resultsAInspect

Log one or more blood test or biomarker results. use: requested logging/import of lab values or a concrete result to record. Review/interpretation alone: no write.

GLUCOSE ROUTING: use this tool for glucose only when it is an actual lab/blood-draw result or part of a reported lab panel. Finger-stick, CGM, home meter, wearable, Apple Health, or Health Connect glucose — including a manual correction to daily glucose — belongs in log_wearable as blood_glucose_mg_dl, not here.

REQUIRED WORKFLOW: 1) call list_lab_markers for canonical names and LOINC codes. 2) for each marker the user provides, find the best match and use its canonical marker_name and loinc_code. 3) if no match exists, use the name as stated and omit loinc_code.

If the user says they have lab results but hasn't shared them, prompt: "You can paste the text from your lab report PDF, or upload a photo of the results page — I'll parse all the values at once."

collection_date: report or explicit user context; missing: ask, never substitute report_date or today. infer: panel_name from matched marker or context. extract: flag (H/L/HH/LL/A), ref_range_low/high and lab_name only when reported.

FASTING: fasting_status is only 'fasting' or 'non_fasting' when the user explicitly states it ("I fasted for this", "hadn't eaten"). Never infer it from time of day, a meal mention in notes, or the draw being in the morning — omit the field (leaves it unknown) whenever it wasn't stated. report_date (when the lab reported results, if given separately from the collection date) and draw_id (only when the user is adding a marker to a draw already logged in this conversation) are also optional, omit otherwise.

Submit all markers from a single lab visit in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
resultsYesArray of individual lab marker results. Required — submit all markers from the visit in a single call.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only state flags (readOnlyHint=false, etc.). The description adds substantial behavioral rules: required workflow order, handling unknown markers (use name as stated, omit LOINC), never inferring fasting status, never substituting report_date for collection_date, and batching all markers in one call. It also explicitly states what is *not* allowed, which goes well beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly organized with clear headings (GLUCOSE ROUTING, REQUIRED WORKFLOW, FASTING) and every sentence earns its place. Some minor repetition of 'omit otherwise' and the explicit prompt text add length, but overall it is efficiently structured and front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is highly complex with many conditional fields and routing rules. The description covers the required pre-step, handling of missing data, field-specific semantics, and batching, and an output schema exists to explain return values. Nothing an agent needs to invoke this correctly is left out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description enriches every parameter. It clarifies that fasting_status must be omitted unless explicitly stated, explains when draw_id applies, specifies that loinc_code must come from list_lab_markers, and gives rules for panel_name inference. This adds meaning far beyond the bare schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Log one or more blood test or biomarker results.' It immediately contrasts with review/interpretation ('Review/interpretation alone: no write') and explicitly routes glucose readings to log_wearable, distinguishing it from the main sibling. An agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use ('requested logging/import of lab values or a concrete result to record') and when-not-to-use ('Review/interpretation alone'). It also names the sibling alternative for glucose (log_wearable), mandates a pre-step with list_lab_markers, and gives a ready-made prompt for incomplete user input. This is textbook usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_mealInspect

RECIPE SAVE: intent=create/save recipe not eaten -> save_as_recipe=true, recipe_only=true intent=log eaten meal + save recipe -> save_as_recipe=true, recipe_only=false/omit recipe_only=true -> food_log_write=false; save_as_recipe must be true recipe_only=true -> recipe_title= recipe_only=true -> recipe_servings=<whole-batch servings; default 1> recipe_only=true -> food_items= recipe_only=true -> calories/protein_g/fat_g/carbs_g=<WHOLE-BATCH totals; estimate, never ask> recipe_only=true -> estimate=

Log a meal to the user's food diary. use: logging request or concrete meal consumed. question/habit/hypothetical alone: no write.

INFER:

  • date: today, or from context ("yesterday", "last night"). For an attached meal photo with , use that photo's local date unless the user explicitly states a different date.

  • meal_time: for an attached meal photo with , send that photo's local HH:MM unless the user explicitly states a different time/date. User-stated timing always wins. Otherwise omit unless the user gave a real clock time.

  • meal_type: when meal_time is supplied, OMIT meal_type unless the user's wording explicitly names or clearly anchors a category; the server assigns the category from meal_time. Without meal_time, for a current-day plain meal/food log omit meal_type and let the server assign it from the user's resolved local clock (00-04 Snack, 04-10 Breakfast, 10-14 Lunch, 14-17 Snack, 17-22 Dinner, 22-24 Snack). Send meal_type when the user's wording explicitly names or clearly anchors a category ("breakfast", "for lunch", "post-workout shake"), or when backdating and neither capture time nor another honest clock signal exists.

MACRO SOURCE: strongest evidence wins. Never replace known stored macros with a fresh estimate.

  1. REPEATS (unless a saved recipe is clearly invoked): if the user refers to a previously logged item ("same", "another", "more", "again", or equivalent in any language), call list_meals for the referenced date/range first. If one row unambiguously matches, reuse its stored calories/protein_g/fat_g/carbs_g/alcohol_g and scale by the quantity ratio when the row's quantity is known. If the match or quantity is ambiguous, ask instead of re-estimating. If nothing matches, continue below.

  2. SAVED RECIPES: pass recipe_name whenever the food phrase plausibly names one of the user's saved recipes (see the profile's "Saved recipes" list) -- not only when the user literally says "saved" or "usual"; a bare "morning coffee" should try recipe_name if "Morning Coffee" is one of theirs. recipe_values: Saved macros and nutrients, including estimates and zeros, are authoritative unless explicitly corrected. Scale with recipe_multiplier; combine added_food_items using final meal macros. Estimate only missing values or changed/added foods. Logging never modifies the saved recipe. Relay no-match/ambiguous errors instead of guessing. A recipe's saturated_fat_g/fiber_g are carried automatically when its stored nutrients have them. They are expanded nutrients, never required to log the meal; do not invent them just to satisfy this tool. If the recipe reports missing core macros, estimate only those core fields and retry. If food_items adds food beyond the recipe, pass the combined core macros; expanded nutrients remain optional and are handled by the gated nutrient pipeline when enabled.

  3. BRAND NAMES: for a branded, restaurant, or specific product, use published macros for the stated size/variant before a generic estimate.

  4. Otherwise estimate every macro from the food description. Never ask the user for macros.

MACROS: final stored meal must contain calories, protein_g, fat_g and carbs_g. recipe_name supplies stored core values; estimate and pass only missing core values or explicit changes. Without a recipe, pass food_items and those four. No partial core macros or placeholders. Saturated fat and fiber are expanded nutrients, never required for log_meal. Pass saturated_fat_g/fiber_g only when already known from strong evidence; otherwise omit them. When expanded nutrient tracking is enabled, the gated nutrient estimator can fill them; when disabled they remain unset. 0 is a real value only when the food truly contains none of that nutrient. Ask only if food_items are absent and no recipe_name applies.

FASTING: if the user ate nothing / fasted all day, log one "Fast day" Snack with every macro 0.

ESTIMATE: pass a compact 3-line block estimating every one of the app's 44 tracked nutrients for the food, in exactly this format, no prose before or after: items::|:|... macros:;;;; nutrients:1:;2:;...;44: Estimate all 44 nutrients. Always return all 44 IDs. Never omit one because you are uncertain -- use your best reasonable estimate from the likely food, quantity, ingredients, preparation and comparable foods; 0 only when a nutrient is reasonably expected to be negligible. Which id is which nutrient, and its unit, is the dictionary on the estimate param. Omit for an unchanged saved recipe; with added_food_items, estimate that addition only. Otherwise supply when available; the server estimates missing values.

DUPLICATES: if this tool returns a duplicate error, tell the user what's already logged and ask whether this is a separate serving (retry with force=true) or should update the existing entry instead (update_meal with adjusted values).

DRINKS / HYDRATION: this is the single write path for consumed drinks too. Known or reasonably inferable drink volume (tea, coffee, milk, juice, shakes, smoothies, etc.) → always report fluids, even if also contributes calories/macros/alcohol. Food moisture → omit. The server decides whether hydration tracking is enabled; never use a separate hydration write tool. A plain fluid-only intake may omit food_items and macros and send only fluids. For alcoholic drinks, include that drink's alcohol_g in its fluid item; if there is exactly one drink, the top-level alcohol_g can stand in.

CAFFEINE / DRINKS: this is the single assistant write path for caffeine too. Drinks only; exclude food moisture. When the user consumed a drink and its volume is known or reasonably inferable, include it in fluids so hydration is recorded; if the drink is caffeinated, include caffeine in the same call too. Do not leave a known-volume drink only in food_items. Each caffeine item needs caffeine_mg and may optionally include source_type. The date comes from the intake date. Preserve exact user-provided milligrams; for vague coffee/tea/energy-drink/pre-workout descriptions, estimate caffeine with the same nutrition-estimation judgment used for meal macros. Never use separate hydration or caffeine write tools. A drink-only or caffeine-only intake may omit food_items/macros.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoYYYY-MM-DD. Omit for today (the server resolves it in the user's own timezone, more reliable than guessing). Send explicitly for any past date.
fat_gNoFat in grams. See MACROS above.
forceNoTrue only when the user has explicitly confirmed a separate entry despite a duplicate warning. Bypasses duplicate detection.
fluidsNoOptional drinks consumed in this intake, including liquid ingredients in shakes or smoothies. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking.
carbs_gNoCarbohydrates in grams. See MACROS above.
fiber_gNoDietary fiber in grams. See MACROS above.
caffeineNoOptional caffeine doses in this intake. Keep each dose simple: caffeine amount is required and type is optional. The dose date follows the intake/meal date. If the user gave exact milligrams, preserve them exactly; otherwise estimate from the described food or drink.
caloriesNoTotal calories(kcal) never kJ
estimateNoThe compact 3-line nutrient estimate for this meal. See ESTIMATE above for the format. Omit to have the server estimate instead -- never blocks the write either way. A block whose parts exceed their whole (saturated fat over fat, fiber over carbs) is discarded and re-estimated. NUTRIENT ID DICTIONARY (id=name, unit is the name's suffix; ug=mcg). Every id means exactly this nutrient, never guess the order: 1=fiber_g 2=sugar_g 3=saturated_fat_g 4=monounsaturated_fat_g 5=polyunsaturated_fat_g 6=trans_fat_g 7=cholesterol_mg 8=sodium_mg 9=potassium_mg 10=calcium_mg 11=iron_mg 12=magnesium_mg 13=phosphorus_mg 14=zinc_mg 15=copper_mg 16=manganese_mg 17=selenium_ug 18=chloride_mg 19=chromium_ug 20=iodine_ug 21=molybdenum_ug 22=vitamin_a_ug 23=vitamin_c_mg 24=vitamin_d_ug 25=vitamin_e_mg 26=vitamin_k_ug 27=thiamin_b1_mg 28=riboflavin_b2_mg 29=niacin_b3_mg 30=pantothenic_acid_b5_mg 31=vitamin_b6_mg 32=biotin_b7_ug 33=folate_b9_ug 34=folic_acid_ug 35=vitamin_b12_ug 36=choline_mg 37=omega3_g 38=omega6_g 39=caffeine_mg 40=water_g 41=starch_g 42=added_sugar_g 43=total_unsaturated_fat_g 44=fluoride_mg
alcohol_gNoAlcohol in grams (not kcal). Only if alcoholic drinks were consumed; unset takes the saved recipe's value when recipe_name matches. 1 standard drink is about 14g.
meal_timeNoOptional local 24-hour HH:MM time. For an attached meal photo, use the current turn's photo capture time when the user did not state a different date/time. User-stated timing always wins.
meal_typeNoOptional. Omit for a current-day plain meal/food log so the server assigns the category from the user's resolved local clock. Send only when the user explicitly names or clearly anchors a category, or when backdating and the current clock cannot apply. A saved recipe's stored meal type never fills this in.
protein_gNoProtein in grams. See MACROS above.
food_itemsNoDescription of the food and drinks consumed. See MACROS above; ask the user only if completely absent and no recipe_name applies.
recipe_nameNoSaved recipe title only, without portion words or additions. See SAVED RECIPES above.
recipe_onlyNotrue=save recipe only and write no food_log row. Requires save_as_recipe=true. See RECIPE SAVE above.
recipe_titleNoRecipe name when recipe_only=true. See RECIPE SAVE above.
save_as_recipeNoTrue only when the user explicitly asks to save a reusable recipe. recipe_only=true saves recipe only; otherwise the final persisted meal is saved after the log succeeds. See RECIPE SAVE above.
recipe_servingsNoWhole-batch serving count when recipe_only=true; default=1. See RECIPE SAVE above.
saturated_fat_gNoSaturated fat in grams. See MACROS above.
added_food_itemsNoOnly food added to a saved recipe; supply final combined macros. Omit when none.
recipe_multiplierNoPortions of the matched saved recipe; default 1, e.g. 0.5 for half.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.
log_recovery_sessionAInspect

Log a completed or skipped recovery/mindfulness session. Use when the user says they did (or skipped) a breathing exercise, meditation, cold plunge, sauna, stretching, or any recovery practice. Also use for one-off standalone sessions not linked to a recurring strategy.

INFER — do not ask:

  • date: default to today

  • category: infer from the practice name

  • strategy_name: use the strategy name if linked, or the user's description

  • duration_minutes: infer if mentioned (omit for skipped sessions)

  • quality: only include if the user rates it (1-5 scale)

  • skipped: true when the user says they skipped, missed, or didn't do a session; false (default) for completed sessions

PREFERRED WORKFLOW: call list_recovery_strategies first to link the session to an active strategy for adherence tracking. If no matching strategy exists, log as standalone.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoDate (YYYY-MM-DD). Default to today.
notesNoSession notes or reason for skipping. Optional.
qualityNoSubjective quality 1-5. Optional.
skippedNotrue if the session was skipped/missed. Default: false.
categoryYesCategory. Required.
strategy_idNoStrategy ID from list_recovery_strategies. Optional — omit for standalone sessions.
strategy_nameYesName of the practice. Required.
duration_minutesNoDuration in minutes. Optional — omit for skipped.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, it details real behavior: infer date/category/strategy_name instead of asking, omit duration_minutes for skipped sessions, include quality only when the user rates it, and map 'skipped/missed/didn't do' to skipped=true. These rules tell the agent exactly how the tool expects the call to be constructed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into a two-sentence purpose, a compact INFER bullet list, and a short workflow callout. There is no filler, and the highest-value guidance is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter write operation, it covers all major decision points: defaults, inference rules, skipped handling, strategy linking, and the standalone fallback. Optional fields like notes are already documented in the schema, and an output schema exists, so the description does not need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description adds inference semantics that the schema cannot express: category comes from the practice name, strategy_name comes from the linked strategy or user phrasing, and duration is omitted when skipped. This materially improves an agent's ability to fill the parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete action and resource: 'Log a completed or skipped recovery/mindfulness session,' and lists representative practices. It also carves out 'one-off standalone sessions not linked to a recurring strategy,' which separates this tool from strategy-management siblings and gives the agent a clear scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trigger is explicit ('Use when the user says they did or skipped...') and the preferred workflow names list_recovery_strategies and a standalone fallback, so the agent knows when to invoke it. It never states a when-not or names a sibling tool as the alternative, leaving some exclusion logic implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_rest_dayA
Idempotent
Inspect

Mark a date with user-selected day context: Rest Day, Sick Day, or Travel Day. Use Rest Day when the user intentionally takes the day off training, Sick Day when they say they are sick/recovering, and Travel Day when they say they are traveling. This context is visible to the AI team and workout-history reads.

Rest/Sick retain the existing rest-day training behavior. Travel is context-only: it does not mean the user skipped training and can coexist with a workout.

INFER — do not ask:

  • date: parse the user's reference; default to today

  • day_type: infer rest, sick, or travel from the user's wording; default to rest

Calling again on the same date updates the status in place. To remove the status, use cancel_rest_day.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoThe date to mark. Format: YYYY-MM-DD. Default: today.
day_typeNoUser-selected day context. Default: rest.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (idempotentHint=true, destructiveHint=false, readOnlyHint=false), and the description reinforces idempotency by stating a repeat call updates status in place. It adds non-obvious behavior: visibility to the AI team and workout-history reads, and that Travel is context-only and can coexist with a workout. Strong given the annotations already carry the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose before branching into day_type semantics, with a clear INFER block and a closing note on idempotency/removal. Every sentence carries information, though it is slightly verbose for a two-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described. The description covers date interpretation, all enum semantics, idempotent re-call behavior, cross-system visibility, and the removal path via cancel_rest_day, leaving no gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns more by explaining how to derive each parameter: parse the user's date reference defaulting to today, and infer day_type from wording defaulting to rest. This adds interpretation guidance the schema does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Mark) and resource (a date with day context), and enumerates the three day types. It distinguishes itself from the sibling cancel_rest_day and is trivially separable from log_wellbeing/log_sleep. An agent knows exactly what this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly maps each day_type to the user condition that selects it (takes day off, sick/recovering, traveling), and names cancel_rest_day for the removal case. It even prescribes inference over asking, which removes ambiguity about behavior at call time.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_runAInspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

RUNNING ONLY: use log_run only for running on foot. A bike/cycling ride, walk, or rowing session is not a run even when it has distance and duration. For those, use log_workout with focus_type Cycling, Walking, or Rowing plus the user's native distance field and duration_sec.

Create an editable running card. Use for a completed run OR a future run plan.

intent:

  • log (default): the run happened. date, distance_mi and duration_sec are required. This writes the completed run and returns a card marker for in-app editing.

  • plan: the run has NOT happened yet. Create a planned run card. Never put a future run in completed history.

RUN TYPE: distinguish easy, long, tempo, interval, recovery, race, fartlek, threshold, progression and hills when the runner or workout structure supports it. If a wearable run is unspecified, do NOT call it easy merely because it was a run.

DETAIL: preserve elapsed time, HR, cadence, power, elevation, RPE, splits and structured segments only when supplied by the user/source. Never invent sensor data or splits.

ParametersJSON Schema
NameRequiredDescriptionDefault
rpeNoRPE 1-10. For a plan this is target RPE.
dateNoYYYY-MM-DD. Completed runs default to today when context allows; planned runs use the intended date when known.
notesNo
avg_hrNoAverage heart rate only when known.
intentYeslog = completed run; plan = future run. Elias uses plan for a workout the runner has not done yet.
splitsNoActual splits/laps only when supplied by the runner/source. Never invent.
surfaceNo
run_typeNoClassify only when the runner or workout structure supports it. Do not turn an unspecified wearable run into easy.
segmentsNoWorkout structure, e.g. warmup, 6x800m threshold, recovery, cooldown. This may describe a plan or what was actually performed.
pain_notesNoPain/discomfort exactly as the runner described it.
start_timeNoLocal HH:MM only when known.
avg_power_wNoAverage running power in watts only when known.
distance_miNoDistance. In mi, or km with input_distance_unit set. See UNIT INPUTS.
elapsed_secNoWall-clock elapsed time including pauses, only when known.
duration_secNoMoving/workout duration in seconds.
avg_cadence_spmNoAverage running cadence in steps/min only when known.
elevation_gain_mNoElevation gain in metres only when known.
perceived_effortNoPost-run check-in only.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the description doesn't need to restate that this is a write operation. It adds behavioral context by explaining what the tool does with the value field (converts once before storage), how intent affects the outcome (writes completed run vs creates planned card), and explicitly says never to invent sensor data or splits. The description goes beyond the annotations by disclosing the conversion convention and the no-invention rule.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact relative to the 19-parameter schema. It front-loads the critical UNIT INPUTS and RUNNING ONLY rules, then intent, then run_type, then detail. Every sentence earns its place – the RUNNING ONLY and UNIT INPUTS sections are high-value disambiguation that prevents wrong tool selection and wrong unit handling. The format uses clear section headers to break up the different concerns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter tool with an output schema and high schema coverage, the description covers the key ambiguities: when to use this tool vs log_workout, how intent changes the call, how units are handled, and what not to invent. It could add a note about the output schema/returned card marker, but the output schema exists and the intent section already mentions it returns a card marker for in-app editing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, which is high, so the baseline is 3. The description adds additional meaning by clarifying the UNIT INPUTS convention (pass value unconverted; tool converts once before storage), explaining the intent parameter semantics in detail (log vs plan, required fields for log), and specifying that run_type should only be set when the runner or workout structure supports it. It also adds the rule that sensor data and splits must never be invented, which adds semantic guidance beyond the schema's parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('log') and resource ('run card'), states that it is for running on foot only, explicitly excludes cycling/walking/rowing sessions, and distinguishes between completed runs (log) and future run plans (plan). It differentiates itself from the sibling log_workout by naming the alternative and the conditions that select it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'RUNNING ONLY: use log_run only for running on foot,' names log_workout as the alternative for cycling/walking/rowing with the specific focus_type values, and gives clear guidance on intent (log vs plan). It also states when run_type should be classified and warns against labeling unspecified wearable runs as 'easy'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_sleepAInspect

Log a sleep entry. use: logging request or concrete sleep event to record, from a tracker or recall. Question/habit/hypothetical alone: no write.

PROACTIVE DATA COLLECTION: If the user says they want to log sleep but hasn't shared numbers, ask: "How many hours did you sleep, and do you have a sleep score or stage breakdown from your tracker?" They can paste or describe the summary screen.

INFER — do not ask:

  • date: date the primary sleep session ended / wake date (night ending on this date); default to today

You may log any subset of fields. One row per day. Calling this tool twice on the same date updates the existing entry (upsert). Entries made through this tool are always tagged as manual — the wearable-provider sources (Fitbit/Oura/Apple Health) are reserved for the actual auto-sync pipelines.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesDate of the sleep entry (night ending on this date). Format: YYYY-MM-DD. Default to today.
bedtimeNoBedtime / primary sleep session start time. Format: ISO 8601 timestamp (e.g. 2026-08-16T22:47:00-04:00) or HH:MM wall-clock time. Optional.
wake_timeNoWake time / primary sleep session end time. Format: ISO 8601 timestamp (e.g. 2026-08-17T06:21:00-04:00) or HH:MM wall-clock time. Optional.
awakeningsNoNumber of times woken during the night. Optional.
sleep_scoreNoSleep quality score on a 0-100 scale (matches wearable scoring). For a 1-10 self-rating, multiply by 10 first. Optional.
total_hoursNoTotal sleep duration in hours (e.g. 7.5). Optional.
rem_sleep_hoursNoREM sleep in hours. Optional — include if the tracker reports it.
deep_sleep_hoursNoDeep/slow-wave sleep in hours. Optional — include if the tracker reports it.
light_sleep_hoursNoLight sleep in hours. Optional — include if the tracker reports it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses critical behavioral traits beyond the annotations: one row per day, calling twice on the same date is an upsert/update, manual tagging, and the exclusion of wearable-provider sources. These details are not present in the input schema or annotations and materially affect how an agent uses the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but well-structured with clear sections and front-loaded purpose. Every section adds necessary operational detail—proactive data collection, inference rules, upsert behavior, and manual tagging—so the length is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a 9-parameter, mutation-prone tool with an output schema. It covers when to call, what to ask, how to infer the required date, update semantics, and provenance constraints. With the output schema present and annotations provided, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents each parameter. The description adds useful semantic guidance not in the schema, such as inferring the date from the wake date rather than asking, defaulting to today, and allowing any subset of fields to be logged.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb and resource, 'Log a sleep entry', and immediately clarifies what qualifies as a loggable event ('logging request or concrete sleep event') and what does not ('Question/habit/hypothetical alone: no write'). It also distinguishes itself from wearable-provider auto-sync sources by noting entries are always tagged as manual.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear guidance on when to use the tool: when a user requests logging or provides a concrete sleep event, including proactive follow-up when numbers are missing. It provides negative guidance for hypotheticals alone and implies that wearable auto-sync pipelines should not use this tool, though it does not name a specific sibling alternative such as log_wearable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_supplement_takenAInspect

Mark a medication or supplement as taken or not taken for a specific date. Only relevant when the user has daily tracking mode enabled. Use when the user says they took (or missed) a medication or supplement on a particular day.

INFER — do not ask:

  • date: default to today

  • taken: default to true (marking as taken)

SELECTOR — pass supplement_id if known, or supplement_name (case-insensitive substring) to resolve it. Exactly one required. If supplement_name matches more than one item, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoDate (YYYY-MM-DD). Default to today.
notesNoOptional note for this check-in.
takenNotrue = taken, false = missed. Default: true.
supplement_idNoSupplement ID. Alternative to supplement_name.
supplement_nameNoAlternative to supplement_id: name substring, case-insensitive (e.g. "magnesium").

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already signal mutation (readOnlyHint=false). The description adds meaningful behavior beyond that: it documents inference defaults (date=today, taken=true), the requirement that exactly one selector be provided, and the error behavior when supplement_name matches multiple items. This is valuable context for invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized. Purpose, usage condition, inference rules, and selector requirements are each given their own section with no redundant or filler content. Every sentence adds operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity, the 100% schema coverage, the presence of an output schema, and annotations covering mutation safety, the description is complete. It covers prerequisites, defaults, selector resolution, and failure behavior, so an agent can invoke it correctly without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds important semantics beyond the schema: the exact-one-required relationship between supplement_id and supplement_name, the inference rules, and the multiple-match error behavior. This compensates well for the schema's lack of required-parameter enforcement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Mark a medication or supplement as taken or not taken for a specific date.' This clearly distinguishes it from sibling tools like list_supplements or manage_supplement, and the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: 'Use when the user says they took (or missed) a medication or supplement on a particular day.' It also notes the daily tracking mode prerequisite. It does not explicitly name alternatives or state when not to use it, but the guidance is clear enough for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_wearableAInspect

Log daily wearable/manual health metrics (RHR, HRV, Zone Minutes / AZM, VO2max, calories eaten / dietary energy, stress, supplemental steps, and physiological vitals including SpO₂, respiratory rate, skin temperature, blood pressure, blood glucose, and core temperature).

VITALS — use the vital fields for manual/home/device readings and corrections, including a finger-stick, CGM, home glucose meter, or wearable/Apple Health/Health Connect value the user explicitly wants stored manually. A glucose value from an actual lab report or blood draw belongs in log_lab_results instead, not here.

STEPS — read before using step_count: manual step_count is ADDITIVE — it adds on top of whatever a connected wearable (Fitbit, Oura, Apple Health, Health Connect) already recorded that day; it never replaces or overrides device data. Only use it when the user explicitly says they walked steps their device did NOT capture (phone left home, battery died, device not worn). If the user says sync is wrong, steps look doubled, or they want to fix/override/replace device data: do NOT pass step_count — explain that manual steps add on top, and sync issues need investigating at the device level.

ALL OTHER FIELDS (RHR, HRV, AZM, VO2max, stress, and physiological vitals) replace the existing manual value for that day and are safe to use for corrections.

INFER — do not ask: date defaults to today unless the user says otherwise.

IDEMPOTENT: if the values you'd log already exist for that date (any source), the tool returns a no-op success — report this as "already had data", not "failed".

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesDate for the entry. Format: YYYY-MM-DD. Default to today.
resting_hrNoResting heart rate in BPM. Optional.
step_countNoSupplemental steps NOT captured by any connected wearable. ADDITIVE — adds to device data, never replaces it. Only use when user explicitly says their device missed these steps.
hrv_morningNoMorning HRV in milliseconds (RMSSD or similar). Optional.
skin_temp_cNoSkin temperature in °C. Manual/home/device reading. Optional.
spo2_avg_pctNoSpO₂ in %. Manual/home/device reading. Optional.
stress_scoreNoStress score (0–100 scale). Optional.
vo2max_estimateNoVO2max estimate (ml/kg/min). Optional.
core_body_temp_cNoCore temperature in °C. Manual/home/device reading. Optional.
dietary_caloriesNoDietary energy / calories eaten for the day in kcal, imported from a connected health app (Apple Health / Health Connect). Optional.
systolic_bp_mm_hgNoSystolic blood pressure in mmHg. Provide with diastolic_bp_mm_hg. Manual/home/device reading. Optional.
diastolic_bp_mm_hgNoDiastolic blood pressure in mmHg. Provide with systolic_bp_mm_hg. Manual/home/device reading. Optional.
blood_glucose_mg_dlNoBlood glucose in mg/dL. Manual/home/device reading. Optional.
active_zone_min_peakNoZone Minutes (also called Active Zone Minutes by Fitbit, Heart Points by Google Fit) at peak intensity. Optional.
confirm_supplementalNoSet to true ONLY after the user has explicitly confirmed these steps were walked without their device and were not captured. Do not set true based on your own inference.
respiratory_rate_bpmNoBreathing in br/min. Manual/home/device reading. Optional.
active_zone_min_moderateNoZone Minutes (also called Active Zone Minutes by Fitbit, Heart Points by Google Fit) at moderate intensity. Optional.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is highly transparent (additive step semantics, replacement behavior, no-op success), but it explicitly claims 'IDEMPOTENT' while the annotation idempotentHint is false. This is a direct annotation contradiction, so per rubric the score must be 1 despite the otherwise rich detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well structured, with front-loaded purpose and uppercase section headers (VITALS, STEPS, ALL OTHER FIELDS, INFER, IDEMPOTENT). Some enumeration of vitals duplicates the schema, creating mild redundancy, but the operational rules justify most of the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with an output schema and 100% schema coverage, the description covers all critical behaviors: defaulting to today, step additive semantics, replacement scope, no-op idempotent success, and lab-result routing. Only minor gaps exist, such as explicit behavior for omitting all fields, but nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds substantial extra meaning: step_count is additive and requires explicit confirmation, vitals replace existing manual values, systolic requires diastolic to accompany, and glucose routing depends on source. These are the exact semantic constraints the schema does not fully capture.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a specific verb and resource: logging daily wearable/manual health metrics across many vitals. It explicitly distinguishes where lab glucose values belong (log_lab_results instead), which separates it from that sibling without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use and when-not-to-use guidance: manual step_count only for device-missed steps, never for sync issues; lab-derived glucose belongs in log_lab_results; all other fields replace existing manual values. This is model-level routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_wellbeingAInspect

Log subjective wellbeing ratings for a day, week, month, or custom date range. use: logging request or concrete wellbeing report to record. Question, general pattern or venting alone: no write.

Supports single-day entries ("how I feel today") and period entries ("this week was stressful", "March was great").

If an overlapping entry already exists for the requested period, returns a warning with the conflicting entry IDs — the user must update or delete existing entries first.

INFER — do not ask:

  • period_start: default to today

  • period_end: default to same as period_start (single day). For "this week" use Monday–Sunday, for "this month" use first–last day.

  • ratings: estimate from description ("exhausted"=2, "great energy"=8, "stressed out"=8 stress, "feeling good"=7 mood)

You may log any subset of rating fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
moodNoMood 1-10 (1=terrible, 10=excellent). Optional.
notesNoFree-text notes about how you feel. Optional.
energyNoEnergy level 1-10 (1=exhausted, 10=wired). Optional.
stressNoStress level 1-10 (1=calm, 10=overwhelmed). Optional.
sorenessNoMuscle soreness 1-10 (1=none, 10=extreme DOMS). Optional.
period_endNoEnd date of period. Format: YYYY-MM-DD. Default: same as period_start (single day). Use for week/month/custom ranges.
period_startNoStart date of period. Format: YYYY-MM-DD. Default: today.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (only hint flags), so the description carries the burden. It discloses important behavior: returns a warning with conflicting entry IDs on overlap, infers period_start/end and ratings from the user's language, and allows logging any subset of rating fields. This goes well beyond the annotations, though it does not describe the success response format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about 150 words and organized into clear segments: usage, period types, conflict behavior, and inference rules. It is front-loaded with the core purpose and uses bullet-like lines for inference. Some redundancy (e.g., repeating defaults) is present, but overall each section earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 optional parameters, a full schema, and an output schema (signal indicates it exists), the description covers the key decision points: when to log, how to handle overlaps, and how to infer unstated parameters. It does not mention authentication or rate limits, but these are not relevant for a client-side logging tool in this context. The agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema documents all 7 parameters. The description adds value by specifying inference rules for period_start/end (default today, week = Monday–Sunday, month = first–last day) and for ratings (e.g., 'exhausted'=2, 'great energy'=8). It also clarifies that any subset of rating fields may be logged. This supplements the schema meaningfully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Log') and a specific resource ('subjective wellbeing ratings') and further clarifies the scope: day, week, month, or custom range. This distinguishes it from sibling tools like update_wellbeing (modify) and list_wellbeing (read), even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'use: logging request or concrete wellbeing report to record' and excludes 'Question, general pattern or venting alone: no write.' That is clear when-not guidance. It does not explicitly name alternative tools (e.g., update_wellbeing for editing), but the positive and negative conditions are sufficient for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_workout
Destructive
Inspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

COMPLETED WORKOUTS, ANY ACTIVITY:

  • With actual exercises/sets, provide them as usual. Session totals and activity metrics can accompany real sets in the SAME log_workout call.

  • With only an activity description (walking, yoga, swimming, a kettlebell circuit, strength, HIIT, cycling, etc.), give focus_type and reported session details; OMIT exercises or pass []. A session does not require fake exercises or sets. Never invent exercises, reps, weights, splits, heart rates or calories.

  • Use notes for reported context, including uncertain durations (e.g. "15-20 minutes, light intensity"). Leave unknown numeric fields unset rather than converting a range into false precision. Mark estimated=true ONLY when numeric values really are estimates; it is not implied by having zero sets.

  • Walking: focus_type "Walking"; cycling/biking: "Cycling"; rowing/RowErg: "Rowing". Other activity names remain as stated. log_run remains the richer dedicated route for completed runs with known distance+duration, running splits or planned runs. Never send a walk, ride or row to log_run.

  • distance_mi/distance_km/distance_meters are alternatives; send the unit the user provided. The server converts, never convert it yourself. duration_sec is total workout/moving seconds, not a set or a repetition. Use reported HR, power, cadence, elevation, actual laps/segments only when known.

Log a complete workout session: exercises, sets, reps, weights, and session metadata. Use when the user describes finishing a workout, lists exercises performed, or asks to log training. workout_to_do: in-app chat uses propose_workout, including save/start requests; discover it if needed. External clients: see SAVED WORKOUT MODE.

EXERCISE NAMES:

  • Reuse known canonical names; otherwise call list_exercises and match each exercise to the closest canonical name. No reasonable match → use the name as stated. Don't ask before logging, match silently and log.

  • "Chest press" (machine) and "bench press" (barbell) are DISTINCT — pass the user's term through so the resolver's aliases pin the right one.

  • name is ONLY the exercise name, never reps/weights/sets — those go in the sets array.

  • LITERAL NAME: literal_name: true keeps the user's exact wording instead of the closest library match, skips the resolver, and gets no NSI score (no benchmark to compare an unmatched name against). Use for "call it exactly X", "not the standard one", "literally X", or a rejected match.

  • The result says when a name was matched to something other than what the user said. Relay it in your own words rather than repeating the line verbatim. If a name matches nothing closely enough, the result names near-miss library exercises; ask the user which they meant rather than accept the unscored custom log silently.

  • EQUIPMENT (load basis): dumbbell_pair is one dumbbell in EACH hand, weight_lb PER HAND (2x for NSI); dumbbell_single is one implement total. Laterality (single-leg/arm) does NOT decide this alone. Set it when the user describes the load (each hand, machine, band); a wrong or missing tag silently halves or doubles NSI. Values: barbell, dumbbell_pair, dumbbell_single, machine, kettlebell, bodyweight, band, cable, trx, other.

SETS:

  • "3 sets of 15 reps" → 3 set objects with reps: 15. "15/12/10" → 3 sets with reps 15, 12, 10.

  • Pure isometric holds (planks, dead hangs, wall sits) have no reps: "30 second plank" = { hold_length_sec: 30 }.

  • Tempo/pause work combines reps + weight_lb + hold_length_sec (seconds per rep) on the same set, never in notes.

  • Loaded carries (farmers carry, sled push, weighted plank) are one set per trip: hold_length_sec + weight_lb, omit reps unless a trip count is given. weight_lb is PER HAND for a two-implement carry, TOTAL for one implement. Distance has no column and is never a duration — put it in notes.

INFER — do not ask:

  • date: today, or from context

  • focus_type: from the exercises (bench/shoulders/triceps=Push, rows/pulldowns/curls=Pull, squats/deadlifts/lunges=Legs, mixed upper=Upper, everything=Full Body)

  • is_bodyweight: true for pull-ups, push-ups, dips, bodyweight squats; missing load alone does not mean bodyweight

  • superset_group: same integer for exercises done back-to-back or as a superset

  • slot_type: 'warmup' for prep at the start, 'finisher' for burnout/cardio at the end, 'working' (default) otherwise

RPE (Rate of Perceived Exertion) — 1-10 scale, half steps allowed (7.5, 8.5):

  • Session-level RPE: overall session difficulty. Infer from user comments like "brutal session" (8-9) or "easy day" (3-4). Optional.

  • Per-set RPE: how hard each individual set felt. Include only if the user explicitly mentions per-set effort or failure. Optional.

  • Scale: 1=minimal effort, 5=moderate, 7=hard, 8=few reps left, 9=one rep left, 10=maximal/failure.

optional_metadata: infer location/focus_type when supported, else omit; heart_points: reported values only. Missing external load: follow LIVE LOADS.

RETURNS the new session's ID (as "[ID NN]"). Pass it to update_workout / delete_workout / get_workout for follow-ups in this conversation.

LIVE LOADS: for user-driven completed-workout logging, obvious bodyweight movements may omit load and are inferred as bodyweight. If a performed set normally uses external resistance, include that set's load. If the user did not provide it, ask one batched clarification before calling this tool. Only set load_unknown=true on a missing set when the user explicitly says they do not know, do not remember, or want to save without that load. Preserve partial known loads positionally; never copy one set's load to another or treat a missing external load as bodyweight.

SAVED WORKOUT MODE: in-app chat uses propose_workout for new workout plans, whether requested to save or start; discover it if needed. External clients without cards: set save_as_saved_workout=true and pass the full prescription through this same tool. In saved mode the workout is NOT logged as completed history, and LIVE LOADS does not apply: prescribed weight may be omitted. To replace an existing Saved Workout, also pass saved_workout_id; the supplied prescription fully replaces its prior prescribed sets/reps/weights. Use saved_workout_title when the reusable name should differ from focus_type.

ParametersJSON Schema
NameRequiredDescriptionDefault
rpeNoSession RPE, 1-10, half steps allowed: 5 moderate, 7 hard, 9 one rep left, 10 failure. Infer from comments about overall difficulty, or omit.
dateYesYYYY-MM-DD. Default to today.
notesNoSession notes: how it went, PRs hit, how they felt.
avg_hrNoReported average heart rate (bpm).
max_hrNoReported maximum heart rate (bpm).
splitsNoActually reported laps/splits, same structure as the in-app run logger; never invent.
caloriesNoReported session calories; never invent.
locationNoGym, Home, Outdoor. Infer from context, else omit.
segmentsNoActually performed intervals/blocks, same structure as the in-app run logger; never invent.
estimatedNoTrue only when a saved numeric value is an AI approximation; describe its source/range in notes.
exercisesNoOnly the actual exercises/sets provided by the user. For a session-level log, omit or use [].
intensityNoQualitative effort as reported (does not manufacture an RPE).
focus_typeNoSpecific completed activity (Walking, Cycling, Yoga, Kettlebell Circuit, Full Body, etc.). Infer from user wording; omit only if notes describe the activity.
start_timeNoLocal HH:MM when explicitly supplied; otherwise omit.
avg_power_wNoReported average power in watts (e.g. cycling).
distance_kmNoSession distance in kilometers. Server converts to miles.
distance_miNoSession distance in miles, when supplied in miles. In mi, or km with input_distance_unit set. See UNIT INPUTS.
elapsed_secNoElapsed seconds including pauses; omit if unknown.
max_power_wNoReported peak power in watts.
duration_secNoTotal workout or moving time in seconds. Convert user-stated minutes to seconds; not a set duration.
avg_cadence_rpmNoReported cycling/rowing cadence, revs/min.
avg_cadence_spmNoReported running/walking cadence, steps/min.
distance_metersNoSession distance in meters. Server converts to miles.
elevation_gain_mNoReported elevation gain in meters.
saved_workout_idNoExisting Saved Workout ID to replace in saved mode. Omit to create a new Saved Workout.
heart_points_peakNoPeak-intensity heart points, if mentioned.
avg_pace_sec_per_miNoReported average seconds per mile. Never invent a pace. In mi, or km with input_distance_unit set. See UNIT INPUTS.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.
saved_workout_titleNoOptional reusable workout name in saved mode. Defaults to focus_type.
heart_points_moderateNoModerate-intensity heart points, Google Fit or equivalent, if mentioned.
save_as_saved_workoutNoTrue when this payload is a reusable Saved Workout prescription, not a completed workout. Defaults to false.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.
manage_locationAInspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Create a training location, or update one: rename it, change its kind, make it the default, or add and remove equipment. Use when the user says where they train, what a place has ("my hotel has dumbbells to 50 and a bench"), or that equipment is missing or new. Call list_locations first to get ids and avoid duplicates.

create: name required. kind defaults to other. Equipment starts from preset (default follows kind: home -> home_gym, gym -> full_gym, hotel -> hotel_gym, else bodyweight) unless add is given, which replaces the preset. The user's first location becomes the default. update: location required (id or exact name). Pass any of name, kind, make_default, add, remove. add upserts an item (with sizes when the user stated them); remove takes items out. Bodyweight is always available and never listed.

Writes immediately to the user's saved location, which Jamie and the On Deck coach read for every plan. There is no delete here.

ParametersJSON Schema
NameRequiredDescriptionDefault
addNoEquipment to add or replace by item.
kindNo
nameNocreate: the new location name. update: a new name to rename it.
presetNocreate: starting equipment. Ignored when add is given.
removeNoEquipment item ids to take out.
locationNoupdate: the location id or exact name from list_locations.
operationYes
make_defaultNoMake this the default location used when none is named.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds substantial context beyond them: writes are immediate, downstream consumers (Jamie and the On Deck coach) read the value for every plan, the first location becomes default, bodyweight is always available and never listed, and there is no delete path. This is rich disclosure that does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with labeled sections (UNIT INPUTS, create, update) and dense, information-bearing sentences. Slightly dense and the UNIT INPUTS block leading the description delays the top-line purpose, but almost every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with 8 params, an output schema present, and annotations covering safety, the description supplies everything an agent needs: create/update branch behavior, defaults, unit handling, dedup guidance, and the no-delete boundary. Return values are covered by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 75% schema coverage the schema already documents many fields, but the description adds real meaning: create vs update semantics for name/kind/location, preset defaults per kind and replacement by add, add upsert behavior with sizes, remove semantics, and the UNIT INPUTS precedence rule. This goes well beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Create a training location, or update one') and enumerates the exact operations: rename, change kind, set default, add/remove equipment. It clearly distinguishes itself from the read-only sibling list_locations, which it explicitly references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete triggering conditions ('Use when the user says where they train, what a place has... or that equipment is missing or new') and an explicit alternative to call first ('Call list_locations first to get ids and avoid duplicates'). It also states a boundary ('There is no delete here'), covering when-not semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_recovery_strategyA
Destructive
Inspect

Add, update, end, or delete a recovery/mindfulness strategy. Use when the user describes a new practice, changes a schedule, stops a practice, or removes one. Infer category from name, start_date defaults to today, infer schedule from context. ASK only if name is missing.

SELECTOR for update/end/delete — pass id if known, or strategy_name (case-insensitive substring, e.g. "sauna") to resolve it. Exactly one of id or strategy_name required. If strategy_name matches more than one strategy, the call errors with candidate IDs to retry with.

AFTER a successful 'add': do NOT just confirm and stop. Reply by (1) restating the assumed schedule (sessions per period, duration, time of day, start date) in plain language, and (2) asking the user to confirm or correct it — especially any optional fields you did NOT set (duration_minutes, time_of_day). Example: "Logged sauna starting today, assuming once per week. Sound right? About how long do you usually go for, and what time of day — morning, evening?" If the user corrects anything, call this tool again with action='update'. The goal is accurate adherence data, not a silent confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoStrategy ID. Required for update, end, delete unless strategy_name is given.
nameNoStrategy name (e.g. 'Box Breathing'). Required for add; on update, sets a new name.
notesNoFree-text notes. Optional.
actionYesWhat to do. Required.
categoryNoCategory. Infer from name.
end_dateNoEnd date (YYYY-MM-DD). Default to today for end action.
start_dateNoStart date (YYYY-MM-DD). Default to today for add.
period_unitNoPeriod unit. Default: 'week'.
time_of_dayNoWhen during the day: ['morning'], ['evening'], etc. Optional.
strategy_nameNoAlternative to id for update/end/delete: strategy name substring, case-insensitive (e.g. "sauna").
duration_minutesNoTarget minutes per session. Optional.
sessions_per_periodNoTarget sessions per period. Default: 1.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only and destructive hints, so the bar for disclosure is lower. The description adds useful behavioral detail beyond that: start_date defaults to today, category is inferred from the name, strategy_name is a case-insensitive substring, and ambiguous matches produce an error with candidate IDs to retry with.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into clear sections: trigger cases, selector rules, and post-add follow-up. It stays front-loaded and every block earns its place, though the post-add paragraph is fairly verbose and could be trimmed without losing the required conversational loop.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter mutation tool with one required field in schema, this description covers the most likely failure points: how to resolve strategy names, what to infer, what defaults to use, how to handle errors, and what to do after a successful add. The output schema exists, so the description does not need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though the schema has 100% parameter coverage, the description adds critical semantics the schema does not express: exactly one of id or strategy_name is required for update/end/delete, name is the only truly required field for add, and optional fields like duration_minutes and time_of_day should be confirmed after an add. This is highly actionable guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line uses a specific verb-resource pairing: 'Add, update, end, or delete a recovery/mindfulness strategy,' and covers the full set of actions. This distinguishes the tool from session-level siblings like log_recovery_session and list_recovery_strategies by making it clear this tool manages the strategy lifecycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete trigger cases: when the user describes a new practice, changes a schedule, stops a practice, or removes one. It also says to ASK only if the name is missing, but it does not explicitly name alternative siblings or state negative 'do not use this when' guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_supplementA
Destructive
Inspect

Add, update, end, or delete a medication or supplement. Use when the user describes their stack, adds a new item, changes a dose or schedule, says they stopped taking something, or wants to remove an entry.

INFER — do not ask:

  • action: 'add' for a new item, 'update' for changing a field, 'end' when they stopped/finished a course, 'delete' only to remove the record entirely

  • category: 'medication' for prescription/OTC drugs and pharmaceuticals, 'supplement' for vitamins/minerals/herbs/other dietary supplements — default 'supplement' if unclear

  • start_date: today for new entries

  • end_date (for 'end'): today unless the user specifies otherwise

ASK the user only if name is missing for a new entry, or 'end' (set end_date) vs 'delete' (remove record) intent is ambiguous.

SELECTOR for 'update', 'end', 'delete' — pass id if known, or supplement_name (case-insensitive substring, e.g. "magnesium") to resolve it. Exactly one of id or supplement_name required. If supplement_name matches more than one item, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoID. Required for update, end, delete unless supplement_name is given.
formNoPhysical form: pill, capsule, tablet, softgel, powder, liquid, gummy, other. Optional.
nameNoName (e.g. 'Magnesium Glycinate' or 'Metformin'). Required for add; on update, sets a new name.
brandNoBrand name. Optional.
notesNoFree-text notes. Optional.
actionYesWhat to do. Required.
categoryNoCategory: 'medication' for drugs/pharmaceuticals, 'supplement' for vitamins/minerals/herbs. Default: 'supplement'.
end_dateNoEnd date (YYYY-MM-DD). Set for 'end' action — default to today. Null means currently active.
dose_unitNoUnit for dose_amount: pills, capsules, tablets, softgels, g, mg, ml, IU, mcg, tbsp, scoop. Required for add.
start_dateNoStart date (YYYY-MM-DD). Required for add — default to today.
unit_labelNoLabel for dose_per_unit (e.g. 'mg', 'IU'). Optional.
dose_amountNoNumeric dose quantity (e.g. 2 for '2 pills'). Required for add.
period_unitNoThe period for times_per_period. Optional.
time_of_dayNoWhen during the day: ['morning'], ['morning','evening'], ['night'], etc. Optional.
dose_per_unitNoAmount per individual unit (e.g. 240 for '240mg per pill'). Optional.
frequency_typeNoFrequency category. Optional — default 'daily'.
supplement_nameNoAlternative to id for update/end/delete: name substring, case-insensitive (e.g. "magnesium").
times_per_periodNoHow many times per period (e.g. 2 for twice per week). Optional — default 1.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses important behavior: actions are inferred, 'end' sets end_date while 'delete' removes the record, and a non-unique supplement_name causes an error with candidate IDs. This gives a clear model of side effects and failure modes, and there is no contradiction with the destructiveHint annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but deliberately structured: purpose, trigger conditions, inference rules, ask conditions, and selector behavior are each clearly separated, with the core verb and resource front-loaded. Every sentence earns its place given the tool's 18-parameter complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with conditional selector requirements, the description covers user-intent mapping, defaults, disambiguation, and error behavior. Since an output schema is present, return values do not need to be spelled out, and nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all 18 parameters at 100% coverage, but the description adds substantial value by specifying inference rules, defaults for start_date, end_date, category, and frequency, and the id/supplement_name selector contract. This is meaningfully more than what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific, multi-operation statement—'Add, update, end, or delete a medication or supplement'—and clearly identifies the resource being managed. It then gives concrete user-intent triggers, making it easy to distinguish from read-only siblings like list_supplements or log_supplement_taken.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Use when...' sentence provides explicit triggering context, and the INFER/ASK guidance tells the agent when to act autonomously versus when to seek clarification. It does not explicitly name sibling alternatives or state when not to use this tool, but the operational guidance is strong enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_empty_dayA
DestructiveIdempotent
Inspect

Set, change, or undo the answer to "why is this day empty?" for a date with no real meals logged. Use when the user wants to flip a fast day to forgotten (or back), or undo either one, in chat instead of the in-app prompt.

There are exactly three states for a date, and this tool is the only way to move between them:

  • fast: writes the 0-kcal "Fast day" food_log entry (a real, counted 0-calorie day).

  • forgot: records that the day was reviewed and simply not logged. Writes nothing to food_log, so the day stays a true blank and is excluded from calorie/TDEE averages, never imputed as 0.

  • unanswered: clears both. The day goes back to being an open question and the in-app prompt may ask about it again.

Setting one answer always clears the other, so a date is never both a fast and a forgotten day at once.

Use list_meals first if unsure whether the date already has real food logged. This tool refuses to touch a day that has actual meals on it (other than an existing fast marker).

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesThe date being answered for. Format: YYYY-MM-DD. Required.
answerYesRequired. "fast" = intentional 0-calorie day. "forgot" = day stays excluded, not imputed as 0. "unanswered" = clear any prior answer.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description thoroughly explains the side effects: writing a 0-kcal food_log entry for 'fast', leaving the day blank for 'forgot', and clearing both for 'unanswered'. It also notes that it refuses to touch days with real meals. This goes beyond the annotations (readOnlyHint=false, destructiveHint=true) by detailing exactly what changes occur to the data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is clear but somewhat repetitive, repeating the three-state definitions twice and the 'setting one clears the other' rule twice. It could be tightened without losing meaning, but it remains organized and not excessively verbose for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers prerequisites (use list_meals first), parameters, exact effects, and failure conditions (refuses days with real meals). It even notes that the tool is for chat usage instead of the in-app prompt. No critical context is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters (date, answer) are described in the schema with full coverage. The description further elaborates on the 'answer' enum values—'fast' means intentional 0-calorie day, 'forgot' means day stays excluded, 'unanswered' clears any prior answer—adding richer semantics beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Set, change, or undo the answer to why is this day empty?' It specifies the resource (a date) and the exact states (fast, forgot, unanswered) it manages. It also distinguishes itself as the only way to move between these states, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is provided: 'Use when the user wants to flip a fast day to forgotten (or back), or undo either one, in chat instead of the in-app prompt.' It also advises to 'Use list_meals first if unsure whether the date already has real food logged' and mentions refusal when real meals exist, giving clear when-to-use and when-not-to-use instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_nutrient_targetA
Idempotent
Inspect

Set or remove the user's OWN daily target for one vitamin, mineral or other tracked nutrient. Use ONLY when the user explicitly asks to set, change or remove their own target ("set my vitamin D target to 50 mcg"). Never call it on your own initiative, to apply a recommendation, or to act on a lab result. A user target is the user's personal preference, never a clinical prescription; do not present it as one. It replaces the reference intake for that nutrient in every nutrient view. Calories and macros are goals, not nutrient targets.

REQUIRED: target_value and unit unless remove is true. Ask for the unit if the user gave none; never guess mcg vs IU or mg vs g. An omitted upper_limit means the reference upper limit applies. remove=true deletes the user's target so the reference intake applies again.

nutrient accepts a name/alias, or the id (1-44) from get_nutrient_summary's nutrient id dictionary -- pass whichever a prior get_nutrient_summary/get_nutrient_history/get_nutrient_contributors compact reply already gave you.

ParametersJSON Schema
NameRequiredDescriptionDefault
unitNoUnit the user stated, e.g. mg, mcg, IU, g.
notesNoWhy the user set it, in their words. Kept when omitted.
removeNotrue deletes the user's target for this nutrient. Default false.
nutrientYesNutrient name or alias, e.g. "vitamin_d", "magnesium", "b12" -- or its dictionary id (1-44, e.g. "24" for vitamin_d).
upper_limitNoThe user's own daily ceiling in unit, at least target_value.
target_valueNoDaily target amount in unit.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false), it discloses consequential behavior: the target 'replaces the reference intake for that nutrient in every nutrient view,' remove=true restores the reference intake, and an omitted upper_limit falls back to the reference upper limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the hardest constraint (only on explicit request), then grouped into a schema-behavior paragraph and a nutrient-identifier paragraph. It is longer than most tool descriptions but nearly every sentence carries a distinct constraint, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description covers the remaining risk surface for a 6-param write tool: required-field conditions, unit ambiguity, default semantics for upper_limit, removal behavior, and cross-tool identifier sourcing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning: it warns never to guess unit (mcg vs IU, mg vs g), explains that remove=true deletes the target, and tells the agent nutrient may be an id from get_nutrient_summary's dictionary. It stops short of restating notes/upper_limit semantics in depth, so it sits just above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Set or remove the user's OWN daily target for one vitamin, mineral or other tracked nutrient') and explicitly demarcates itself from siblings: 'Calories and macros are goals, not nutrient targets,' ruling out create_goal/update_goal and set_supplement_nutrients.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when ('ONLY when the user explicitly asks to set, change or remove their own target') and when-not ('Never call it on your own initiative, to apply a recommendation, or to act on a lab result'), plus the framing constraint that a target is a preference, not a prescription.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_supplement_nutrientsA
Idempotent
Inspect

Set the nutrient content of ONE configured dose of a medication or supplement, as printed on its label. Use when the user shares a Supplement Facts / Nutrition Facts panel, states specific nutrient amounts from a label, or corrects a previously stored value.

VALUES MUST COME FROM THE PRODUCT LABEL OR THE USER'S EXPLICIT STATEMENT. Never estimate or guess a nutrient amount. If the user hasn't given real label numbers, ask for the label instead of calling this.

Amounts are per the dose already configured on the row (dose_amount/dose_unit, e.g. "2 capsules"), not per individual unit and not a daily total. Daily totals are derived separately from frequency_type/times_per_period. A nutrient name or unit this tool doesn't recognize is skipped and reported back in the result rather than silently dropped.

SELECTOR: pass id if known, or name: resolved as an exact case-insensitive match, then a unique case-insensitive prefix match. Ambiguous or unmatched name errors listing candidate ids, without writing. Exactly one of id or name is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoSupplement/medication ID. Alternative to name.
nameNoAlternative to id: resolved as an exact match if one exists, else a unique prefix match, case-insensitive.
replaceNotrue (default) replaces all stored nutrients for this item; false merges these into the existing set, overwriting only the given nutrient keys.
nutrientsYesThe label's Supplement Facts per serving, for one configured dose. See the VALUES MUST COME FROM THE PRODUCT LABEL note above.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations (readOnly=false, idempotent=true, destructive=false): unrecognized nutrient names/units are skipped and reported in the result rather than silently dropped, and ambiguous or unmatched names error out listing candidate ids without writing. It also clarifies the dose-basis semantics and the replace-vs-merge default that the schema only sketches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the strongest constraint (values must come from the label) are front-loaded, and the paragraphing separates values, dose basis, and selector logic. It runs long for four parameters, and the all-caps warning in mid-text is a touch heavy, but nearly every sentence carries actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description still notes the useful result signal (skipped nutrients reported back). Guidance, guardrails, dose semantics, and selector/error behavior are all covered for a 4-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description goes further: amounts are per the already-configured dose (dose_amount/dose_unit), not per individual unit and not a daily total, and it spells out the selector resolution order (id, else exact case-insensitive name, else unique prefix). Both points reinforce, rather than merely restate, the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+scope: setting nutrient content for ONE configured dose of a supplement/medication, as printed on the label. This clearly separates it from siblings like set_nutrient_target (daily targets), manage_supplement (supplement CRUD), and log_supplement_taken. An agent can identify the tool from the first sentence alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit triggering conditions are given (user shares a Supplement Facts panel, states label amounts, or corrects a stored value) plus a hard exclusion: never estimate, ask for the label instead if the user hasn't given real numbers. The gap is that it never names the sibling to use instead for adjacent tasks (e.g. manage_supplement to create the item, set_nutrient_target for daily targets).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_body_compositionA
Read-onlyIdempotent
Inspect

Show body composition over time with an interactive metric picker for weight, body fat, lean and muscle mass, hydration, visceral fat, BMI, waist, and related scale metrics, with a 7-day/30-day/90-day/1-year range toggle. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 90d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds genuinely new behavioral context beyond them: the tool renders an interactive MCP app, exposes a metric picker plus a range toggle, and 'Returns a short text summary alongside the visual view.' It does not cover pagination or data-freshness details, which is minor here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the capability and followed by the usage preference; the long metric enumeration is dense but each item is meaningful. No filler beyond the somewhat lengthy metric list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and it still gives a one-line note about the text-plus-visual return. For a single-parameter, read-only visualization tool, everything an agent needs to select and invoke it is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'range' parameter is fully documented with its enum and default in the schema. The description repeats the 7d/30d/90d/1y options but omits the 90d default and adds no syntax or behavioral nuance, so it does not surpass the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb ('Show') with a well-defined resource ('body composition over time') and enumerates the covered metrics and range toggle, so the agent knows exactly what data it surfaces. It does not explicitly contrast itself with close siblings like show_body_weight or list_body_metrics, leaving differentiation implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear directive: 'When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text.' That establishes the triggering context and a rendering preference, but it names no alternative tools or exclusions, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_body_weightA
Read-onlyIdempotent
Inspect

Visual weight and body-fat trend for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds useful return-format context by stating it returns an interactive chart plus a short summary, which goes beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the tool's purpose, states the available time windows, and communicates the output format efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a low-complexity tool with one optional enum parameter, rich annotations, and an output schema. The description covers the core behavior and return format, leaving no significant gaps for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter, and the schema already fully describes it with an enum and a clear description. The description repeats the same range values without adding new semantic meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool visualizes weight and body-fat trends across specific time ranges and returns a chart plus summary. It does not explicitly distinguish itself from the sibling show_body_composition, which could overlap on body-fat, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for viewing weight and body-fat trends, and the enum ranges suggest when it can be used. However, there is no explicit guidance on when to choose this over alternatives like show_body_composition or list_body_metrics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_exercise_progressionA
Read-onlyIdempotent
Inspect

Show the user's estimated 1-rep-max progression for a lift over time as an interactive line chart, with filters for date range and muscle group, and an exercise picker. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y.
exerciseNoExercise name to chart (e.g. "Bench Press"). Optional; defaults to the most-logged lift in the window.
muscle_groupNoOptional muscle-group filter for the exercise picker (e.g. "Chest", "Legs").

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark the tool as read-only, idempotent, and non-destructive, and the description does not contradict any of these. It additionally discloses the output format (short text summary alongside visual view) and the UI preference, making the behavior fully transparent without adding any misleading side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundant words. It packs the key purpose, filtering capabilities, UI preference, and output summary into a compact but complete statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives the agent enough context to decide when to call the tool (when the user asks about progression) and what to expect (interactive chart, summary text). It also differentiates itself from text-based alternatives, covering the essential contextual needs for correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides detailed descriptions for all three parameters, including the enum for range and the default behavior for exercise. The function description reinforces their purpose as filters/picker, and the schema coverage is 100%, ensuring the agent understands each parameter's role.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the function shows 1-rep-max progression over time as an interactive line chart with filters for date range and muscle group, and an exercise picker. This makes the tool's purpose unambiguous and distinct from other workout-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance: when the user asks about this topic, prefer this tool and render the interactive app rather than describing rows in text. This directly tells the agent when and how to invoke the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_health_overviewA
Read-onlyIdempotent
Inspect

Show a rich overview of wearable health signals including steps, Zone Minutes, resting heart rate, HRV, VO2max, and stress, with a 7-day/30-day/90-day/1-year range toggle. Prefer this for broad wearable or overall health-trend questions. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 30d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it read-only, idempotent, and non-destructive, so the safety profile is covered. The description adds genuinely useful behavior beyond that: it renders an interactive MCP app rather than raw rows and returns a short text summary alongside the visual view, which materially affects how an agent should present results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The metric list and range toggle are front-loaded in the first sentence, with usage and rendering guidance following. It is efficient overall, though 'When the user asks about this, prefer calling this tool' mildly restates the preceding preference sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and annotations carry the safety profile. The description covers purpose, usage, and the interactive rendering behavior; the only omission is explicit routing to the narrower week-level siblings an agent might otherwise pick.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single enum-constrained parameter, so the schema already documents 'range' and its 7d/30d/90d/1y values. The description restates the same enum values without adding syntax, defaults, or semantics beyond the schema — the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Show') and resource ('health overview') and enumerates the metrics covered (steps, Zone Minutes, resting HR, HRV, VO2max, stress) plus the range toggle. It gestures at differentiation by framing itself as the 'broad' wearable/health-trend tool versus the narrower siblings, but never names an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use condition ('Prefer this for broad wearable or overall health-trend questions') and a rendering preference over text output. It does not state when NOT to use it or name the competing narrower tools (e.g. show_week_steps), so the routing guidance is directional but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_meal_diaryA
Read-onlyIdempotent
Inspect

Show the meals logged on a day as a rich diary with daily calories and macros versus targets. Prefer this for what-did-I-eat and daily food-log review questions. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoDiary date (YYYY-MM-DD). Optional; defaults to today.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey read-only, idempotent, and non-destructive behavior. The description adds useful behavioral context by noting that it returns 'a short text summary alongside the visual view' and that it renders an interactive MCP app.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: what the tool does, when to prefer it, and what it returns. The key purpose is front-loaded and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter read-only tool with a full output schema, the description is complete. It explains the visual rendering behavior, the summary return, and the intended use case, leaving no important gap for an agent deciding to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents the only parameter (`date`) with type, format, optionality, and default behavior. The description adds no additional parameter-level meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Show'), a specific resource ('meals logged on a day'), and a distinctive format ('rich diary with daily calories and macros versus targets'). This clearly separates it from sibling tools like list_meals and log_meal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to prefer this tool for 'what-did-I-eat and daily food-log review questions' and instructs rendering the MCP app over describing rows in text. It establishes clear usage context, though it does not explicitly name sibling alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_recoveryA
Read-onlyIdempotent
Inspect

Visual resting-heart-rate and HRV trend for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description is consistent with them. The description adds that output is an interactive chart plus summary, but provides no further behavioral detail such as authentication requirements or default behavior when no range is supplied. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence that front-loads the core purpose and enumerates all valid range options. There is no redundant phrasing or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only visualization tool with an output schema, the description largely suffices: it names the metric, the range choices, and the return form. It does not say what happens when optional range is omitted, which is a minor but real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the only parameter, and the range enum is fully documented. The description merely repeats the allowed values (7d, 30d, 90d, 1y) without adding new meaning or clarifying the optional default behavior, so it stays at the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Visual resting-heart-rate and HRV trend' and 'Returns an interactive chart plus short summary.' It is clear enough to distinguish from list/log siblings, but it does not explicitly differentiate itself from sibling visualization tools like show_health_overview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the intended use case: retrieving a visual RHR/HRV trend over a selected time window. It does not provide explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives, so the agent must infer the right context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_runsA
Read-onlyIdempotent
Inspect

Visual running-mileage trend for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 30d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, idempotent, and non-destructive hints. The description adds useful behavioral context beyond annotations by disclosing the output form (interactive chart plus short summary) and the supported time ranges. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The main capability and time ranges are front-loaded, and the output format is stated second. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one optional parameter and an output schema, the description is nearly complete. It clearly communicates purpose and return format; the only minor gap is not explicitly stating the 30d default, though the schema already covers it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter is fully documented in the schema. The description merely repeats the enum values without adding new semantic detail such as how the chart renders or what the summary includes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Visual running-mileage trend' with explicit time windows (7d, 30d, 90d, 1y). This clearly distinguishes it from siblings like list_runs, log_run, and other show_* tools by emphasizing the visual chart output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it—when the user wants a visual mileage trend—but it does not explicitly state when not to use it or name alternatives such as list_runs for raw data. The context is clear but exclusions and alternative routing are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_sleep_detailA
Read-onlyIdempotent
Inspect

Show one night of sleep in detail with duration, score, stages, bedtime, wake time, awakenings, and recent-night context. Prefer this for last-night or specific-night sleep questions. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoNight ending date (YYYY-MM-DD). Optional; defaults to the most recent sleep entry.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly and destructive annotations, the description discloses the output behavior (returns a short text summary alongside a visual view) and the preference to render an interactive app, giving full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet covers purpose, usage, and output in three sentences with no redundant or vague phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It provides sufficient context for an agent to decide when to use the tool, what it returns, and how to invoke it. The included field list compensates for the absence of an explicit output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'date' is fully described with format (YYYY-MM-DD), optionality, and default behavior (most recent sleep entry), providing complete semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: showing one night of sleep in detail with specific attributes (duration, score, stages, etc.) and identifies it as preferred for last-night or specific-night queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly gives when-to-use guidance ('prefer this for last-night or specific-night sleep questions') and instructs to render the interactive MCP app instead of describing rows in text, making alternatives implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_week_fit_scoreA
Read-onlyIdempotent
Inspect

Visual Fit Score trend and component breakdown for 7d, 30d, 90d, or 1y. Returns an interactive card plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 7d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral detail by specifying the return type: an interactive card plus short summary. This goes beyond the structured annotations and helps set output expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the core purpose and time options, then specify the output format. Every word earns its place; there is no filler or redundant restating of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter, read-only visualization tool with an output schema, annotations, and a clear return description, nothing essential is missing. An agent has enough to invoke it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, including the enum values and default for 'range'. The description repeats the same time windows and does not add extra semantic detail beyond what the schema already provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Visual') and a clear resource ('Fit Score trend and component breakdown'), and it enumerates the supported time windows. It clearly distinguishes this from the sibling tools by naming the unique Fit Score capability rather than a generic health or recovery view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for use: viewing Fit Score trends and components over selectable ranges. It does not explicitly name alternatives or state when-not-to-use, but the Fit Score focus is distinctive enough among siblings that an agent can infer the right selection without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_week_macrosA
Read-onlyIdempotent
Inspect

Visual calories and macros versus targets for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 7d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already convey read-only, idempotent, and non-destructive behavior. The description adds that it returns an interactive chart and summary, but provides no other behavioral details such as data freshness or error cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences, front-loaded with the core purpose and then the return format. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one parameter, the description is sufficient. It could be slightly more complete by clarifying what 'macros' includes or that targets are user-defined, but those are not essential for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already explains the range parameter with its enum and default. The description repeats the enum values in context but does not add significantly deeper meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool visualizes calories and macros versus targets across selectable time ranges, which distinguishes it from sibling tools like show_week_sleep or show_week_steps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by describing the visualization, but it does not explicitly state when to use this tool versus other nutrition-related tools (e.g., list_meals or show_meal_diary), nor does it mention any prerequisites or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_week_sleepA
Read-onlyIdempotent
Inspect

Visual sleep duration and score trend for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 7d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the non-obvious behavioral detail that it returns an interactive chart plus a short summary, rather than raw JSON. No contradictions with annotations were found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first covers purpose and ranges, the second covers return value. There is no filler or redundant restating of the tool name. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only chart tool with one optional parameter, an output schema, and clear annotations, the description is complete. It specifies the visual nature, the selectable time windows, and the expected return shape. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the range parameter and its default are already fully documented. The description repeats the allowed values without adding additional syntax, format, or edge-case guidance. Per baseline, a 3 is appropriate when the schema handles parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (visualize/show) with a clear resource (sleep duration and score trend) and scope (7d, 30d, 90d, or 1y). This distinguishes it from list_sleep and show_sleep_detail, which imply raw or detailed views. The 'interactive chart plus short summary' makes the visualization purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Visual sleep duration and score trend' clearly signals when this tool is appropriate: when the user wants a chart/trend rather than raw sleep data. It does not explicitly mention alternatives like list_sleep or show_sleep_detail, but the visual-vs-list/detail distinction is strongly implied by the wording.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_week_stepsA
Read-onlyIdempotent
Inspect

Visual step-count trend versus goal for 7d, 30d, 90d, or 1y. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 7d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds useful behavioral context beyond annotations by specifying the interactive chart and short summary return format, and by framing the output as a trend versus goal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences communicate the core behavior, the supported time windows, and the return format without redundancy. The main action is front-loaded, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with one optional parameter, a complete schema, and an output schema, the description provides all necessary selection and invocation context. The interactive chart and summary are specified, and no additional behavioral caveats are essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the sole parameter 'range' is fully documented with an enum and default. The description repeats the allowed range values but adds no new meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: a visual step-count trend compared to a goal. It clearly differentiates this from sibling tools like show_week_sleep or show_week_macros by naming the exact metric and view type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when the user wants a step-count trend versus goal over 7d, 30d, 90d, or 1y. It does not explicitly name alternative tools or exclusion criteria, but the context is unambiguous for a simple visualization tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_week_workoutsA
Read-onlyIdempotent
Inspect

Visual training trend for 7d, 30d, 90d, or 1y. Includes user-marked Rest Day, Sick Day, and Travel Day context even when no workout exists. Returns an interactive chart plus short summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 7d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds genuinely useful behavioral context beyond that: it surfaces user-marked Rest/Sick/Travel Days even when no workout exists, and states the return is an interactive chart plus summary.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with zero filler; the scoping detail (time windows) leads, followed by the notable behavioral nuance and the return shape.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with an output schema and full annotation coverage, the description supplies the important remaining context (rest/sick/travel day inclusion, chart output). The only meaningful gap is the absence of routing guidance versus sibling visualization tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the enum values, default (7d), and meaning of 'range'. The description only repeats the same window values, adding no syntax or semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: produces a visual training trend, with the supported time windows enumerated. It's clearly a chart/aggregate tool rather than a raw-data tool, distinguishing it implicitly from list_workouts and show_workout, but it never names a sibling to make the boundary explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'visual training trend' and the range parameter, but the description offers no when-to-use guidance or exclusions relative to the many sibling show_*/list_* tools. An agent must infer that this is for a charted overview rather than per-workout detail.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_wellbeingA
Read-onlyIdempotent
Inspect

Show energy, mood, stress, and soreness together with recent context and an overall wellbeing trend, with a 7-day/30-day/90-day/1-year range toggle. Prefer this for how-I-have-been-feeling and subjective recovery questions. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeNoTime window. One of 7d, 30d, 90d, 1y. Default 30d.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds behavior the annotations do not: it renders an interactive MCP app and returns a short text summary alongside the visual view, which tells the agent what output mode to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what is shown and the range toggle, then usage guidance. Slightly long in the third sentence ('prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text'), but each sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, yet the description still tells the agent a text summary accompanies the visual. With annotations covering safety and a single well-documented parameter, the definition is nearly complete; only sibling differentiation is thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with the single enum parameter and its default fully documented, so the baseline is 3. The description restates the toggle values but adds no format or semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Show) and resource (energy, mood, stress, soreness plus trend) and lives in the same namespace as list_wellbeing/update_wellbeing while clearly being the visual/aggregate view rather than the raw listing. An agent can tell it apart from list_wellbeing without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the query class it serves ('how-I-have-been-feeling and subjective recovery questions') and instructs preferring the interactive MCP app over text descriptions of rows. It does not, however, contrast itself against the very close siblings list_wellbeing or show_recovery, which would have removed the remaining ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

show_workoutA
Read-onlyIdempotent
Inspect

Show a single logged workout session's exercises and sets as an inline card. Defaults to the most recent workout; can target a specific date. When the user asks about this, prefer calling this tool and rendering the interactive MCP app over describing the underlying rows in text. Returns a short text summary alongside the visual view.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoWorkout date (YYYY-MM-DD). Optional; defaults to the most recent workout.

Output Schema

ParametersJSON Schema
NameRequiredDescription
kindYesWidget discriminant identifying the payload shape.
generatedAtNoISO timestamp the snapshot was built.

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only and idempotent behavior. The description adds that it returns a short text summary alongside the visual view, which is useful behavioral detail beyond the annotations, though it does not cover error cases or side effects (which are already implied safe by annotations).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and information-dense: it states the output, default behavior, targeting capability, and a usage preference, all in two sentences without redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read-only tool with an output schema, the description provides sufficient context: what the user sees, how to target a date, and the relationship to the interactive app. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'date' parameter is fully described in both schema and tool description, including format (YYYY-MM-DD), optionality, and default behavior. Schema coverage is 100%, and the description adds meaningful usage context beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states action ('Show'), resource ('single logged workout session's exercises and sets'), and format ('inline card'), distinguishing it from list tools and get_workout by emphasizing the visual/interactive rendering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to prefer this tool over text descriptions when the user asks about a workout, and contrasts it with alternative presentation methods. The optional date parameter and default behavior are also specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_cycleA
Idempotent
Inspect

Update an existing period record. Use to correct dates, add a missing end date, or clear an end date (resume).

Common uses:

  • "my period ended on the 9th not the 8th" → update ended_on

  • "actually my period started the 2nd not the 3rd" → update started_on

  • "I'm still on my period" → reopen: true to clear ended_on and mark it active

SELECTOR — pass id if known, or date (the period's start date, or any date that falls within it) to resolve it. Exactly one required. If date matches more than one record, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRecord ID. Alternative to date.
dateNoAlternative to id: a date (YYYY-MM-DD) that identifies the period — its start date, or any day within it. Resolves only when exactly one record matches.
reopenNotrue clears ended_on and marks the period active again.
ended_onNoUpdated end date. Format: YYYY-MM-DD.
started_onNoUpdated start date. Format: YYYY-MM-DD. Optional.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, idempotentHint=true), the description discloses real behavioral details: reopen clears ended_on and marks the period active, the date selector resolves any day within the period, exactly one selector is required, and multiple matches produce an error with candidate IDs. These are meaningful operational traits not present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently organized into purpose, common-use examples, and selector rules. Every sentence carries information, the examples are compact, and the critical selector constraint is placed at the end as a distinct section. No filler or redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five parameters, an output schema, and a nontrivial selector ambiguity, the description covers everything needed to call it correctly: what it does, which fields map to which intents, how to resolve the record, and the failure mode for ambiguous dates. The presence of an output schema means return values do not need explanation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100% and each parameter has its own description, the tool description adds significant meaning: id and date are mutually exclusive selectors, date can be a start date or any day within the period, and the examples map natural-language requests to specific parameters (ended_on, started_on, reopen). This goes well beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Update an existing period record,' then enumerates the concrete actions it supports (correct dates, add a missing end date, clear an end date/resume). The word 'existing' implicitly distinguishes it from log_cycle, and 'update' separates it from delete_cycle, so an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use the tool ('Use to correct dates, add a missing end date, or clear an end date') and gives realistic user-intent examples. It does not explicitly name alternatives like log_cycle or delete_cycle or state when not to use it, so it stops one step short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_goal
Destructive
Inspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Change, complete, pause, stop/end, reopen, or delete an existing goal. Call this tool directly for ordinary goal changes. It already loads the user's current goals and resolves a unique goal from goal_id or a natural-language goal_ref, so do NOT call list_goals first just to find an ID.

For edit, pass only the fields to change. For complete/pause/end/reopen/delete, no edit fields are required. If goal_ref genuinely matches multiple goals, this tool returns the candidates and changes nothing. Formal goal end/delete follows the existing Goals UI cancel lifecycle; standard-target end/pause turns that target off while delete removes it.

WEIGHT INPUTS: For a body_weight_target_lb standard target or body_comp metric lean_mass_lb, edit with target_value_lb. This alias is unit-aware; in-app calls use the preferred unit, while external calls use pounds unless input_weight_unit is set. Legacy generic target_value remains canonical pounds.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoConcise title, inferred from the goal inputs.
weeksNoFor weight_loss. Duration in weeks when target_date is not supplied. For body_comp. Duration in weeks when target_date is not supplied.
actionYesWhat to do. 'end' when the user is stopping or canceling a goal; for a formal goal that follows the app's cancel behavior and removes it.
metricNoFor consistency. What consistency behavior to track. Required on create. Set at create, not editable later. For body_comp. Body composition metric. Required on create. Set at create, not editable later. For nutrition. Legacy nutrition metric.
goal_idNoFormal goal ID, when known.
goal_refNoNatural-language reference when no ID is known, e.g. "protein", "10K", "weight loss". This tool resolves it against current goals itself.
new_nameNoFor n1_experiment. Name for a new supplement when not using supplement_id.
new_brandNoFor n1_experiment. Optional brand for a new supplement.
race_dateNoFor race. Race date in YYYY-MM-DD format. Required on create.
start_dateNoYYYY-MM-DD. Default: today.
start_valueNoFor body_comp. Starting value when the goal begins. Required on create. Set at create, not editable later.
target_dateNoFor weight_loss. Target date in YYYY-MM-DD format. Use this when the user names a deadline. For body_comp. Target date in YYYY-MM-DD format. Use this when the user names a deadline.
week_windowNoFor consistency. How the week this goal is measured against is bounded: rolling = the last 7 days, sunday/monday = a calendar week that resets on that day. Default: rolling.
start_1rm_lbNoFor strength. Estimated 1RM when the goal starts. Required on create. Set at create, not editable later. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_hoursNoFor consistency. Nightly sleep target in hours. Required when metric is sleep_duration.
target_valueNoFor body_comp. Target body composition value. Required on create. For nutrition. Legacy nutrition target value. For standard targets, this is the numeric target value.
exercise_nameNoFor strength. Exercise name. Required on create. Set at create, not editable later.
new_dose_unitNoFor n1_experiment. Dose unit for a new supplement.
supplement_idNoFor n1_experiment. Existing supplement ID. Use either supplement_id or the new-supplement fields.
target_1rm_lbNoFor strength. Target 1RM. Required on create. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
start_value_lbNoStarting lean mass in body_comp. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
new_dose_amountNoFor n1_experiment. Dose amount for a new supplement.
start_weight_lbNoFor weight_loss. Starting body weight. Required on create. Set at create, not editable later. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_per_weekNoFor consistency. Target occurrences per week. Required on create.
target_time_secNoFor race. Target finish time in seconds. Required on create.
target_value_lbNoWeight goal target. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
target_weight_lbNoFor weight_loss. Target body weight. Required on create. In lb, or kg with input_weight_unit set. See UNIT INPUTS.
input_weight_unitNoSet to kg when the user gave kg for the _lb fields in this object. Omit when they are already lb.
intervention_daysNoFor n1_experiment. Intervention duration in days.
baseline_directionNoFor n1_experiment. Use already logged previous 14 days or collect the next 14 days.
target_distance_miNoFor race. Target race distance. Required on create. In mi, or km with input_distance_unit set. See UNIT INPUTS.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.
acknowledged_warningsNoWarning keys the user explicitly acknowledged after a guarded create attempt. Omit otherwise.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.
update_injuryA
Idempotent
Inspect

Update an existing injury entry. Use when the user reports an injury is improving, worsening, resolved, or wants to change details. When severity changes, the new value is automatically tracked in the severity history for trend analysis. Only send fields that need to change. Setting end_date automatically marks the injury as Resolved. Use severity_date to backfill historical severity changes (e.g., "it was a 7 in January, dropped to 4 by March").

SELECTOR — pass id if known, or injury (a body part or injury type substring, case-insensitive, e.g. "shoulder") optionally narrowed by date (an injury active on that day). Exactly one of id or injury required. If injury matches more than one entry, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoInjury ID. Alternative to injury.
dateNoOptional, narrows the injury selector to one active on this date (YYYY-MM-DD). Ignored when id is given.
sideNoUpdated side. Optional.
notesNoUpdated notes (replaces existing). Optional.
injuryNoAlternative to id: body part or injury type substring, case-insensitive (e.g. "shoulder"). Optionally narrow with date.
statusNoUpdated status. Optional.
end_dateNoDate injury resolved. Format: YYYY-MM-DD. Auto-sets status to Resolved.
severityNoUpdated severity 1-10. Optional. Change is tracked in severity history.
start_dateNoUpdated start date. Format: YYYY-MM-DD. Optional.
severity_dateNoDate for the severity entry in the history log. Format: YYYY-MM-DD. Default: today. Use to backfill past severity changes.
affected_movementsNoUpdated list of affected movements. Optional.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations by disclosing side effects: severity changes are automatically tracked in history, setting end_date auto-marks Resolved, and severity_date backfills history. Also explains that ambiguous selectors cause an error with candidate IDs. These behavioral details are not available in the minimal idempotent/destructive hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy due to the selector explanation, but it is well-structured into three logical blocks (purpose/usage, side effects, selector rules) and includes a concrete example for substring matching. It is front-loaded with the core purpose, and every sentence adds necessary information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers essential operational details: default value for severity_date, auto-status behavior, selector ambiguity and error handling, and precedence of id over injury. While it does not describe the response shape, an output schema exists (context signal indicates has output schema: true), so that omission is acceptable. The description is sufficiently complete for a caller to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers each parameter (100% coverage), so the baseline is 3. The description adds meaningful semantic relationships not evident from individual field descriptions: end_date ↔ status, severity ↔ severity_date, and the selector precedence (date ignored when id is given). This raised the score from baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action and resource: 'Update an existing injury entry.' The description also explicitly lists use cases ('improving, worsening, resolved, or wants to change details'), which fully distinguishes it from siblings like log_injury, list_injuries, and delete_injury without needing to name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance ('Use when the user reports an injury is improving, worsening, resolved, or wants to change details') and explains parameter selection rules (e.g., 'Only send fields that need to change', selector ambiguity and error behavior). It does not explicitly name alternative tools, but the existing-entry phrasing and use-case list make the boundary clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_lab_resultA
Idempotent
Inspect

Update one or more fields on an existing lab result. Use when the user wants to correct a result already logged, most often its collection date. Only send the fields that need to change; omit all others.

SELECTOR, pass exactly one of id, date, date+marker, or draw_id:

  • id: addresses one marker's row. Any editable field may change.

  • draw_id: addresses every result sharing that draw_id at once, unambiguous by construction. Only new_date, panel_name_new, lab_name, fasting_status, and report_date may change this way — those are draw-level fields.

  • date (optionally narrowed by panel_name): addresses every result from that draw at once, but ONLY when exactly one draw exists on that date — see AMBIGUITY below. Same draw-level fields as draw_id.

  • date + marker (a marker-name substring, case-insensitive, optionally narrowed by panel_name): resolves to one marker's row, same as id. Errors with candidate IDs if more than one marker on that date matches.

AMBIGUITY: a bare date (optionally + panel_name) selector is rejected, with nothing changed, if it would match more than one physical draw — an explicit draw_id from one source plus legacy rows with none, two distinct draw_ids, or two differently-named legacy sources on the same day. The error names every draw found; retry with draw_id, marker, or a narrower panel_name.

FASTING: fasting_status only changes to 'fasting' or 'non_fasting' when the user explicitly states it for that draw — never infer from time of day or a notes mention. 'unknown' is a valid explicit value too, for undoing a mistaken confirmation.

NOTES: changing result_value, result_unit, or result_comparator by id or date+marker automatically regenerates the "Reported as <9 IU/mL"-style note from the new value/unit/comparator, replacing only a note it previously auto-generated — any other free text in notes survives. Pass notes yourself only to add or change unrelated free text; it is folded in alongside the regenerated line rather than overwriting it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoLab result ID. Selects a single marker row. Alternative to date/draw_id.
dateNoCollection date of the draw to update. Format: YYYY-MM-DD. Selects every result from that draw (see AMBIGUITY above), or (with marker) one row. Alternative to id/draw_id.
flagNoNew lab flag: "H", "L", "HH", "LL", or "A". Optional, omit if not changing. Only allowed when selecting by id or date+marker.
notesNoNew notes. Optional, omit if not changing. Only allowed when selecting by id or date+marker. See NOTES above.
markerNoOptional with date: marker-name substring, case-insensitive (e.g. "LDL"), narrowing the date selector to a single marker row so per-marker fields can be edited without an id. Ignored when id or draw_id is given.
draw_idNoOpaque label of your choosing grouping a set of results from one visit, normalized server-side. Selects every result sharing that label, unambiguous by construction. Alternative to id/date. Ignored when id is given.
lab_nameNoNew lab name (e.g. "Quest Diagnostics", "LabCorp"). Optional, omit if not changing. Works with any selector.
new_dateNoNew collection date. Format: YYYY-MM-DD. Optional, omit if not changing. This is the main reason to call this tool, and it works with any selector.
panel_nameNoOptional, narrows a date (or date+marker) selector to one panel within that draw (e.g. "Lipid Panel"). Ignored when id is given.
marker_nameNoNew marker name. Optional, omit if not changing. Only allowed when selecting by id or date+marker.
report_dateNoNew report/result date, separate from the collection date. Format: YYYY-MM-DD. Optional, omit if not changing. Works with any selector.
result_unitNoNew unit of measurement (e.g. "mg/dL"). Optional, omit if not changing. Only allowed when selecting by id or date+marker. See NOTES above.
result_valueNoNew numeric result value. Optional, omit if not changing. Only allowed when selecting by id or date+marker. See NOTES above.
fasting_statusNoNew fasting status for the whole draw. See FASTING above. Optional, omit if not changing. Works with any selector.
panel_name_newNoNew panel name. Optional, omit if not changing. Works with any selector.
result_comparatorNoNew comparator for a reporting-limit result (e.g. "<9"), or "none" to clear it back to an exact value. Optional, omit if not changing. Only allowed when selecting by id or date+marker. See NOTES above.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes far beyond the annotations by disclosing nuanced behaviors: ambiguous date selectors are 'rejected, with nothing changed,' fasting status must be explicitly stated and 'never infer from time of day,' and changing result values 'automatically regenerates' notes while preserving unrelated free text. These details materially shape how an agent invokes the tool and interprets outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every section earns its place given 16 optional parameters and tricky selector semantics. It is front-loaded with the core purpose and use case, then organized with clear headers (SELECTOR, AMBIGUITY, FASTING, NOTES) and bullet lists that make the complex rules scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity mutation tool with 16 parameters, an output schema, and annotations already covering read-only/destructive/idempotent traits, the description is remarkably complete. It covers all selector modes, ambiguity handling, field-level restrictions, fasting inference rules, and note regeneration behavior. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the schema already documents every parameter, but the description adds significant meaning: which fields are draw-level, how marker substrings work case-insensitively, the exact ambiguity failure modes, and the notes regeneration rule. It also explains the idempotent-like 'only send the fields that need to change' contract.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Update one or more fields on an existing lab result,' and immediately clarifies the main use case: 'correct a result already logged, most often its collection date.' This clearly separates it from sibling create/delete/list tools such as log_lab_results, delete_lab_result, and list_lab_results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly says when to use the tool: when the user wants to correct an already-logged lab result. It also gives strong within-tool selector guidance (id, date, date+marker, draw_id) and warns when a selector will be rejected. However, it does not explicitly say 'do not use for new results' or name log_lab_results as the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_mealA
Destructive
Inspect

Update an existing meal. Use only when the current message explicitly changes, corrects, or adds to a meal already logged. Never infer an update from earlier chat history. A plain food statement ("coffee with milk") is a new entry: use log_meal, even if that meal type already exists today.

A meal imported from Cronometer, Fitbit, Apple Health, or Health Connect can't be edited here; the call refuses and names where to edit it instead. Relay that to the user rather than retrying.

FIND THE MEAL: call directly, no preliminary list for an ID or additive totals. Use id if known (it is in list_meals output and in the chat history's saved-records note after a log). Otherwise use date (YYYY-MM-DD, defaults to today) plus name and/or target_meal_type to narrow the existing row. target_meal_type finds the current type and is never written; meal_type sets a new type. If the match is not exactly one row, nothing changes. On multiple matches, ask the user which meal they mean; never select a candidate id yourself. Send only fields that change.

THREE MODES: add_saved: add_recipe_name appends recipe food and ADDS its macros to current totals. Match: case-insensitive exact title, then substring; relay no-match/ambiguous errors. Stored recipe macros, including 0, are authoritative; explicit macros replace the recipe's contribution only for a correction or changed food/quantity. Missing recipe macros: estimate only the missing fields from returned food text and retry with the same selector and add_recipe_name. add_unsaved: add_food_items plus the four core macros (calories/protein_g/fat_g/carbs_g) for ONLY the new food; estimate those core values, never ask for them. Text appends; core macros ADD to current totals. alcohol_g: optional, omitted adds nothing. add_both: set add_recipe_name and add_food_items; ordinary core macro fields describe ONLY the unsaved addition. Recipe contributes its stored core macros. correct: both add_* unset; supplied fields REPLACE stored values. Send corrected totals, omit unchanged fields. Changed food_items: the four core macros for the WHOLE corrected meal; list_meals only if needed stored values are unavailable this turn. Additions and explicit value corrections need no pre-read. food_items: when supplied, replaces the description even in add modes. SATURATED FAT / FIBER: expanded nutrients, never required for an update. The estimator or a saved recipe's nutrient panel can fill omitted values, and core-macro updates work without them. 0 is a real value only when known, never a placeholder. ambiguous_add_vs_correct: ask before updating.

ESTIMATE: estimate (same format as log_meal's ESTIMATE section) supplies the nutrient estimate for whatever food this call changes -- the add_unsaved/add_both addition's own food, or, for a mode-3 food_items replacement, the WHOLE corrected meal. Omit to let the server estimate instead; never required.

MOVE TO A DIFFERENT DATE -> move_to_date. "move Tuesday's lunch to Wednesday": date/name/target_meal_type only SELECT which meal to update; they never move it. Set move_to_date to actually change the stored date, keeping the same id.

REMOVE ONE ADDED COMPONENT -> remove_item_name. Only works for a food previously recorded on THIS meal (add_recipe_name, add_food_items, or log_meal's own estimated items). If the meal has no such recorded item, this throws telling you to use mode 3 with corrected totals instead. An estimated item (log_meal's own food, no add_recipe_name/add_food_items involved) carries a name and gram estimate but no recorded calorie/macro split for that one item, so a precise subtraction isn't possible: this throws asking you to also send calories/protein_g/fat_g/carbs_g (the meal's corrected TOTALS after removing it) in the SAME call, which performs the removal and applies those totals together, in that order -- do this rather than a separate correcting call. Removing an item re-estimates the meal's nutrient panel from what remains (its updated food text and items), rather than leaving a stale estimate for food that's gone -- pass estimate here too to supply that re-estimate yourself instead of letting the server do it.

DRINKS / HYDRATION WHEN UPDATING: a food update must not silently erase an already-linked hydration event. For changes unrelated to drinks, omit fluids and the existing hydration is preserved. When ADDING a drink with add_food_items/add_recipe_name, fluids contains only the newly added drink(s) and they are appended to the meal's hydration. When CORRECTING the meal's drinks with both add_* fields omitted, fluids is the complete corrected drink list and replaces the linked hydration only after replacement rows have been safely inserted. If the correction removes every drink, explicitly send fluids: [] and the linked hydration is deleted. Whenever a drink volume is known or reasonably inferable, include it. Drinks only; exclude food moisture.

CAFFEINE / DRINKS WHEN UPDATING: keep the two sidecars explicit in the same update_meal call. fluids is the hydration side and caffeine is the caffeine side; when a drink change affects both, send both. Omit either one when that side is unchanged so existing linked data is preserved. When ADDING caffeinated food/drink with add_food_items/add_recipe_name, caffeine contains only the new dose(s) and they are appended. When CORRECTING the meal's caffeine with both add_* fields omitted, caffeine is the complete corrected dose list and replaces linked caffeine only after replacement rows have been safely inserted. If the correction removes every caffeine dose, explicitly send caffeine: []. Each retained/new dose needs caffeine_mg and may optionally include source_type; its date follows the meal date.

SAVE AS RECIPE: when the user asks to save an already-logged meal as a reusable recipe, set save_as_recipe=true. This may be the only requested action: identify the real meal with id or the normal selectors and do not invent a food or macro edit. A follow-up like "save that as a recipe" after a successful log should use the known meal id plus save_as_recipe=true. Do not tell the user recipes cannot be saved from chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoMeal ID, if already known. Alternative to date + name/target_meal_type, see FIND THE MEAL above.
dateNoDate the meal was logged. Format: YYYY-MM-DD. Used with name and/or target_meal_type to find the meal when id is omitted; defaults to today if id and date are both omitted. This only SELECTS which meal to update -- see move_to_date below to actually change a meal's stored date.
nameNoSubstring of the food description (case-insensitive) to disambiguate multiple meals on the same date. Only used when id is omitted.
fat_gNoFat (g). See THREE MODES.
fluidsNoOptional drinks consumed in this intake, including liquid ingredients in shakes or smoothies. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking.
carbs_gNoCarbohydrates (g). See THREE MODES.
fiber_gNoOptional expanded nutrient: dietary fiber (g).
caffeineNoOptional caffeine doses in this intake. Keep each dose simple: caffeine amount is required and type is optional. The dose date follows the intake/meal date. If the user gave exact milligrams, preserve them exactly; otherwise estimate from the described food or drink.
caloriesNoCalories (kcal) never kJ
estimateNoThe compact 3-line nutrient estimate for the food this call adds or the WHOLE corrected meal on a food_items replacement. See ESTIMATE above for the format. Omit to have the server estimate instead -- never blocks the write either way. A block whose parts exceed their whole (saturated fat over fat, fiber over carbs) is discarded and re-estimated. NUTRIENT ID DICTIONARY (id=name, unit is the name's suffix; ug=mcg). Every id means exactly this nutrient, never guess the order: 1=fiber_g 2=sugar_g 3=saturated_fat_g 4=monounsaturated_fat_g 5=polyunsaturated_fat_g 6=trans_fat_g 7=cholesterol_mg 8=sodium_mg 9=potassium_mg 10=calcium_mg 11=iron_mg 12=magnesium_mg 13=phosphorus_mg 14=zinc_mg 15=copper_mg 16=manganese_mg 17=selenium_ug 18=chloride_mg 19=chromium_ug 20=iodine_ug 21=molybdenum_ug 22=vitamin_a_ug 23=vitamin_c_mg 24=vitamin_d_ug 25=vitamin_e_mg 26=vitamin_k_ug 27=thiamin_b1_mg 28=riboflavin_b2_mg 29=niacin_b3_mg 30=pantothenic_acid_b5_mg 31=vitamin_b6_mg 32=biotin_b7_ug 33=folate_b9_ug 34=folic_acid_ug 35=vitamin_b12_ug 36=choline_mg 37=omega3_g 38=omega6_g 39=caffeine_mg 40=water_g 41=starch_g 42=added_sugar_g 43=total_unsaturated_fat_g 44=fluoride_mg
alcohol_gNoAlcohol (g). See THREE MODES.
meal_typeNoUpdated meal type to WRITE onto the meal (e.g. reclassify a Snack as Dinner), only when the user asks to change the type. A "For my breakfast:"-style label at the start of the message is the pill the user had open, not a request to change it. Optional, omit if not changing. Never inferred from add_recipe_name. Distinct from target_meal_type above, which FINDS a meal by its current type and is never written.
protein_gNoProtein (g). See THREE MODES.
food_itemsNoReplacement food description; omit to preserve or append. See THREE MODES.
move_to_dateNoMove this meal to a different date. Format: YYYY-MM-DD. Distinct from date above, which only finds the meal; this is what actually changes it, keeping the same id. Optional, omit if not moving the meal.
add_food_itemsNoUnsaved food description to add; omit unless adding unsaved food. See THREE MODES.
recipe_fiber_gNoOptional recipe-component fiber (g) for add_both.
save_as_recipeNoTrue only when the user explicitly asks to save this meal as a reusable recipe. The recipe is copied from the final persisted meal. On update_meal, this can be the only requested action; use the real meal id or normal selectors and do not invent an edit.
add_recipe_nameNoSaved recipe name or close phrase to add; omit unless adding saved food. See THREE MODES.
saturated_fat_gNoOptional expanded nutrient: saturated fat (g).
remove_item_nameNoName or substring of a previously-recorded item to remove from this meal (see REMOVE ONE ADDED COMPONENT above). Normally exclusive of every other field below -- only id/date/name/target_meal_type may accompany it, to select the meal, plus optionally estimate for the remaining meal -- EXCEPT when the tool has already refused this exact removal asking for corrected totals: then also send calories/protein_g/fat_g/carbs_g together with remove_item_name in the retry.
target_meal_typeNoWhich meal type to FIND on the date, e.g. Breakfast, to disambiguate multiple meals logged that day -- "update today's breakfast" is target_meal_type: "Breakfast". Case-insensitive, only used when id is omitted. This is NEVER written to the meal; it only narrows the search, exactly like name above. Distinct from meal_type below, which SETS the new type to write. If no meal of this type is logged on the date, the call throws naming the meal types that ARE logged that day and changes nothing -- it never falls back to whichever meal the date happens to match.
recipe_saturated_fat_gNoOptional recipe-component saturated fat (g) for add_both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/not-idempotent/not-readonly, and the description goes well beyond them: refusal on imported meals with instructions to relay rather than retry, no-op on ambiguous matches with ask-the-user policy, replace-vs-append semantics for each mode, hydration/caffeine preservation rules, and ordered insert-then-replace behavior. Side effects and failure paths are unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the triggering condition and route-away rules, and the ALL-CAPS section headers (FIND THE MEAL, THREE MODES, ESTIMATE) make it scannable despite its length. It is verbose and repeats the append/replace rule for fluids and caffeine, but for a 23-parameter tool with four modes and two sidecars most sentences carry real load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Selection (id vs date+name/target_meal_type), all three add modes, correction, removal, date move, recipe save, sidecar handling, and both error paths are covered; an output schema exists, so return values need no explanation. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description supplies the semantics the schema defers to it for: the add vs correct modes (adds ADD to totals, corrections REPLACE), meal_type (writes) vs target_meal_type (finds only), fluids/caffeine append-when-adding vs full-list-replace-when-correcting, and the remove_item_name retry exception. This is meaning well beyond the field types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update an existing meal') and immediately carves out its boundaries against siblings: plain food statements go to log_meal, date changes go to move_to_date, component removal goes to remove_item_name. An agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('only when the current message explicitly changes, corrects, or adds'), when-not ('never infer an update from earlier chat history'), and named alternatives for adjacent operations (log_meal, move_to_date, remove_item_name, and mode 3 corrected totals instead of remove when no recorded item exists). This is about as complete as routing guidance gets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_recovery_sessionA
Idempotent
Inspect

Update one or more fields on an existing recovery session log entry. Use when the user wants to correct or change something already logged (e.g. wrong duration, quality rating, category, or notes). Only send the fields that need to change; omit all others.

SELECTOR — pass id if known, or session_date (+ optional session_category to narrow) to resolve it. Exactly one of id or session_date required. If it matches more than one session, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoRecovery session ID. Alternative to session_date.
dateNoUpdated date (YYYY-MM-DD). Optional.
notesNoUpdated notes. Optional.
qualityNoUpdated quality 1-5. Optional.
skippedNoUpdated skipped status. Optional.
categoryNoUpdated category. Optional.
strategy_idNoUpdated strategy ID. Optional — set null to unlink.
session_dateNoAlternative to id: the date (YYYY-MM-DD) the session was logged on. Optionally narrow with session_category.
strategy_nameNoUpdated practice name. Optional.
duration_minutesNoUpdated duration in minutes. Optional.
session_categoryNoOptional, narrows session_date to one category when more than one session shares that date. Ignored when id is given.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly=false, idempotent=true, destructive=false), the description reveals selector resolution behavior: exactly one of id or session_date is needed, session_category can narrow the match, and ambiguity errors with candidate IDs for retry. This is valuable behavioral context an agent would otherwise infer only from failures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly scoped paragraphs: the first gives the use case and patch style, the second the selector contract. It is front-loaded with the actionable verb and includes an example list without bloat. Every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter update tool with an output schema available, the description covers the purpose, patch semantics, selector requirements, disambiguation, and error behavior. The only small omission is an explicit 'do not pass both id and session_date' rule, but the 'or' selector language makes that sufficiently clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already documents all 11 parameters (100% coverage), the description adds critical relational semantics: the exact-one-of id/session_date selector rule, optional narrowing by session_category, that session_category is ignored when id is supplied, and the patch-style 'omit unchanged fields' convention. These relationships are absent from the individual parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and object: 'Update one or more fields on an existing recovery session log entry.' It distinguishes from related siblings like log_recovery_session and delete_recovery_session by emphasizing 'existing' and 'already logged,' and lists concrete fields (duration, quality, category, notes) that make the scope unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool: when the user wants to correct or change something already logged, and it gives updating guidance ('Only send the fields that need to change'). It does not name alternatives such as log_recovery_session for new entries, but the 'existing'/'already logged' framing and examples supply clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_wellbeingA
Idempotent
Inspect

Update an existing wellbeing entry. Only updates fields that are provided — omitted fields remain unchanged.

SELECTOR — pass id if known, or date (any day within the entry's period) to resolve it. Exactly one of id or date required. If date matches more than one entry, the call errors with candidate IDs to retry with.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoWellbeing entry ID. Alternative to date.
dateNoAlternative to id: a date (YYYY-MM-DD) that falls within the entry's period. Resolves only when exactly one entry matches.
moodNoUpdated mood 1-10. Optional.
notesNoUpdated notes (replaces existing). Optional.
energyNoUpdated energy 1-10. Optional.
stressNoUpdated stress 1-10. Optional.
sorenessNoUpdated soreness 1-10. Optional.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly false, idempotent, non-destructive), the description discloses that only provided fields are updated and omitted fields remain unchanged, and that ambiguous date matches cause an error with candidate IDs. This gives the agent accurate expectations about partial updates and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short paragraphs with no filler: the operation is stated first, then the selector rule. Every sentence carries necessary information for correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all non-obvious invocation logic (selector requirement, ambiguity handling, partial update behavior). With the output schema present and annotations covering safety/idempotence, nothing essential is missing for an agent to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already documents all 7 parameters at 100% coverage, the description adds the crucial constraint that exactly one of id/date is required and explains the resolution error behavior. It also clarifies partial update semantics that apply to the optional fields, which the schema alone does not state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update') and resource ('existing wellbeing entry'), making the operation unambiguous. The word 'existing' implies modification rather than creation, separating it from log_wellbeing, though it does not explicitly name sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit invocation conditions: exactly one of id or date is required, with a fallback rule for dates that match multiple entries. It does not explicitly contrast against create/delete alternatives, but the context for when to use this update tool is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_workoutA
Destructive
Inspect

UNIT INPUTS: value: pass the user's number unconverted; tool converts once before storage. alternate_unit: set the field's matching input_* companion; canonical_unit: omit companion. precedence: overrides instructions to convert manually.

Update a workout session: correct metadata, fix set values, rename/add/remove exercises or individual sets, or move exercises between supersets. Use for any post-log correction.

FIND THE SESSION: call directly, no preliminary list for an ID. session_id: positive ID from a workout tool, never invented. Otherwise pass session_date (YYYY-MM-DD, defaults to today) and, only if more than one session was logged that day, name (a substring of the workout's focus/type, case-insensitive) to narrow it down. A match that isn't exactly one session returns an error explaining why, with nothing changed — retry with session_id or a narrower name, never guess. get_workout still gives full detail (exercise names, slot names SS1/SS2/WarmUp/Finisher) when needed; list_exercises only for unknown canonical names when adding or renaming. Call with only the fields that change — operations can combine in one call.

OPERATIONS:

  • Metadata: date, focus_type, location, notes, rpe, heart_points_moderate/peak.

  • set_updates: patch reps/load/bodyweight/notes/equipment on a set, addressed by set_id OR by exercise_name + set_position (1-based, matches get_workout's "Set N"). Use clear_weight=true when a stored load is wrong but the real external load is unknown; use is_bodyweight=true when the corrected set was genuinely bodyweight.

  • remove_sets: delete sets, same set_id-or-exercise_name+set_position addressing; remaining sets renumber; an emptied exercise/slot is removed automatically.

  • rename_exercises: renames every set of an exercise in place (preserves set IDs, RPE, notes; rebuilds NSI), never remove + add.

  • remove_exercises: deletes all sets for named exercises; empty slots removed automatically.

  • add_exercises: new exercises with sets; to_superset_slot joins an existing slot, omit for standalone.

  • move_exercises: reassigns an exercise to a different slot; "new" makes it standalone. Use from_superset_slot from get_workout when the same exercise name appears in multiple slots.

SUPERSET SLOTS: rename_exercises/remove_exercises use superset_slot (or { name, superset_slot } for remove_exercises) from get_workout when a name appears in multiple slots.

LITERAL NAME: literal_name: true keeps the user's exact wording instead of the closest library match, skips the resolver, and gets no NSI score (no benchmark to compare an unmatched name against). Use for "call it exactly X", "not the standard one", "literally X", or a rejected match. Applies below.

EQUIPMENT (load basis): dumbbell_pair is one dumbbell in EACH hand, weight_lb PER HAND (2x for NSI); dumbbell_single is one implement total. Laterality (single-leg/arm) does NOT decide this alone. Set it when the user describes the load (each hand, machine, band); a wrong or missing tag silently halves or doubles NSI. Values: barbell, dumbbell_pair, dumbbell_single, machine, kettlebell, bodyweight, band, cable, trx, other.

A set_id or exercise_name+set_position matching more than one set (the same exercise in two superset slots) is ambiguous and errors rather than guessing — use the exact set_id from get_workout to disambiguate.

The result discloses a mismatched name from rename_exercises/add_exercises; relay it in your own words. If a name matches nothing closely enough, the result names near-miss library exercises; ask the user which they meant rather than accept the unscored custom log silently.

INFER — do not ask: session_date defaults to today, set positions count from 1 per exercise. Slot names and set_ids beyond what's inferable come from get_workout; canonical exercise names come from list_exercises.

SAVED WORKOUTS: pass saved_workout_id to edit a reusable Saved Workout instead of completed workout history. Use saved_workout_title, saved_exercise_updates, and/or add_exercises. add_exercises keeps its normal payload shape; to_superset_slot accepts the Saved Workout slot label returned by get_workout or its SS1-style alias. For progression requests, inspect real exercise history first rather than applying a deterministic formula.

cardio_totals: distance_mi/distance_km/distance_meters, duration_sec, calories are update fields; never put corrected totals only in notes. existing_exercise: use add_sets to add sets; add_exercises is for new exercises. ambiguous_slot: rename/remove/move requires the source slot when the exercise appears in multiple slots.

ParametersJSON Schema
NameRequiredDescriptionDefault
rpeNoSession RPE, 1-10, half steps allowed: 5 moderate, 7 hard, 9 one rep left, 10 failure. Infer from comments about overall difficulty, or omit.
dateNoNew session date, YYYY-MM-DD.
nameNoSubstring of the focus/type, e.g. "Push", case-insensitive, to pick between sessions on session_date. Only without session_id.
notesNoNew session notes.
add_setsNoAdd one or more sets to an exercise already present in this workout. Use add_exercises only for a brand-new exercise.
caloriesNoCorrected session calories when known.
locationNoNew location, e.g. Gym, Home.
focus_typeNoNew category, e.g. Push, Pull, Legs.
session_idNoPositive session ID returned by a workout tool. If unknown, omit and use session_date + name; never guess.
distance_kmNoCorrected session distance in kilometers. Server converts it; do not convert it yourself.
distance_miNoCorrected session distance in miles. Use only when the user gave miles. In mi, or km with input_distance_unit set. See UNIT INPUTS.
remove_setsNoSets to delete. See OPERATIONS above.
set_updatesNoIndividual set corrections.
duration_secNoCorrected total session duration in seconds.
session_dateNoYYYY-MM-DD, defaults to today. Finds the session when session_id is omitted. Distinct from `date`, which CHANGES the stored date.
add_exercisesNoNew exercises to add. Reuse known canonical names; list_exercises if unknown.
move_exercisesNoMove exercises between superset slots. When the same exercise name appears in multiple slots, from_superset_slot is required.
distance_metersNoCorrected session distance in meters. Server converts it; do not convert it yourself.
remove_exercisesNoExercises to remove; matched sets deleted, empty slots removed automatically. See SUPERSET SLOTS above.
rename_exercisesNoRename logged exercises in place.
saved_workout_idNoSaved Workout ID to update instead of a completed workout session.
heart_points_peakNoNew peak heart points.
input_distance_unitNoSet to km when the user gave km for the _mi fields in this object. Omit when they are already mi.
saved_workout_titleNoOptional new title for the Saved Workout.
heart_points_moderateNoNew moderate heart points.
saved_exercise_updatesNoOptional prescription edits matched first by exact Saved Workout exercise name, then by a unique canonical library match.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYesHuman-readable result text returned by the tool.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this destructive, and the description adds substantial behavioral detail: set renumbering after removal, automatic removal of emptied slots, preservation of set IDs/RPE/notes on rename, no NSI scoring for literal names, server-side unit conversion, and errors with "nothing changed" on ambiguous matches. None of this contradicts the annotations, and it meaningfully clarifies post-conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is densely structured into labeled sections and every sentence carries operational value, which is appropriate for a 26-parameter destructive tool. However, it opens with UNIT INPUTS before stating the tool's purpose, so it is not optimally front-loaded, and some points repeat schema descriptions. The length is largely justified by complexity, but it could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers session selection, all five operation groups, superset slot handling, literal naming, equipment semantics, saved workouts, cardio totals, and explicit inference rules for defaults. It also tells the agent when to consult get_workout and list_exercises. Since an output schema exists, the absence of return-value detail is acceptable; nothing needed to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds crucial operational semantics beyond the schema: unit conversion rules, the clear_weight versus is_bodyweight decision, set_id versus exercise_name+set_position addressing, superset_slot disambiguation, and the laterality caveat for equipment. These details materially change how parameters should be populated and prevent silent errors like halved NSI scores.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource — "Update a workout session" — and enumerates the exact operation families: correct metadata, fix set values, rename/add/remove exercises or sets, and move exercises between supersets. This clearly distinguishes it from siblings like log_workout, get_workout, and delete_workout. Even though UNIT INPUTS appears first, the core purpose is explicit and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives thorough routing guidance: call directly with session_id or session_date+name, use get_workout only for detail, and use list_exercises only for unknown canonical names. It also draws explicit lines between add_sets and add_exercises, completed workouts versus saved_workout_id, and when to ask the user versus when to infer. Exclusions such as "never guess" session ids and ambiguity errors are clearly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedlog_workout21 fields changed
      • addedInput schema / properties / avg_cadence_rpm
        Added value: +{
        +  "description": "Reported cycling/rowing cadence, revs/min.",
        +  "maximum": 300,
        +  "minimum": 1,
        +  "type": "number"
        +}
      • addedInput schema / properties / avg_cadence_spm
        Added value: +{
        +  "description": "Reported running/walking cadence, steps/min.",
        +  "maximum": 400,
        +  "minimum": 1,
        +  "type": "number"
        +}
      • addedInput schema / properties / avg_hr
        Added value: +{
        +  "description": "Reported average heart rate (bpm).",
        +  "maximum": 260,
        +  "minimum": 20,
        +  "type": "integer"
        +}
      • addedInput schema / properties / avg_pace_sec_per_mi
        Added value: +{
        +  "description": "Reported average seconds per mile. Never invent a pace. In mi, or km with input_distance_unit set. See UNIT INPUTS.",
        +  "maximum": 7200,
        +  "minimum": 1,
        +  "type": "number"
        +}
      • addedInput schema / properties / avg_power_w
        Added value: +{
        +  "description": "Reported average power in watts (e.g. cycling).",
        +  "maximum": 5000,
        +  "minimum": 0,
        +  "type": "number"
        +}
      • changedInput schema / properties / calories / description
        Previous value: -"Session calories for simple cardio when supplied by the user/device."New value: +"Reported session calories; never invent."
      • changedInput schema / properties / distance_km / description
        Previous value: -"Session distance in kilometers. Server converts it to storage units; do not convert it yourself."New value: +"Session distance in kilometers. Server converts to miles."
      • changedInput schema / properties / distance_meters / description
        Previous value: -"Session distance in meters. Server converts it to storage units; do not convert it yourself."New value: +"Session distance in meters. Server converts to miles."
      • changedInput schema / properties / distance_mi / description
        Previous value: -"Session distance in miles for simple cardio. Use only when the user supplied miles. In mi, or km with input_distance_unit set. See UNIT INPUTS."New value: +"Session distance in miles, when supplied in miles. In mi, or km with input_distance_unit set. See UNIT INPUTS."
      • changedInput schema / properties / duration_sec / description
        Previous value: -"Total session duration in seconds for simple cardio when known."New value: +"Total workout or moving time in seconds. Convert user-stated minutes to seconds; not a set duration."
      • addedInput schema / properties / elapsed_sec
        Added value: +{
        +  "description": "Elapsed seconds including pauses; omit if unknown.",
        +  "maximum": 86400,
        +  "minimum": 1,
        +  "type": "integer"
        +}
      • addedInput schema / properties / elevation_gain_m
        Added value: +{
        +  "description": "Reported elevation gain in meters.",
        +  "maximum": 100000,
        +  "minimum": 0,
        +  "type": "number"
        +}
      • addedInput schema / properties / estimated
        Added value: +{
        +  "description": "True only when a saved numeric value is an AI approximation; describe its source/range in notes.",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / exercises / description
        Previous value: -"Every exercise performed, in order."New value: +"Only the actual exercises/sets provided by the user. For a session-level log, omit or use []."
      • changedInput schema / properties / focus_type / description
        Previous value: -"Infer from exercises: Push, Pull, Legs, Upper, Lower, Full Body, Cardio, Mobility. Omit if unclear."New value: +"Specific completed activity (Walking, Cycling, Yoga, Kettlebell Circuit, Full Body, etc.). Infer from user wording; omit only if notes describe the activity."
      • addedInput schema / properties / intensity
        Added value: +{
        +  "description": "Qualitative effort as reported (does not manufacture an RPE).",
        +  "enum": [
        +    "light",
        +    "moderate",
        +    "vigorous"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / max_hr
        Added value: +{
        +  "description": "Reported maximum heart rate (bpm).",
        +  "maximum": 260,
        +  "minimum": 20,
        +  "type": "integer"
        +}
      • addedInput schema / properties / max_power_w
        Added value: +{
        +  "description": "Reported peak power in watts.",
        +  "maximum": 5000,
        +  "minimum": 0,
        +  "type": "number"
        +}
      • addedInput schema / properties / segments
        Added value: +{
        +  "description": "Actually performed intervals/blocks, same structure as the in-app run logger; never invent.",
        +  "items": {
        +    "properties": {
        +      "distance_mi": {
        +        "description": "In mi, or km with input_distance_unit set. See UNIT INPUTS.",
        +        "type": "number"
        +      },
        +      "duration_sec": {
        +        "type": "integer"
        +      },
        +      "input_distance_unit": {
        +        "description": "Set to km when the user gave km for the _mi fields in this object. Omit when they are already mi.",
        +        "enum": [
        +          "mi",
        +          "km"
        +        ],
        +        "type": "string"
        +      },
        +      "kind": {
        +        "enum": [
        +          "warmup",
        +          "easy",
        +          "steady",
        +          "tempo",
        +          "threshold",
        +          "interval",
        +          "recovery",
        +          "hill",
        +          "cooldown",
        +          "other"
        +        ],
        +        "type": "string"
        +      },
        +      "name": {
        +        "type": "string"
        +      },
        +      "notes": {
        +        "type": "string"
        +      },
        +      "repeats": {
        +        "type": "integer"
        +      },
        +      "target_hr_max": {
        +        "type": "integer"
        +      },
        +      "target_hr_min": {
        +        "type": "integer"
        +      },
        +      "target_pace_sec_per_mi": {
        +        "description": "In mi, or km with input_distance_unit set. See UNIT INPUTS.",
        +        "type": "number"
        +      }
        +    },
        +    "required": [
        +      "name",
        +      "kind"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 50,
        +  "type": "array"
        +}
      • addedInput schema / properties / splits
        Added value: +{
        +  "description": "Actually reported laps/splits, same structure as the in-app run logger; never invent.",
        +  "items": {
        +    "properties": {
        +      "avg_cadence_spm": {
        +        "type": "number"
        +      },
        +      "avg_hr": {
        +        "type": "integer"
        +      },
        +      "avg_power_w": {
        +        "type": "number"
        +      },
        +      "distance_mi": {
        +        "description": "In mi, or km with input_distance_unit set. See UNIT INPUTS.",
        +        "type": "number"
        +      },
        +      "duration_sec": {
        +        "type": "integer"
        +      },
        +      "elevation_change_m": {
        +        "type": "number"
        +      },
        +      "index": {
        +        "type": "integer"
        +      },
        +      "input_distance_unit": {
        +        "description": "Set to km when the user gave km for the _mi fields in this object. Omit when they are already mi.",
        +        "enum": [
        +          "mi",
        +          "km"
        +        ],
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "index"
        +    ],
        +    "type": "object"
        +  },
        +  "maxItems": 100,
        +  "type": "array"
        +}
      • addedInput schema / properties / start_time
        Added value: +{
        +  "description": "Local HH:MM when explicitly supplied; otherwise omit.",
        +  "type": "string"
        +}
  2. 3 tool updates
    • Changedcreate_goal2 fields changed
      • addedInput schema / properties / start_value_lb
        Added value: +{
        +  "description": "Starting lean mass in body_comp. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.",
        +  "type": "number"
        +}
      • addedInput schema / properties / target_value_lb
        Added value: +{
        +  "description": "Weight goal target. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.",
        +  "type": "number"
        +}
    • Changedlog_body_metrics1 field changed
      • changedInput schema / properties / weight_lb / description
        Previous value: -"Body weight in pounds, between 50 and 700. Convert from kilograms if the user spoke in kg (kg x 2.2046). In lb, or kg with input_weight_unit set. See UNIT INPUTS."New value: +"Body weight, between 50 and 700. In lb, or kg with input_weight_unit set. See UNIT INPUTS."
    • Changedupdate_goal2 fields changed
      • addedInput schema / properties / start_value_lb
        Added value: +{
        +  "description": "Starting lean mass in body_comp. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.",
        +  "type": "number"
        +}
      • addedInput schema / properties / target_value_lb
        Added value: +{
        +  "description": "Weight goal target. See WEIGHT INPUTS above. In lb, or kg with input_weight_unit set. See UNIT INPUTS.",
        +  "type": "number"
        +}
  3. 1 tool update
    • Changedlog_meal3 fields changed
      • addedInput schema / properties / added_food_items
        Added value: +{
        +  "description": "Only food added to a saved recipe; supply final combined macros. Omit when none.",
        +  "type": "string"
        +}
      • addedInput schema / properties / recipe_multiplier
        Added value: +{
        +  "description": "Portions of the matched saved recipe; default 1, e.g. 0.5 for half.",
        +  "type": "number"
        +}
      • changedInput schema / properties / recipe_name / description
        Previous value: -"Name or close phrase for one of the user's saved recipes, e.g. \"protein oats\", \"chicken bowl\", or \"breakfast\" for their saved Breakfast recipe. See SAVED RECIPES above."New value: +"Saved recipe title only, without portion words or additions. See SAVED RECIPES above."
  4. 2 tool updates
    • Changedlog_meal1 field changed
      • changedInput schema / properties / fluids / description
        Previous value: -"Optional drinks consumed in this intake. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking."New value: +"Optional drinks consumed in this intake, including liquid ingredients in shakes or smoothies. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking."
    • Changedupdate_meal1 field changed
      • changedInput schema / properties / fluids / description
        Previous value: -"Optional drinks consumed in this intake. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking."New value: +"Optional drinks consumed in this intake, including liquid ingredients in shakes or smoothies. Omit when no drink amount is known. Hydration is persisted only when the user enabled hydration tracking."
  5. 1 tool update
    • Changedlog_meal4 fields changed
      • addedInput schema / properties / recipe_only
        Added value: +{
        +  "description": "true=save recipe only and write no food_log row. Requires save_as_recipe=true. See RECIPE SAVE above.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / recipe_servings
        Added value: +{
        +  "description": "Whole-batch serving count when recipe_only=true; default=1. See RECIPE SAVE above.",
        +  "type": "number"
        +}
      • addedInput schema / properties / recipe_title
        Added value: +{
        +  "description": "Recipe name when recipe_only=true. See RECIPE SAVE above.",
        +  "type": "string"
        +}
      • changedInput schema / properties / save_as_recipe / description
        Previous value: -"True only when the user explicitly asks to save this meal as a reusable recipe. The recipe is copied from the final persisted meal. On update_meal, this can be the only requested action; use the real meal id or normal selectors and do not invent an edit."New value: +"True only when the user explicitly asks to save a reusable recipe. recipe_only=true saves recipe only; otherwise the final persisted meal is saved after the log succeeds. See RECIPE SAVE above."
  6. 2 tool updates
    • Addedlist_locations
    • Addedmanage_location
  7. 1 tool update
    • Changedlist_lab_results2 fields changed
      • changedInput schema / properties / marker_name / description
        Previous value: -"Optional — filter to a specific marker (e.g. \"LDL Cholesterol\")."New value: +"Optional — filter to a partial marker name (e.g. \"LDL Cholesterol\")."
      • changedInput schema / properties / panel_name / description
        Previous value: -"Optional — filter to a specific panel (e.g. \"Lipid Panel\")."New value: +"Optional — filter to a partial panel name (e.g. \"Lipid Panel\")."
  8. 8 tool updates
    • Addedget_nutrient_contributors
    • Addedget_nutrient_history
    • Addedget_nutrient_summary
    • Changedlist_lab_results1 field changed
      • changedInput schema / properties / start_date / description
        Previous value: -"Start of date range. Format: YYYY-MM-DD. Default: 365 days ago."New value: +"Start of date range. Format: YYYY-MM-DD. Default: all time."
    • Changedlog_meal3 fields changed
      • addedInput schema / properties / estimate
        Added value: +{
        +  "description": "The compact 3-line nutrient estimate for this meal. See ESTIMATE above for the format. Omit to have the server estimate instead -- never blocks the write either way. A block whose parts exceed their whole (saturated fat over fat, fiber over carbs) is discarded and re-estimated.\n\nNUTRIENT ID DICTIONARY (id=name, unit is the name's suffix; ug=mcg). Every id means exactly this nutrient, never guess the order:\n1=fiber_g\n2=sugar_g\n3=saturated_fat_g\n4=monounsaturated_fat_g\n5=polyunsaturated_fat_g\n6=trans_fat_g\n7=cholesterol_mg\n8=sodium_mg\n9=potassium_mg\n10=calcium_mg\n11=iron_mg\n12=magnesium_mg\n13=phosphorus_mg\n14=zinc_mg\n15=copper_mg\n16=manganese_mg\n17=selenium_ug\n18=chloride_mg\n19=chromium_ug\n20=iodine_ug\n21=molybdenum_ug\n22=vitamin_a_ug\n23=vitamin_c_mg\n24=vitamin_d_ug\n25=vitamin_e_mg\n26=vitamin_k_ug\n27=thiamin_b1_mg\n28=riboflavin_b2_mg\n29=niacin_b3_mg\n30=pantothenic_acid_b5_mg\n31=vitamin_b6_mg\n32=biotin_b7_ug\n33=folate_b9_ug\n34=folic_acid_ug\n35=vitamin_b12_ug\n36=choline_mg\n37=omega3_g\n38=omega6_g\n39=caffeine_mg\n40=water_g\n41=starch_g\n42=added_sugar_g\n43=total_unsaturated_fat_g\n44=fluoride_mg",
        +  "type": "string"
        +}
      • addedInput schema / properties / fiber_g
        Added value: +{
        +  "description": "Dietary fiber in grams. See MACROS above.",
        +  "type": "number"
        +}
      • addedInput schema / properties / saturated_fat_g
        Added value: +{
        +  "description": "Saturated fat in grams. See MACROS above.",
        +  "type": "number"
        +}
    • Addedset_nutrient_target
    • Addedset_supplement_nutrients
    • Changedupdate_meal5 fields changed
      • addedInput schema / properties / estimate
        Added value: +{
        +  "description": "The compact 3-line nutrient estimate for the food this call adds or the WHOLE corrected meal on a food_items replacement. See ESTIMATE above for the format. Omit to have the server estimate instead -- never blocks the write either way. A block whose parts exceed their whole (saturated fat over fat, fiber over carbs) is discarded and re-estimated.\n\nNUTRIENT ID DICTIONARY (id=name, unit is the name's suffix; ug=mcg). Every id means exactly this nutrient, never guess the order:\n1=fiber_g\n2=sugar_g\n3=saturated_fat_g\n4=monounsaturated_fat_g\n5=polyunsaturated_fat_g\n6=trans_fat_g\n7=cholesterol_mg\n8=sodium_mg\n9=potassium_mg\n10=calcium_mg\n11=iron_mg\n12=magnesium_mg\n13=phosphorus_mg\n14=zinc_mg\n15=copper_mg\n16=manganese_mg\n17=selenium_ug\n18=chloride_mg\n19=chromium_ug\n20=iodine_ug\n21=molybdenum_ug\n22=vitamin_a_ug\n23=vitamin_c_mg\n24=vitamin_d_ug\n25=vitamin_e_mg\n26=vitamin_k_ug\n27=thiamin_b1_mg\n28=riboflavin_b2_mg\n29=niacin_b3_mg\n30=pantothenic_acid_b5_mg\n31=vitamin_b6_mg\n32=biotin_b7_ug\n33=folate_b9_ug\n34=folic_acid_ug\n35=vitamin_b12_ug\n36=choline_mg\n37=omega3_g\n38=omega6_g\n39=caffeine_mg\n40=water_g\n41=starch_g\n42=added_sugar_g\n43=total_unsaturated_fat_g\n44=fluoride_mg",
        +  "type": "string"
        +}
      • addedInput schema / properties / fiber_g
        Added value: +{
        +  "description": "Optional expanded nutrient: dietary fiber (g).",
        +  "type": "number"
        +}
      • addedInput schema / properties / recipe_fiber_g
        Added value: +{
        +  "description": "Optional recipe-component fiber (g) for add_both.",
        +  "type": "number"
        +}
      • addedInput schema / properties / recipe_saturated_fat_g
        Added value: +{
        +  "description": "Optional recipe-component saturated fat (g) for add_both.",
        +  "type": "number"
        +}
      • addedInput schema / properties / saturated_fat_g
        Added value: +{
        +  "description": "Optional expanded nutrient: saturated fat (g).",
        +  "type": "number"
        +}
  9. 2 tool updates
    • Changedcreate_goal1 field changed
      • changedInput schema / properties / goal_type / enum
        Previous value: -[
        -  "weight_loss",
        -  "strength",
        -  "consistency",
        -  "body_comp",
        -  "race",
        -  "n1_experiment",
        -  "daily_calories",
        -  "daily_protein_g",
        -  "daily_fat_g",
        -  "daily_carbs_g",
        -  "weekly_workouts",
        -  "daily_steps",
        -  "daily_cycling_distance_mi",
        -  "weekly_azm_target",
        -  "sleep_hours_target",
        -  "body_weight_target_lb",
        -  "body_fat_pct_target"
        -]New value: +[
        +  "weight_loss",
        +  "strength",
        +  "consistency",
        +  "body_comp",
        +  "race",
        +  "n1_experiment",
        +  "daily_calories",
        +  "daily_protein_g",
        +  "daily_fat_g",
        +  "daily_carbs_g",
        +  "weekly_workouts",
        +  "daily_steps",
        +  "daily_cycling_distance_mi",
        +  "daily_running_distance_mi",
        +  "weekly_azm_target",
        +  "sleep_hours_target",
        +  "body_weight_target_lb",
        +  "body_fat_pct_target"
        +]
    • Changedlist_workouts1 field changed
      • addedInput schema / properties / all_history
        Added value: +{
        +  "description": "Set true for all stored workout history instead of the recent default. Finds the earliest stored workout automatically; activity_type narrows that earliest-date lookup when present.",
        +  "type": "boolean"
        +}
  10. 2 tool updates
    • Changedlog_meal1 field changed
      • changedInput schema / properties / calories / description
        Previous value: -"Total calories (kcal). See MACROS above."New value: +"Total calories(kcal) never kJ"
    • Changedupdate_meal1 field changed
      • changedInput schema / properties / calories / description
        Previous value: -"Calories (kcal). See THREE MODES for totals vs additions and conditional requirements."New value: +"Calories (kcal) never kJ"
  11. 1 tool update
    • Changedlog_meal1 field changed
      • addedInput schema / properties / meal_time
        Added value: +{
        +  "description": "Optional local 24-hour HH:MM time. For an attached meal photo, use the current turn's photo capture time when the user did not state a different date/time. User-stated timing always wins.",
        +  "type": "string"
        +}
  12. 1 tool update
    • Changedlog_meal2 fields changed
      • changedInput schema / properties / meal_type / description
        Previous value: -"Required. Infer from time of day or context, even when recipe_name is used (a saved recipe's own stored meal type never fills this in)."New value: +"Optional. Omit for a current-day plain meal/food log so the server assigns the category from the user's resolved local clock. Send only when the user explicitly names or clearly anchors a category, or when backdating and the current clock cannot apply. A saved recipe's stored meal type never fills this in."
      • removedInput schema / required
        Removed value: -[
        -  "meal_type"
        -]
  13. 1 tool update
    • Changedlist_workouts2 fields changed
      • addedInput schema / properties / activity_type
        Added value: +{
        +  "description": "Stored workout focus type to match case-insensitively, for example Stair Stepper. Defaults to no activity-type filter.",
        +  "type": "string"
        +}
      • addedInput schema / properties / exercise_name
        Added value: +{
        +  "description": "Exercise name that must appear in the workout sets, matched case-insensitively. Defaults to no exercise filter.",
        +  "type": "string"
        +}

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to query live Garmin Connect health and fitness data, including daily metrics, activities, sleep analysis, and trends via natural language.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI assistants such as Claude, ChatGPT and Gemini to read wearable health data — daily activity, sleep stages, heart rate, SpO2, HRV, workouts, weight and nutrition — through a self-hosted, OAuth-protected Cloudflare Worker endpoint. It also lets users log weight, body fat, water, food, exercises, mood and symptoms by voice or chat, while leaving sensor-recorded data read-only.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.