Skip to main content
Glama
andronaft

health-os

health-os

tests Listed on mcpservers.org

Local-first personal health record, exposed over MCP. Your labs, diagnoses, medications, wearable data and food log live in your own Postgres; any MCP client — one running a local model or a cloud assistant — can read and update them through guarded tools. Critical values, drug-safety rules and screening schedules are deterministic code, not LLM judgement.

Medical disclaimer. This is not a medical device and does not give medical advice. Critical-value alerts and screening reminders are only a signal to contact a doctor — never a diagnosis and never a reason to delay care. Use at your own risk.

What it does

  • Lab results — drop a PDF or photo into your MCP client; the model extracts the values, health-os normalizes names (uk/ru/en/Latin synonyms) and units, and stages the panel as pending. Nothing counts as fact until you approve it.

  • Safety net in code — critical values alert immediately (log, macOS notification, optional Telegram); critical findings in narrative reports are flagged; drug-interaction questions are refused and redirected to a doctor/pharmacist (only deterministic checks run: total daily paracetamol across products, biotin before lab tests); a crisis tool returns a fixed response with hotlines, independent of the model.

  • Trends and analytics — Mann-Kendall trends, personal baselines and anomalies, age-gated risk calculators, a screening calendar, a weekly report, a doctor-visit brief.

  • Food log — meals with a 41-nutrient profile, %RDA, deficiency/excess flags, meal templates.

  • Devices — Apple Health export and Garmin import.

  • 28 MCP tools + server instructions — the safety rules are sent to every client on connect; see mcp_server/README.md.

Related MCP server: indaga-agent

Try it in one command

Only Docker needed. Starts a throwaway database with a fictional patient — two years of labs (LDL creeping up), blood pressure, medications, a food log and a lab panel awaiting approval:

git clone https://github.com/andronaft/health-os && cd health-os/demo
docker compose up -d --build

Point your MCP client at it:

{
  "mcpServers": {
    "health-os-demo": {
      "command": "docker",
      "args": ["run", "-i", "--rm", "--network", "health-os-demo",
               "-e", "DATABASE_URL=postgresql+psycopg://health:demo@db:5432/health_os",
               "health-os:local"]
    }
  }
}

Ask "show my health summary", "is my LDL trending up?", "what's pending review?", "what am I short on nutritionally?". Remove it all with docker compose down -v.

Install for your own data

Requires Docker and Python 3.12+.

git clone https://github.com/andronaft/health-os && cd health-os
cp .env.example .env                  # set the passwords
docker compose up -d db               # Postgres 16 + pgvector
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/alembic upgrade head        # schema
.venv/bin/python -m seed.load         # marker catalog, synonyms, units, nutrients
.venv/bin/python -m seed.demo         # optional: a fictional demo patient to play with

Then connect an MCP client — config for LM Studio, Open WebUI, Ollama CLI and Claude is in mcp_server/README.md. Try: "show my health summary", "LDL trend", "what am I short on nutritionally this week?".

Local models: the server speaks standard MCP over stdio, so any MCP client that runs a local model can use it. Verified so far: the server itself with the official MCP Python client (CI + the Docker demo). Not yet verified end-to-end with a local model — see #9; reports welcome.

How it works

MCP client (local or cloud model)
        │ stdio
   mcp_server/  ── read tools ──▶ approved views (read-only role, 5s timeout, row limits)
        │        ── write tools ─▶ core/services: normalize → status → critical rules → pending
        │
   safety/    critical values, narrative flags, interactions, crisis, alerts
   analytics/ trends, baselines, calculators, screening, nutrition, weekly report
        │
   PostgreSQL 16 + pgvector  ◀── ingestion/ (Apple Health, Garmin, embeddings)

Directory

What's inside

core/

config, DB, normalization, services, dedup, health summary

mcp_server/

MCP server, read and write tools

safety/

deterministic safety rules and alert delivery

analytics/

trends, baselines, calculators, screening, nutrition, reports

ingestion/

extraction schema, confidence scoring, device importers, embeddings

migrations/

Alembic schema

seed/

reference catalog + the demo patient

evals/

red-team scenarios (injections, hidden critical values, unit tricks)

scripts/

backup/restore (restic; install_launchd.sh schedules them on macOS), read-only role setup, importers

Privacy / local-first

  • Your data stays in your own database. Postgres runs locally in Docker; data/ and .env are outside git. Nothing is sent anywhere by health-os itself.

  • What leaves the machine depends on the MCP client you connect. With a local model (LM Studio, Open WebUI + Ollama, …) nothing does. With a cloud assistant, whatever the tools return is sent to that provider — use one whose terms fit medical data (no training on your data, zero/short retention).

  • The goal is fully local: local models for chat and extraction, local embeddings for search (already supported via fastembed). Cloud clients remain optional.

  • Optional alert channel (Telegram) sends only a generic "check your health system" text, never values.

  • Encrypt the disk (FileVault / LUKS / BitLocker) — the database files are plaintext at rest.

  • Never put real medical data in issues, PRs or tests — synthetic data only.

Tests

make test              # everything (needs Postgres for the integration part)
make test-unit         # pure unit tests — no database needed
make test-integration  # only tests marked `integration`

Integration tests never touch the working database: tests/conftest.py drops and recreates <POSTGRES_DB>_test on the same server (migrations + seed) on every run. Override with TEST_DATABASE_URL (the name must end in _test). Without Postgres, integration tests are skipped locally; CI sets REQUIRE_DB=1 so they fail instead.

Development history: PROGRESS.md.

License

AGPL-3.0-or-later. You may use, modify and fork health-os; if you distribute it or run a modified version as a network service, you must publish your source under the same license.

Want to use it in a closed-source or commercial product without those obligations? A separate commercial license is available from the author — reach out via GitHub (@andronaft).

Contributions are welcome — see CONTRIBUTING.md (includes a short CLA).

Available Tools

28 tools
approve_staged_sourceA
DestructiveIdempotent

Approve a staged panel (pending→approved). ONLY on an explicit instruction from the user in the current message — do not call right after stage_lab_panel. Refuses when numeric values have an unconvertible unit (they'd be invisible to trends); allow_missing_canonical=true only if the user accepts that.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idYes
allow_missing_canonicalNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructive, idempotent, non-read-only), so the bar is lower, yet the description still adds real behavior: it will refuse panels with unconvertible units and hides them from trends. It does not spell out post-approval effects on downstream data, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the action and state change, then the critical gating rule, then the parameter caveat. No filler, though the parenthetical aside slightly interrupts flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need no explanation, and the description covers the confirmation requirement, the refusal condition, and the one non-obvious parameter. Complete enough to call safely; only minor downstream-effect detail is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden. It explains allow_missing_canonical's consequence ('only if the user accepts that' the values are invisible to trends), which is meaningful beyond the bare boolean. source_id is left implicit but is self-evident from the tool's subject.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource plus the state transition (pending→approved), which an agent cannot get from the name alone. It also implicitly differentiates from the staging sibling by naming stage_lab_panel in the guidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit precondition ('ONLY on an explicit instruction from the user in the current message') and an explicit exclusion ('do not call right after stage_lab_panel'). This is exactly the when/when-not guidance the dimension rewards.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_medication_safetyA
Read-only

Call for ANY question about medications: "can I take X", combining drugs, doses. Returns the standard refusal to assess interactions (a doctor/pharmacist must check) plus deterministic checks: total daily paracetamol across products and biotin interference with lab tests. paracetamol_products: [{name, mg_per_dose, doses_per_day}] — include combination cold/flu remedies; if omitted, current medications are used. planned_tests: marker codes (e.g. ["tsh", "ferritin"]).

ParametersJSON Schema
NameRequiredDescriptionDefault
planned_testsNo
taking_biotinNo
paracetamol_productsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint. The description goes well beyond by disclosing the actual return behavior: a standard refusal to assess interactions (a doctor/pharmacist must check) plus deterministic checks for total daily paracetamol across products and biotin interference with lab tests. The fallback behavior when paracetamol_products is omitted is also disclosed, which is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The trigger condition is front-loaded in the first clause, followed by return behavior and parameter notes. It is dense but every sentence carries information; the parameter-level detail in the same paragraph slightly muddies the structure but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-required-arg tool with an output schema, the description covers the trigger, the return semantics, the refusal caveat, and defaults for omitted parameters. It is nearly complete; minor omissions are the purpose of taking_biotin being stated only implicitly and behavior when no medication data exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it gives the paracetamol_products shape ({name, mg_per_dose, doses_per_day}), warns to include combination cold/flu remedies (a non-obvious edge case), and gives planned_tests format plus an example (['tsh','ferritin']). taking_biotin is only indirectly implied by the biotin-interference mention, leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource and an explicit trigger scope ('Call for ANY question about medications: can I take X, combining drugs, doses'), then states what it returns: a standard refusal to assess interactions plus deterministic checks. This clearly differentiates it from siblings like get_medications and record_medication, which merely read or write medication data rather than answer safety questions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides strong usage context ('ANY question about medications') and explains a conditional behavior — if paracetamol_products is omitted, current medications are used instead. It also implicitly tells the agent this is not a substitute for medical advice via the refusal note. It lacks an explicit 'do not use for X, use sibling Y' exclusion, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crisis_resourcesA
Read-only

Call IMMEDIATELY on any sign of crisis, suicidal thoughts or self-harm. Returns a fixed response with hotlines and the user's trusted contact. Reply with it as is — no analytics.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior the annotations cannot convey: the response is fixed and pre-authored, and the agent must reply with it verbatim with 'no analytics' — an important constraint on how the result is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, urgency front-loaded in the first clause, with each sentence carrying distinct information: when to call, what comes back, and how to handle it. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still usefully characterizes the return value, and it covers the trigger and the required handling behavior. The only real gap is the unexplained 'message' input, which is a minor omission for an optional parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single optional 'message' parameter has no title-level explanation. The description never mentions the parameter or what the agent should pass in it (e.g., the user's triggering message), so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool produces — a fixed response containing hotlines and the user's trusted contact — which is unambiguous and clearly distinct from every sibling health-data tool. An agent can tell at a glance that this is not a data-retrieval or analytics tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit, unmissable trigger: 'Call IMMEDIATELY on any sign of crisis, suicidal thoughts or self-harm.' The when-not condition is only implied (i.e., do not call otherwise), and no alternative tool is named — though none exists among the siblings — so this falls just short of the full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_allergiesB
Read-only

Allergies, including unverified (fail-safe: treated as an allergy).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — unverified allergies are fail-safe treated as allergies — but says nothing about how unverified vs verified entries are surfaced or distinguished in the response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler and the fail-safe caveat front-loaded. It is arguably under-specified rather than verbose, but nothing in it wastes space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a safety-relevant read tool with an empty input schema, the description gives no usage context or relational guidance to the sibling record_allergy, leaving the definition minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case; there is no parameter semantics to explain and the description correctly does not invent any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase ('Allergies, including unverified...') rather than a verb+resource statement, so the agent must infer that this retrieves allergy data. It does distinguish the scope (includes unverified entries), but it never names the sibling record_allergy or clarifies the read/write boundary between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as record_allergy or query_observations. The agent is left to infer that this is the read path for allergy data purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_diagnosesC
Read-only

Diagnoses with two status axes (clinical_status + verification_status). Advice — only on confirmed; suspected — in question.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds almost nothing beyond that: it does not say whether all diagnoses are returned or how they are scoped, ordered, or paginated. The only extra signal is that records carry two status axes, which is output detail the output schema likely covers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no padding, which is appropriately sized. But the second sentence is cryptic and ambiguous — 'Advice' and 'in question' are not standard status labels — so it does not clearly earn its place alongside the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema present, the description is not required to explain return values, and the read-only annotations cover safety. It is still incomplete in an important respect: it never states the scope of the result set (all diagnoses for the user?), leaving the agent unable to predict what it will receive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline this scores 4. The description makes no misleading claims about inputs, though it also has no parameters to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names the resource (diagnoses) and one of its key attributes (the clinical_status + verification_status axes), which partly distinguishes it from siblings like get_medications and get_allergies. However, it never states the action — no verb such as 'list' or 'retrieve' — so the agent must infer it is a read operation for the user's diagnoses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence reads as a gloss on status values ('Advice — only on confirmed; suspected — in question') rather than guidance on when to call this tool. There is no indication of when to prefer it over get_health_summary, get_timeline, or query_observations, nor any prerequisite or exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_health_summaryA
Read-only

Deterministic health summary: profile, allergies (including unverified), active diagnoses, current medications, recency of exams. A guide — exact values via query_observations.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine behavioral context beyond that: the output is 'deterministic' (stable for a given state) and deliberately includes unverified allergies, which tells the agent how to interpret data quality. It doesn't mention freshness/pagination, but for a parameterless read that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste. The content enumeration is front-loaded and the routing hint to query_observations comes last as the qualifier it is.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description is not obliged to explain return values, and it does list the sections an agent should expect. The only gap is the absence of any statement about how this aggregate relates to the single-domain siblings (get_allergies, get_diagnoses, get_medications), which an agent choosing between them would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema imposes no documentation burden and the baseline is 4. The description correctly spends no words on parameter syntax and instead describes the content of the result.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource and enumerates exactly what the summary contains: profile, allergies (including unverified), active diagnoses, current medications, and exam recency. It also distinguishes itself from query_observations by labeling itself 'a guide' rather than a source of exact values, so an agent can separate it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context (overview/guide use) and names the alternative for exact values: 'exact values via query_observations.' That is an explicit alternative with a selection condition, though it never states when NOT to use this tool or how it relates to the narrower get_allergies/get_diagnoses/get_medications siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_medicationsA
Read-only

Current medications with doses (status='taking').

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered by structured data. The description adds the meaningful scoping trait that only medications with status='taking' are returned, but says nothing about ordering, recency, or whether historical meds are excluded outright.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence carrying verb, resource, payload, and filter with zero waste. Nothing extraneous and the key constraint is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and for a zero-param read tool this is nearly sufficient. The only residual gap is a lack of routing context against the many sibling read tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description's parenthetical restates the only implicit input concept (status='taking'), but no parameter semantics are actually needed here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('get_medications') plus explicit scope: current medications with doses filtered to status='taking'. An agent can distinguish it from the write-side sibling record_medication. It does not, however, explicitly differentiate itself from check_medication_safety or get_allergies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The filter condition (status='taking') implies when this returns data, but there is no explicit guidance on when to prefer this over siblings such as check_medication_safety or get_health_summary, and no exclusions stated. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_screening_recommendationsA
Read-only

Screening calendar for the user's profile (age-gate, "you don't need this yet"). Statuses: due/overdue/up_to_date/not_yet. Requires a filled-in profile.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral context on top: the age-gate behavior ('you don't need this yet'), the enumerated status values, and a hard precondition that the profile must be filled in.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the resource statement front-loaded and the prerequisite trailing. The parenthetical age-gate phrase is slightly idiomatic but conveys real information, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description need only frame the tool and its precondition, which it does. It is complete enough to invoke correctly, though it omits any guidance on what to do when the profile is unfilled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The mention of statuses describes output semantics rather than inputs, which is harmless but not a parameter clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource ('screening calendar') scoped to the user's profile, with the caveat that it is age-gated. It is clearly distinguishable from data-retrieval siblings like get_health_summary or query_observations, but it never explicitly contrasts itself with any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case (viewing a screening calendar) is implied, and one prerequisite is stated: 'Requires a filled-in profile.' However no alternatives or when-not conditions are given, e.g. nothing says what to call instead if the profile is empty (set_profile is a sibling).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_timelineC
Read-only

Chronology of health events (diagnoses, visits, panels, medications, hospitalizations, vaccinations).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds nothing behavioral beyond the content list: it does not say how results are ordered, how far back the default window reaches, or whether the list is capped. Listing event types is purpose information, not behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler, and the resource is stated up front. The parenthetical enumeration is long but each item earns its place by defining scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required. However, for a tool with a single undocumented time-window parameter and many overlapping siblings, the definition leaves the agent without enough to call it confidently versus alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'days' parameter, so the schema supplies only a name and a default of 3650. The description never mentions the time-window parameter at all, leaving an agent to guess whether the default ~10-year span is intended and whether the argument filters or paginates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (a health-event chronology) and enumerates exactly what it contains (diagnoses, visits, panels, medications, hospitalizations, vaccinations), which implicitly distinguishes this aggregate view from single-domain siblings like get_diagnoses or get_medications. It is a noun phrase rather than a verb+resource, so the retrieval action itself is only implied, but the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this over get_health_summary, get_trend, get_diagnoses, or query_observations, all of which overlap in subject matter. No prerequisites, time-window conventions, or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_trendC
Read-only

Marker trend (Mann-Kendall): increasing/decreasing/no_trend + significance.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
type_codeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and an output schema covers return values, so the bar is lower. The description does add value by naming the statistical method (Mann-Kendall) and the discrete outcome categories, but says nothing about lookback behavior, data requirements, or how significance is reported.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause carrying the method and the output vocabulary, with zero filler. It is efficient, though its brevity shades into under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema handles return values, but with a required, undocumented type_code, an unexplained days window, and no usage context, the description is not sufficient for an agent to invoke this correctly. More must be said about the required identifier and the time window.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with two parameters, one required. The description never explains what type_code identifies (a marker id? a code system?) or that days is a lookback window (default 1825 ≈ five years). It leaves both parameters semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation (marker trend via Mann-Kendall) and names the returned categories (increasing/decreasing/no_trend + significance), which is far more specific than the bare name 'get_trend'. It does not, however, distinguish itself from siblings like query_observations or get_timeline that also surface health data over time.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The agent cannot tell from the description when a statistical trend call is preferable to the query/timeline siblings, nor what 'marker' covers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_reportB
Read-only

Deterministic weekly report + health metrics of the system itself (pending-queue size).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuine behavioral trait — 'Deterministic' — implying repeatable, non-stochastic output, plus the content of the pending-queue metric. It says nothing about time window, caching, or cost, but with annotations and an output schema present the bar is lower.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the primary deliverable and keeping the secondary metric subordinate in parentheses. No filler, though the telegraphic '+' construction is slightly terse relative to how much ambiguity remains.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with an output schema, the description does not need to enumerate return values, and it correctly signals that one of them is system pending-queue size. The main gap is selection guidance against the many report-like siblings, which the output schema cannot compensate for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The description implies the report period ('weekly') without needing any argument, which is consistent with the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a concrete deliverable ('weekly report') plus a specific extra payload ('health metrics of the system itself (pending-queue size)'), so an agent knows this is an operational/system report rather than a clinical one. The phrase 'of the system itself' partially disambiguates from clinical siblings like get_health_summary and nutrition_report, but it never states the relationship explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use statement, no prerequisites, and no alternatives named. With siblings such as get_health_summary, nutrition_report, get_timeline and get_trend available, the agent gets no signal about when this tool is the right choice over those.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_meal_templatesA
Read-only

Saved templates for frequent meals (for quick log_from_template).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds only that the returned records are saved templates for frequent meals; it says nothing about ordering, limits, or whether all templates are returned, though the output schema carries return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource front-loaded and no filler. It is telegraphic, but every word earns its place and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read tool with an output schema and annotations, the definition is minimally adequate - it tells the agent what the records are and hints at the downstream workflow. It stops short of confirming scope (all templates vs. filtered) or the role it plays alongside save_meal_template.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies for a parameterless listing tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource precisely ('Saved templates for frequent meals'), which lets an agent distinguish it from log_from_template, save_meal_template, and log_meal. It is a noun phrase rather than an explicit verb, but the tool name supplies 'list' and the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(for quick log_from_template)' implies the intended workflow - fetch templates before logging from one - but states no explicit when-to-use condition or alternative to choose instead. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pending_reviewsB
Read-only

Markers in the review queue (NOT confirmed — do not cite as fact).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=false), so the bar is lower, and the description adds genuinely useful context the annotations cannot: these entries are unconfirmed and must not be treated as fact. It stops short of explaining how items enter or leave the queue, but the confirmation caveat is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that front-loads the resource and puts the critical caveat in parentheses where it cannot be missed. Nothing is wasted, though the terseness trades away some clarification an agent might need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no prose, and there are no parameters to document. However, the description never explains the relationship to the sibling that resolves staged items (approve_staged_source), leaving the workflow around this queue ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline is 4 by rule. Schema description coverage is 100% regardless.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Markers in the review queue' identifies the resource being listed, but it is a noun fragment that relies on the tool name for its verb and never clarifies what a 'marker' actually is. It also does not distinguish this queue-view from the sibling approve_staged_source, which appears to act on the same staging area.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to reach for this tool versus query_observations, get_timeline, or approve_staged_source, and no prerequisites or trigger conditions. The only guidance is the output-handling caveat about not citing the data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_from_templateB

Log a meal from a saved template (nutrients × portion_factor). eaten_at=ISO (default — now). Quick entry of frequent meals in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
eaten_atNo
symptomsNo
wellbeingNo
portion_factorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (not read-only, not idempotent, not destructive), so the description only needs to add context. It discloses the portion scaling semantics and that eaten_at defaults to now, but says nothing about what happens when the template name doesn't exist or whether logging is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense clauses with no filler, and the core action is front-loaded. The parenthetical multiplication is compact and informative rather than bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the write semantics are covered by annotations. Still missing the failure mode for an unknown template name and the timezone/format expectation for eaten_at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden and only partially compensates: it explains portion_factor as a multiplier and eaten_at as an ISO timestamp defaulting to now, but leaves symptoms and wellbeing entirely undefined in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Log a meal from a saved template'), and the parenthetical '(nutrients × portion_factor)' clarifies the computation. The 'from a saved template' qualifier implicitly separates it from log_meal and save_meal_template, though neither sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Quick entry of frequent meals in one call' implies the intended use case (repeat meals) but never states when to prefer this over log_meal or how to discover valid template names via list_meal_templates. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_mealA

Log a meal into the diary (auto-approved). description — the meal description (required); meal_type: breakfast/lunch/dinner/snack/drink; eaten_at=ISO 'YYYY-MM-DD HH:MM' (default — now). nutrients — {code: amount} per nutrient_types (energy_kcal/protein/carbs/fat/fiber/sugar/ added_sugar/saturated_fat/omega3/sodium/potassium/calcium/iron/magnesium/zinc/vitamin_a/ vitamin_c/vitamin_d/vitamin_b12/folate_b9/water/caffeine/alcohol/...). added_sugar — sugar ADDED to the dish (not natural from fruit/milk). The full-profile estimate is made by the model (including from a photo). glycemic_index/glycemic_load — per meal; symptoms/wellbeing — reaction after eating. Review — query_food; daily norms — query_nutrition.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNo
portionNo
eaten_atNo
symptomsNo
meal_typeNo
nutrientsNo
wellbeingNo
descriptionYes
glycemic_loadNo
glycemic_indexNo
nutrient_sourceNomodel_estimate

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, non-idempotent mutation; the description adds genuinely new context by disclosing that entries are auto-approved, which matters given the existence of approve_staged_source and list_pending_reviews siblings. It also explains that nutrient estimation is model-generated 'including from a photo'. It stops short of describing edit/delete or duplicate behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core action and the required parameter lead, followed by per-parameter semantics in a compact semicolon-delimited form. The em-dash 'Review — query_food' fragments are slightly awkward but every line carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter mutation tool with an output schema (so return values need no explanation), the description covers nearly all parameters and the approval behavior. Missing minor coverage of notes, portion, and nutrient_source, but nothing an agent needs to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load and largely does: it documents meal_type values, the eaten_at ISO format and default, the {code: amount} shape of nutrients with a long code list, the added_sugar vs natural-sugar distinction, glycemic_index/load scope, and symptoms/wellbeing semantics. Only notes, portion, and nutrient_source (default model_estimate) go unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Log a meal into the diary') and adds the consequential detail that the entry is auto-approved. It does not explicitly name which sibling it supersedes (save_meal_template, log_from_template, list_pending_reviews), though it does point to query_food and query_nutrition at the end.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the closing pointers ('Review — query_food; daily norms — query_nutrition'), which route the agent for follow-up reads but never state when to choose this tool over log_from_template or save_meal_template. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nutrition_reportA
Read-only

Nutrition analytics over N days: top deficiencies/excesses (%RDA+flags) + food's link to wellbeing (average GI/sugar/sodium by wellbeing category). Association, not causation.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the description adds real value by disclosing the analysis window semantics and, crucially, the interpretation caveat 'Association, not causation,' which prevents the agent from over-claiming causal links. It does not mention computation cost or data-sufficiency requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the analytic subject comes first, outputs second, and the caveat last. The phrasing is dense with domain shorthand (%RDA+flags, GI), but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter analytics tool with an output schema present, the description covers purpose, window semantics, output composition, and interpretation limits. The only gap is routing guidance against the many sibling query/report tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the single 'days' parameter has no documentation in the schema beyond its title. The description compensates by framing it as the analytics window ('over N days'), clarifying that it drives the lookback period rather than acting as a filter — the schema default of 30 is the only remaining detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Nutrition analytics over N days') and enumerates the concrete outputs (deficiencies/excesses with %RDA flags, GI/sugar/sodium grouped by wellbeing category). It implicitly separates itself from raw-data siblings like query_nutrition, though it never names them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to use this report versus query_nutrition, get_trend, or get_weekly_report. The analytic nature is implied by the output description, but the agent must infer the selection condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_doctor_visitB
Read-only

Preparation package for a visit: summary + recent abnormalities + screening due + pending queue.

ParametersJSON Schema
NameRequiredDescriptionDefault
specialtyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful composition context by listing the four aggregated sections, but says nothing about patient scoping, latency of an aggregate call, or whether the package is filtered by the specialty parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, and the enumerated components are easy to scan. It is arguably too terse for an aggregate tool, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the return shape need not be explained. However, for a composite aggregator the description leaves the specialty parameter's effect and the patient/context scope unexplained, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'specialty' parameter has no schema description, yet the description never mentions it at all. An agent cannot tell whether specialty filters the whole package, selects screening guidelines, or is merely a label, which is a real gap for a 1-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a concrete deliverable ('preparation package for a visit') and enumerates its four components (summary, recent abnormalities, screening due, pending queue), which lets an agent distinguish it from single-purpose siblings like get_health_summary or get_screening_recommendations. It does not explicitly name an alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a visit' weakly implies the pre-appointment scenario, but there is no explicit when-to-use, when-not-to-use, or routing to siblings such as get_health_summary plus get_screening_recommendations for a manual equivalent. An agent must infer the use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_foodB
Read-only

Food log over N days: meals (gi/gl/wellbeing) + nutrients per meal. Optional meal_type filter (breakfast/lunch/dinner/snack/drink).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
meal_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results include gi/gl/wellbeing plus nutrients per meal, which is useful content context, but says nothing about ordering, limits, or pagination behavior for the N-day window.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the resource and scope, with no filler. Minor drag from unexpanded abbreviations (gi/gl) that force the agent to guess, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, and annotations carry the read-only profile. The remaining gaps are the lack of differentiation from query_nutrition/nutrition_report and the undefined 'days' semantics, which leave an agent in a crowded sibling set with real ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It does supply the valid meal_type values (breakfast/lunch/dinner/snack/drink), which the schema lacks as an enum, but it gives no gloss for 'days' beyond the phrase 'over N days' and no default hint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete resource and scope: a food log over N days with meals and per-meal nutrients. That is more specific than the bare name, but it never distinguishes this tool from near-identical siblings such as query_nutrition and nutrition_report, so an agent cannot route confidently among them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the many food/nutrition siblings (query_nutrition, nutrition_report, log_meal, list_meal_templates). The only usage-like detail is that the meal_type filter is optional, which is a parameter fact rather than guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_nutritionB
Read-only

"Healthiness" over N days: average daily intake of each nutrient, %RDA and flags deficient (<70% of norm)/excess (>upper limit). Vitamins, minerals, sodium, sugar, fats.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered. The description adds useful behavioral detail beyond that: the flag thresholds (<70% of norm = deficient, >upper limit = excess). It does not disclose aggregation semantics, auth needs, or rate limits, so it is solid but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense, front-loaded sentence with no filler; the headline computation and the flag rule both appear immediately. Slightly packed with parentheticals, but every clause carries meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be explained, and the description instead conveys what the metrics mean (average, %RDA, deficiency/excess thresholds). For a one-parameter read-only query this is nearly complete, with only the usage-vs-sibling gap remaining.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single 'days' parameter has no schema description, so the description must compensate. It does imply the parameter's role via 'over N days' and a default-adjacent window, but never states the default of 7 or the accepted range, so compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific computation (average daily intake of each nutrient over N days) with concrete outputs (%RDA and deficiency/excess flags). It is clear what the tool does, though it does not explicitly distinguish itself from the sibling nutrition_report, leaving that distinction to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this versus alternatives. The sibling set contains nutrition_report, query_food, and get_health_summary, which plausibly overlap, yet the description offers no routing condition, exclusions, or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_observationsA
Read-only

Values of a marker (e.g. 'cholesterol_total') over N days. >90 days → weekly aggregation min/avg/max. Only confirmed (approved) values.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
type_codeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint, so the description carries real weight here and delivers: it discloses the >90-day weekly min/avg/max aggregation rule and that unconfirmed values are excluded. That is genuine behavioral context beyond the annotations, though response shape/pagination is left to the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse fragments, core purpose front-loaded, no filler. The telegraphic style is efficient though it reads more like shorthand notes than a polished description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. For a two-parameter read-only query the description covers scope, aggregation behavior, and data filtering; only the type_code vocabulary and the days default remain implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It gives an example value for type_code and explains that days controls the window and triggers weekly aggregation past 90 days, but says nothing about the default of 365 or valid code formats, leaving part of the burden unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (marker values) and the shape of the query (over N days) with a concrete example code ('cholesterol_total'), so the agent knows exactly what it retrieves. It does not distinguish itself from similar-looking siblings such as get_trend or get_health_summary, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to prefer this tool over alternatives like get_trend or query_nutrition, and no prerequisites. The 'only confirmed (approved) values' clause hints at data scope but is a filter, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_allergyC

Add an allergy (verified=false is ALSO treated as an allergy — fail-safe).

ParametersJSON Schema
NameRequiredDescriptionDefault
allergenYes
reactionNo
severityNo
verifiedNo
allergen_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this as a non-read-only, non-destructive, non-idempotent write, so the safety profile is covered. The description adds one genuinely useful behavioral rule — that verified=false is still treated as an allergy (fail-safe) — but says nothing about what gets created, whether duplicates are merged, or permission needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the action, with the fail-safe caveat parenthetically attached. It is efficient, though its brevity is partly under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter write tool with 0% schema coverage, the description is too thin — it omits the semantics of most parameters and any mention of the required 'allergen' field. The output schema existing means return values need not be described, but the input side remains incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description carries the burden and only addresses one of them (verified). The meanings of reaction, severity, and allergen_type are left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Add an allergy'), which clearly separates it from the read-side sibling get_allergies. It does not, however, distinguish it from the other record_* writers (record_diagnosis, record_medication) beyond the resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as record_diagnosis or record_medication, and no prerequisites or context are given. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_diagnosisC

Add a diagnosis (verification_status: suspected/…/confirmed/refuted; advice only on confirmed).

ParametersJSON Schema
NameRequiredDescriptionDefault
severityNo
icd10_codeNo
diagnosed_atYes
diagnosis_nameYes
clinical_statusNoactive
verification_statusNoconfirmed

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, non-destructive, non-idempotent write. The description adds one genuine behavioral nuance beyond that: advice is only surfaced for confirmed diagnoses, plus the verification_status vocabulary. It says nothing about duplicate handling, permissions, or what happens to previously recorded diagnoses.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single front-loaded sentence with no filler, which is good, but the truncated enum 'suspected/…/confirmed/refuted' wastes a slot on an ellipsis and the brevity edges into under-specification rather than tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter mutation tool with 0% schema description coverage and a rich sibling set of record_* tools, the definition is far too thin. An output schema exists so return values need not be explained, but input semantics and routing guidance are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters, so the description must compensate and largely does not. It only illuminates verification_status (values suspected/…/confirmed/refuted) and leaves diagnosis_name, diagnosed_at, severity, icd10_code, and clinical_status completely unexplained, including required-vs-optional expectations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Add) and resource (diagnosis), which is unambiguous on its own. It does not, however, differentiate itself from the sibling get_diagnoses or from other record_* tools like record_allergy and record_medication, so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no mention of prerequisites, and no named alternative. The parenthetical about advice only on confirmed is a downstream consequence, not a usage rule for invoking this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_medicationC

Add current medications/supplements (product_type: prescription/otc/supplement/herbal).

ParametersJSON Schema
NameRequiredDescriptionDefault
atc_codeNo
dose_unitNo
start_dateYes
dose_amountNo
product_typeNoprescription
times_per_dayNo
prescribed_forNo
medication_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a write (readOnlyHint=false) that is non-idempotent and non-destructive, so the safety profile is covered structurally. The description adds nothing beyond that: it does not disclose that repeated adds may create duplicate entries, whether existing records are updated or a new one is appended, or any auth/validation requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the most decision-relevant detail (scope plus product_type values) comes first. It is efficient, though the terseness leans toward under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for an 8-parameter, 0%-coverage mutation tool the description is far too thin: no usage context, no input semantics beyond one field, and no behavioral notes for a non-idempotent write.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters and there are no enums in the schema, so the description carries the full explanatory burden. It supplies only product_type's four accepted values (prescription/otc/supplement/herbal), leaving atc_code, dose_unit, dose_amount, times_per_day, prescribed_for, start_date format, and medication_name semantics entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Add') and resource ('current medications/supplements'), which is enough to distinguish it from read-side siblings like get_medications and the safety checker check_medication_safety. However, it does not explicitly name those siblings or clarify the write-vs-read boundary, so differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. Notably, the sibling set contains a staging/approval workflow (stage_lab_panel, list_pending_reviews, approve_staged_source) and check_medication_safety, yet the description does not indicate whether medications should be recorded directly or staged/reviewed first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_meal_templateB
Idempotent

Save a template for a frequent meal (by name). nutrients — {code: amount} per 1 portion. Then logged in one call via log_from_template(name).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
meal_typeNo
nutrientsNo
descriptionNo
glycemic_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a write (readOnlyHint=false), idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds the meaningful detail that nutrients are expressed as {code: amount} per 1 portion, but says nothing about whether saving over an existing name overwrites it or what persistence side effects occur on repeat calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines, front-loaded with the core action, and the nutrients format hint earns its space. The dangling 'Then logged in one call...' clause is slightly awkward as written but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is unnecessary. Still, for a 5-parameter write tool with zero schema descriptions, the definition leaves meal_type and glycemic_index unexplained and does not resolve the overwrite question implied by idempotentHint, so it is only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load, and it explains only the nutrients format ({code: amount} per 1 portion). The four other parameters — meal_type, description, and glycemic_index in particular — get no semantic explanation, leaving the agent to guess their acceptable values and units.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: saving a template for a frequent meal, keyed by name. It is distinguishable from log_meal (which logs an actual meal) and list_meal_templates, though it relies on the reader to infer those contrasts from sibling names rather than stating them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the pointer 'Then logged in one call via log_from_template(name)', which sketches the intended workflow. However, it never says when to create a template versus logging directly with log_meal, and it gives no prerequisites or exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_profileB
Idempotent

Create/update the profile (date_of_birth=YYYY-MM-DD, sex=male/female).

ParametersJSON Schema
NameRequiredDescriptionDefault
sexYes
height_cmNo
blood_typeNo
date_of_birthYes
emergency_contactNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the write/upsert safety profile is covered. The description's 'Create/update' reinforces but does not extend that. It never states whether omitted optional fields (height_cm, blood_type, emergency_contact) are left untouched or reset to their defaults, which is the main behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the two most error-prone parameters are annotated inline. The parenthetical is dense but readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Safety and idempotency are covered by annotations and an output schema exists, so return values need no explanation. Still, for a five-parameter write tool with 0% schema coverage, the description leaves the optional fields and their null/default semantics unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden. It usefully supplies the date format (YYYY-MM-DD) and the accepted values for sex (male/female), which the schema itself does not constrain via enum. However, three of five parameters (height_cm, blood_type, emergency_contact) remain entirely undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create/update the profile'. An agent immediately knows this writes the user's health profile. It does not name a related sibling (e.g. get_health_summary) for differentiation, but no sibling performs profile writes, so ambiguity is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Create/update' implicitly signals an upsert, but there is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives. The agent must infer that this is the sole entry point for profile data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sql_queryD
Read-only
ParametersJSON Schema
NameRequiredDescriptionDefault
sqlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Tool has no description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness1/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tool has no description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool has no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Tool has no description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tool has no description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tool has no description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stage_lab_panelA

Stage an extracted lab panel (PENDING). rows: a list of {raw_name, value, unit, ref_min, ref_max}. Critical values are alerted immediately. Afterwards — show the table to the user and wait for an explicit approve_staged_source.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsYes
facilityNo
panel_dateYes
panel_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true. The description adds genuinely useful behavioral context beyond that: the PENDING state, immediate alerting of critical values, and the requirement of explicit human approval before downstream approval. It stops short of specifying idempotency or re-staging behavior, which matters given idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and state, followed by the row shape and the required follow-up. No filler, though the trailing dash construction is slightly informal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The description covers purpose, the staged/PENDING lifecycle, critical-value behavior, and the mandatory approval follow-up, leaving only minor gaps around the non-rows parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden. It documents the rows item shape ({raw_name, value, unit, ref_min, ref_max}) well, but leaves facility, panel_type, and even the required panel_date format/expectations unaddressed, so it only partially compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Stage) and resource (extracted lab panel) plus the resulting state (PENDING), and it clearly positions itself before approve_staged_source. It does not name a sibling to avoid, but the two-step relationship is explicit enough to distinguish it from the read/report siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: stage now, then show the table to the user and wait for an explicit approve_staged_source. This tells the agent when the tool fits in the workflow and what must follow, though it doesn't state exclusions for panels already staged or alternative staging paths.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 28 tool updatesv0.2.0
    • First observedapprove_staged_source
    • First observedcheck_medication_safety
    • First observedcrisis_resources
    • First observedget_allergies
    • First observedget_diagnoses
    • First observedget_health_summary
    • First observedget_medications
    • First observedget_screening_recommendations
    • First observedget_timeline
    • First observedget_trend
    • First observedget_weekly_report
    • First observedlist_meal_templates
    • First observedlist_pending_reviews
    • First observedlog_from_template
    • First observedlog_meal
    • First observednutrition_report
    • First observedprepare_doctor_visit
    • First observedquery_food
    • First observedquery_nutrition
    • First observedquery_observations
    • First observedrecord_allergy
    • First observedrecord_diagnosis
    • First observedrecord_medication
    • First observedsave_meal_template
    • First observedsearch
    • First observedset_profile
    • First observedsql_query
    • First observedstage_lab_panel

TDQS

C2.8/5.0

Scored across 28 tools

Disambiguation4/5

Tool boundaries are mostly clear, with descriptions distinguishing overlapping areas like query_food vs query_nutrition vs nutrition_report and health_summary vs weekly_report vs prepare_doctor_visit. The main ambiguity is sql_query, which has no description and could overlap with search or query_observations.

Naming Consistency4/5

Names consistently use snake_case with verb_noun patterns such as get_, query_, record_, log_, list_, and save_. A few tools like search, sql_query, and crisis_resources deviate slightly from the verb_noun convention but remain readable.

Tool Count3/5

28 tools is heavy for a health assistant, and several summary/report tools could potentially be consolidated. However, the broad domain of records, labs, nutrition, safety, and screening means most tools have a distinct purpose, making the count borderline rather than excessive.

Completeness4/5

The surface covers core health workflows: profile setup, records, lab staging/approval, observations, trends, nutrition logging, reports, medication safety, and crisis resources. It lacks explicit update/delete operations for allergies, diagnoses, medications, and meals/templates, leaving minor lifecycle gaps.

Maintenance

ActivityMaintained
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    A local-first MCP server that enables AI agents to read user-authorized Google Health API v4 data from Fitbit, Pixel Watch, and partners via OAuth, with tokens never leaving the machine.
    26
    493 npm
    62
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local-first MCP server for querying multi-omic personal health data (genome, labs, wearables) with an honesty contract and progressive disclosure skills.
    1
    AGPL 3.0
  • A
    license
    C
    quality
    D
    maintenance
    A local-first, model-agnostic MCP server that stores personal health data in a SQLite file and provides analysis-ready views for any AI client to log, retrieve, and reason over health records.
    79
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI clients to securely access, query, and mutate normalized user-controlled health data (e.g., from Apple Health or Supabase) through a bounded set of MCP tools, with optional OAuth and sandboxed deployment.
    MIT