Skip to main content
Glama
skiitk2510

CareCompanion

by skiitk2510

CareCompanion

Medication safety guardrails for Alexa+. A self-hosted MCP server that lets a voice assistant say no safely: it refuses a duplicate or too-soon dose until the person confirms and gives a reason, flags drug interactions and allergies when a caregiver adds a medication, escalates symptoms to the family by severity, and ships the family dashboard as an MCP App that Alexa+ renders inline. Packaged with an Agent Skill, a classic Alexa Skill front end, and a simulated Alexa+ experience for the demo.

Built for the Build, Ship, Shape: Amazon Developer Hackathon (Alexa+ track · AWS Builder mini · Open Source mini). MIT licensed; created from scratch during the hackathon window.

Status: MCP server, guardrails, Bedrock brain, Agent Skill and both UIs are built and tested. Live deployment, demo video and final polish are in progress (see Status).

What ships

Piece

Where

What it is

MCP server (Streamable HTTP + stdio)

src/

8 tools · 1 resource · 1 prompt on POST|GET|DELETE /mcp (spec 2025-11-25, sessionful), plus a carecompanion-mcp stdio binary.

Guardrails

src/domain/guardrails/

Duplicate/too-soon dose guard with explicit confirmation; informational interaction + allergy warnings; fail-safe symptom escalation.

MCP App screens

ui/ → ui://carecompanion/dashboard.html

One view bundle, three screens: the family dashboard (caregiver_summary, inline + fullscreen), the elder's today card (get_todays_plan) and the dose-guard card (log_dose, "Record anyway" needs a reason). Rendered inline by MCP App hosts (Alexa+, MCP Inspector, basic-host).

Live alerts

src/http/mcpRoutes.ts

Every new alert is pushed to all open MCP sessions as notifications/message, so a host holding a session hears about a missed dose or a symptom without polling.

Agent Skill

skills/carecompanion/SKILL.md

Teaches any agent the five safe workflows against the server (validated with skills-ref).

Simulated Alexa+ experience

web/ + src/agent/

An Echo-Show-style web app with browser voice; its brain is Amazon Bedrock (Claude Haiku 4.5) calling the MCP tools through a real MCP client.

Classic Alexa Skill

src/alexa/ + skill-package/

A real Alexa Skills Kit skill ("open care companion") whose endpoint is this server: every utterance goes through the same agent loop and MCP tools, so Amazon's own Alexa simulator or an Echo device can drive the demo.

Conformance probe · add-on package · evals

scripts/verify-live.mjs, addon-package/, docs/evals/

A 20-rule graded check of the deployed /mcp (stand-in for Amazon's Local Inspector); an Alexa+ add-on manifest with icons and legal docs ready for alexa-ai new mcp; 12 persona evaluations with committed transcripts.

Related MCP server: Ansim Dolbom Assistant

Quickstart

Requires Node 22.19+ (24 recommended). No AWS account is needed for the MCP server or the web app: without Bedrock a rule-based brain answers the demo conversations.

git clone https://github.com/skiitk2510/carecompanion.git
cd carecompanion
npm install
npm run build          # server + dashboard view + web app
npm start              # http://127.0.0.1:3000  (MCP at /mcp, web app at /, health at /healthz)

Then, in another terminal:

npm run verify:live -- http://127.0.0.1:3000   # graded conformance probe: 20 rules, writes docs/conformance/verdict.json
npm run mcp:lifecycle  # curl walk-through of the Streamable HTTP session lifecycle
npm run mcp:inspect    # MCP Inspector UI against the running server
npm run agent:smoke    # the five demo beats through POST /api/agent (prints which brain + tools answered)
npm test               # 180+ tests: guardrails, seed, MCP tools over the SDK client, agent loop, REST routes, conformance

The conformance probe (docs/conformance/README.md) is our own stand-in for Amazon's Local Inspector, which is not available to hackathon participants: it grades the sessionful lifecycle, the 2025-03-26 and 2025-11-25 handshakes, the MCP App resource, the 500 ms round-trip budget, error semantics, CORS and the Tier-1 auth behaviour, and it exits non-zero on any MUST failure so it can gate a deployment.

Local MCP hosts can use stdio instead: node dist/server/bin/stdio.js. Opening the repository in Claude Code gives you both the server (.mcp.json) and the Agent Skill (.claude/skills/carecompanion) — ask for "Eleanor's morning briefing" and watch the tools run; docs/skill-walkthrough.md has the full script.

Copy .env.example to .env to change the household timezone, seed behaviour, Bedrock settings or the demo token; every variable is documented there.

The MCP surface

Every tool returns text written to be spoken aloud (content[0].text) and typed structuredContent — hosts that compose their own reply (Alexa+ does) get clean data; hosts that read text get a sentence that already works. Invalid arguments never reach a handler; every handler is pure in-memory work (a few milliseconds — no LLM runs inside the server, as a voice host expects).

Tool

Who

Purpose

get_todays_plan

elder

Today's doses, what is next, anything missed, appointments, whether the elder has checked in.

log_dose

elder

Record a dose by whatever the elder calls it. The guard refuses duplicates/too-soon doses with requiresConfirmation; overrides need a reason and alert the caregiver.

skip_dose

elder

Record a deliberate skip of the next due dose; critical medications alert the caregiver.

daily_checkin

elder

Mood + symptoms in the elder's words; escalation bands (watch / urgent / emergency) notify caregivers and return emergencyGuidance to speak first.

call_for_help

elder

Alert every caregiver at once; return emergency guidance.

add_medication

caregiver

Always adds; returns interaction/allergy warnings with a disclaimer (informational, never medical advice).

caregiver_summary

caregiver

Adherence (7/30 d), today's doses, open alerts, check-in trend, appointments — the MCP App tool (_meta.ui.resourceUri).

resolve_alert

caregiver

Acknowledge / resolve with a note; idempotent audit trail.

Resource carecompanion://elder/{elderId}/adherence (JSON, listed per elder) · Prompt morning_briefing. Argument details: skills/carecompanion/references/tools.md.

Architecture

flowchart LR
  subgraph hosts[MCP hosts]
    A[Alexa+ / Inspector / basic-host / Claude Code]
  end
  subgraph server[CareCompanion server · Node 24]
    M[/mcp Streamable HTTP<br/>sessionful, spec 2025-11-25/]
    T[8 tools · 1 resource · 1 prompt]
    G[guardrails<br/>dose guard · interactions · escalation]
    S[(household store<br/>today-relative seed, tz-correct)]
    V[[ui:// dashboard resource<br/>single-file React view]]
    AG[/api/agent<br/>Bedrock Converse loop/]
    X[MCP client over loopback HTTP]
  end
  W[web/ simulated Alexa+<br/>Web Speech STT/TTS]
  B[(Amazon Bedrock<br/>Claude Haiku 4.5)]
  SK[skills/carecompanion<br/>SKILL.md]
  A -->|tools/call| M --> T --> G --> S
  T --> V
  W -->|POST utterance| AG --> B
  AG --> X -->|initialize · tools/call · DELETE| M
  SK -.teaches.-> A
  • Transport: the sessionful NodeStreamableHTTPServerTransport pattern (session map, GET stream, DELETE) — the SDK's stateless default answers GET/DELETE with 405, which voice hosts treat as a broken lifecycle.

  • Domain: instants are UTC, "today" and schedule slots are computed in the household timezone with DST-safe Intl math (src/domain/time.ts); dose slots are virtual, outcomes are recorded; missed doses and missing check-ins are swept lazily on read, so the server needs no timers.

  • Guardrails are pure functions with their own tests: doseGuard.ts (max daily / min interval / inactive → refuse with confirmation), interactions.ts (curated pair table + brand aliases + allergy classes), escalation.ts (phrase bands with word boundaries; help → everyone).

  • MCP App: caregiver_summary is registered with registerAppTool and the view with registerAppResource; the view talks to the server only through the host bridge (callServerTool) and ships with an empty CSP. It follows the Alexa+ design guide: in inline mode it renders a wider-than-tall summary block (stat tiles, next doses, open alerts) with a control that requests fullscreen for the full dashboard; tokens sit on one root scale (16 px body, 8/12 px radii, 48 px touch targets) and the host's theme variables win over the fallbacks.

Runtime MCP calls (how the agent uses the server)

The Alexa+ track requires the track technology to be called at runtime. The simulated Alexa+ brain does not shortcut into the domain layer: every turn opens a real MCP session over loopback HTTP, lists the tools, calls them, and terminates the session — the same wire protocol an external host uses, visible in the server log as session opened → mcp tools/call → session closed.

  • src/agent/mcpExecutor.ts — openMcpExecutor(): Client + StreamableHTTPClientTransport, listTools() → Bedrock tool specs (schemas sanitized), callTool() per tool use, terminateSession().

  • src/agent/bedrockBrain.ts — the Converse tool-use loop (all results of a round returned in one user turn, round cap, result truncation, emergency guidance harvested from structured results).

  • src/agent/ruleBrain.ts — the offline intent router that drives the same tools when Bedrock is unavailable.

  • src/agent/service.ts — brain selection, 5-minute Bedrock pause after access/credential failures, daily turn cap.

Demo

The simulated Alexa+ experience: the elder's Echo-Show-style device beside the family dashboard

The duplicate-dose guard refusing a second Lisinopril and asking for an explicit confirmation and reason

The dashboard rendered as an MCP App inside the ext-apps basic-host in inline mode, dark host theme

The same MCP App after the view requested fullscreen from the host

The elder's today card, rendered from the get_todays_plan result in the same view bundle

The dose-guard card: the refusal in the guard's own words, a required reason, and "Record anyway"

Figures are reproducible: npm run figures drives the local Chrome through the web app and basic-host (scripts/figures.mjs).

  • Live server + web experience: coming with the deployment (/mcp, /, /healthz).

  • Demo video: coming.

  • Script and recording checklist: docs/demo-script.md. The demo menu in the web app can re-seed the household and set the local time (POST /api/demo/reset, token-gated) so the "mid-morning" and "evening" beats are reproducible at any hour.

What is real: the MCP server, tools, guardrails, MCP App view, Agent Skill, Bedrock loop, tests. What is simulated: the Alexa+ device (a web app with browser speech), caregiver notifications (recorded on the alert and logged — no SMS/email is sent), and the household itself (synthetic data, resets on restart).

The first request after ~15 minutes idle on the free hosting tier takes up to a minute (cold start); the app is otherwise stateless per session, so MCP clients simply re-initialize after a restart (404 / -32001).

Try it on real Alexa (classic skill)

The hackathon organizers confirmed that the Alexa+ add-on toolkit and web simulator are not available to participants and suggested demoing "your MCP being called from" a classic Alexa Skill. CareCompanion ships one: POST /alexa is an Alexa Skills Kit endpoint (request signatures verified over the raw body) that turns intents back into natural utterances, runs them through the same Bedrock or rule-based loop as the web app, and answers with speech, a card, and an APL screen on devices with a display. Ten minutes of setup with a free Amazon developer account, no device needed: docs/alexa-skill.md. In the console's Test tab:

open care companion
what's my plan today
I took my lisinopril
yes record it anyway because the doctor told me to double it
I'm feeling a bit dizzy
I need help right now

Each turn appears in the server log as session opened → mcp tools/call … → session closed → alexa request.

Onboarding to Alexa+ (optional)

The server meets the Alexa+ MCP Toolkit requirements as documented: Streamable HTTP, MCP 2025-11-25, a remote URL, tools well inside the 500 ms round-trip budget, and visuals via MCP Apps. With a US developer account the deployed URL can be registered as a dev-stage add-on and tried in Amazon's web simulator:

alexa-ai configure                                   # Login with Amazon (browser)
alexa-ai new mcp --name "CareCompanion" --locale en-US --mcp-server-url "https://<your-host>/mcp"
alexa-ai deploy                                      # development stage → Add-on ID → web simulator

Amazon's Local Inspector (addon-local-inspector https://<your-host>/mcp, distributed through the Developer Console) renders the caregiver_summary MCP App in Small (≤10") and Large (≥11") device frames and writes a certification verdict; the view was built for exactly those breakpoints.

Service-level authentication (Alexa+ Tier 1)

The public demo is unauthenticated on purpose (synthetic household, no PHI). The toolkit's Tier-1 model — the add-on fetches a short-lived Bearer token with the OAuth 2.0 client-credentials grant and sends it on every MCP request — is built in and switched on with three variables:

MCP_AUTH=client_credentials MCP_CLIENT_ID=alexa-addon MCP_CLIENT_SECRET=<long random string> PUBLIC_URL=https://<host>

With that set: GET /.well-known/oauth-authorization-server and /.well-known/oauth-protected-resource/mcp describe the server (RFC 8414 / 9728); POST /oauth/token (HTTP Basic client auth, grant_type=client_credentials, optional resource=<PUBLIC_URL>/mcp) returns a token valid for up to an hour; /mcp answers 401 JSON without a WWW-Authenticate header to anything else, exactly as the Alexa+ checklist asks. Tokens are stateless HMAC envelopes, so restarts and multiple instances need no shared store. The simulated Alexa+ brain mints its own token for its loopback calls. User-level account linking (Tier 2) is out of scope for a synthetic household.

Why it matters

The case for the guardrails, with every figure fetched and quoted from CDC, JAMA, NEJM, AARP/NAC, Pew and Amazon sources: docs/EVIDENCE.md. In short: a third of 60-to-79-year-olds take five or more medications; adverse drug events send older adults to the emergency department at more than twice the rate of younger adults, with anticoagulants such as warfarin among the top culprits; a warfarin-plus-NSAID combination roughly doubles the odds of a gastrointestinal bleed; one in four older adults falls each year and blood-pressure medicines are a named risk factor; most family caregivers manage medications, and more than one in ten live an hour or more away. Amazon's own elder-care subscription for Alexa, Alexa Together, is "no longer available".

Safety & scope

CareCompanion is a coordination and reminder tool, not medical advice. Interaction warnings come from a small, curated, informational table and always carry a disclaimer; escalation errs toward alerting a human; nothing is overridden without an explicit confirmation and a reason from the person. Demo data is synthetic — no real patients, no PHI. A production deployment would add authentication (the SDK's OAuth middleware is ready to mount), a real datastore, and clinically maintained interaction data.

AWS / Bedrock setup (optional)

The conversational brain uses Amazon Bedrock (Claude Haiku 4.5 via the us. cross-region inference profile). Without it, POST /api/agent answers through the rule-based brain and reports brain: "rules".

  1. In the Bedrock console for us-east-1, open Model access → Anthropic and submit the one-time use case details form, then enable Claude Haiku 4.5. Until then Bedrock answers ResourceNotFoundException: Model use case details have not been submitted… and the app falls back to rules.

  2. Check access without spending anything:

    aws bedrock get-foundation-model-availability --model-id anthropic.claude-haiku-4-5-20251001-v1:0 --region us-east-1

    agreementAvailability.status must be AVAILABLE.

  3. Provide credentials with bedrock:InvokeModel (a local profile, or AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY on the host) and keep BEDROCK_REGION=us-east-1 — the region is pinned explicitly because a profile's default region may differ.

  4. npm run agent:smoke shows which brain answered each turn; /api/agent/status shows usage and caps.

Project layout

src/            server: config, logging, domain (model, time, store, seed, guardrails, actions), mcp (tools,
                resources, prompts, schemas), agent (Bedrock loop, rule brain, MCP executor), http (routes)
ui/             the dashboard MCP App view (React, single-file build → dist/ui/mcp-app.html); ui/src/components
                is the shared Dashboard used by both UIs
web/            the simulated Alexa+ web app (Vite + React + Tailwind; Web Speech API)
skills/         the Agent Skill (SKILL.md, references, install script)
test/           vitest suites (domain, MCP over the SDK client, agent, REST)
scripts/        curl lifecycle walk-through, agent smoke
docs/           friction log, product feedback, demo script, Devpost text

Development: npm run dev (server with reload) + npm run dev:web (Vite dev server proxying /api and /mcp); npm run typecheck, npm run lint, npm run format, npm run skill:validate.

Boot prints Warning: Server is binding to 0.0.0.0 without DNS rebinding protection. That is expected: host-header validation is applied to /mcp only, so platform health checks on /healthz keep working (friction log FL-03).

Status

  • M0 walking skeleton: sessionful Streamable HTTP, stdio, lifecycle tests, CI

  • M1 domain, guardrails, seeded household (145 tests)

  • M2 MCP surface: 8 tools · 1 resource · 1 prompt, MCP App registration

  • M3 Bedrock Converse brain over a loopback MCP client, rule-brain fallback, /api/agent

  • M4 simulated Alexa+ web app with browser voice

  • M5 dashboard MCP App view verified in basic-host (initialized, tool result delivered, auto-resize)

  • Repositioning after the competitive survey: guardrails first; elder screens (today card, dose guard); server-push alert notifications; Alexa+ inline/fullscreen display modes

  • Classic Alexa Skill front end · live conformance probe · Alexa+ add-on package · evidence document

  • M6 Agent Skill walk-through, docs, product feedback

  • M7 live deployment · M8 video · M9 submission

Feedback for the tool makers

License

MIT — see LICENSE.

Available Tools

8 tools
add_medicationAdd a medicationA

Caregiver adds a medication to the elder's schedule. The medication is always added; the result also carries informational interaction and allergy warnings against the elder's current medications — read every warning and the disclaimer to the caregiver. Not medical advice.

ParametersJSON Schema
NameRequiredDescriptionDefault
doseYese.g. "200 mg"
formNotablet, capsule, liquid, …
nameYesBrand or generic name, e.g. "Ibuprofen"
addedByNoCaregiver id making the change.
elderIdNoElder id. Omit for the household elder.
purposeNoPlain-language purpose: "pain", "blood pressure"
criticalNoMissing it should raise a critical alert. Default false.
graceMinNoMinutes after a slot before it counts as missed. Default 90.
instructionsNoe.g. "with food"
maxDailyDosesNoDefault: number of schedule times.
scheduleTimesYes
minIntervalMinNoMinimum minutes between doses. Default 240.

Output Schema

ParametersJSON Schema
NameRequiredDescription
alertIdYes
warningsYes
medicationYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses critical behavioral traits beyond the annotations: 'The medication is always added' signals non-idempotent, always-effective behavior, and it goes further by explaining that the result carries interaction/allergy warnings and instructing the agent to read them aloud. This is rich, actionable context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler. The primary action is front-loaded, followed by essential behavioral caveats. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter tool with an output schema and annotations, the description provides the missing context: side-effect certainty, warning-reading obligations, and medical disclaimer. An agent has enough to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 92%, so the schema already documents parameter meaning. The description adds no parameter-specific guidance beyond the general medication concept, which is appropriate but not above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('adds a medication') and the resource ('the elder's schedule'), and attributes it to a caregiver. It is distinct from sibling tools like log_dose or skip_dose, which handle existing doses rather than adding new medications.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes clear usage context: this is the tool for a caregiver to add a medication to the schedule. It does not explicitly name when-not-to-use or alternatives, but its role is self-evident against the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call_for_helpCall for helpA

The elder asked for help or said something alarming outside a check-in ("help", "I fell", "I can't breathe"). Immediately alerts ALL caregivers and returns emergency guidance to speak verbatim.

ParametersJSON Schema
NameRequiredDescriptionDefault
elderIdNoElder id. Omit for the household elder.
messageNoThe elder's words, if any.

Output Schema

ParametersJSON Schema
NameRequiredDescription
alertIdYes
notifiedYes
emergencyGuidanceYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a mutating operation (readOnlyHint=false), but the description adds key behavioral detail: it alerts ALL caregivers immediately and returns emergency guidance to speak verbatim. This goes beyond the annotation flags and helps an agent understand the impact and output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence that front-loads the trigger condition and clearly states the action and output. There is no redundant phrasing; every clause contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are covered. The description includes the trigger, the side effect (alerting all caregivers), and the output (emergency guidance). It does not mention prerequisites or reversibility, but for an emergency tool this is acceptable and the core information is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% coverage for both parameters (elderId and message), each with descriptive text. The description adds minimal extra semantic value—it provides example phrases for the message but does not explain parameter relationships or additional constraints beyond the schema. Baseline 3 is appropriate given the schema's completeness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('alerts') and resource ('ALL caregivers'), and defines the trigger conditions with concrete examples ('help', 'I fell', 'I can't breathe'). It clearly distinguishes from sibling tools like daily_checkin and resolve_alert by emphasizing the 'outside a check-in' context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear trigger condition ('the elder asked for help or said something alarming outside a check-in') which implicitly excludes regular check-in scenarios. It does not explicitly name alternatives or say 'use this instead of daily_checkin', but the context is strong enough for an agent to infer the appropriate use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

caregiver_summaryCaregiver summary and dashboardA
Read-onlyIdempotent

Caregiver-facing summary of the elder's week: medication adherence (7 or 30 days), today's doses, open alerts, the check-in trend and upcoming appointments. Hosts that support MCP Apps also render the live family dashboard.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoAdherence window. Default 7.
elderIdNoElder id. Omit for the household elder.

Output Schema

ParametersJSON Schema
NameRequiredDescription
todayYes
openAlertsYes
generatedAtYes
checkedInTodayYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is established. The description adds value by listing the summary's contents and disclosing that hosts supporting MCP Apps may render a live family dashboard. This conditional rendering behavior is useful context beyond the structured annotations, though error or availability behavior is not discussed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the tool's purpose and content in a dense list, and the second adds a meaningful conditional about dashboard rendering. It is not padded, though the 'MCP Apps' terminology could be slightly unclear without additional context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, idempotent tool with only two well-documented parameters and an output schema, the description covers the essential behavior and content. The main gap is the absence of explicit guidance on when to choose this tool over get_todays_plan, but the audience and weekly scope largely convey that. Overall it is sufficient for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented in the schema. The description echoes the 7/30-day adherence window but adds no new semantic meaning beyond the schema. With full schema coverage, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a caregiver-facing summary of the elder's week and enumerates the exact data areas: medication adherence, today's doses, open alerts, check-in trend, and appointments. This specificity distinguishes it from siblings like get_todays_plan, which is more narrowly about today's plan. The tool's role as a read-only dashboard is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The caregiver-facing and weekly-summary framing implies when this tool is appropriate, but it never explicitly states when to use it versus get_todays_plan or the alert/dose tools. No exclusions or alternative routing are provided. The dashboard rendering note adds some context but does not give clear selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

daily_checkinDaily check-inA

Record how the elder feels today: their mood plus any symptoms in their own words. Runs symptom escalation and alerts caregivers automatically for urgent or emergency symptoms. If the result contains emergencyGuidance, say it to the elder FIRST and verbatim.

ParametersJSON Schema
NameRequiredDescriptionDefault
moodYesHow the elder says they feel: "good", "okay", "not great", "awful"
notesNo
elderIdNoElder id. Omit for the household elder.
symptomsNoAnything they mention, verbatim: ["a bit dizzy", "slept badly"]

Output Schema

ParametersJSON Schema
NameRequiredDescription
flaggedYes
alertIdsYes
notifiedYesCaregiver names notified (simulated).
severityYes
checkInIdYes
emergencyGuidanceNoWhen present, say this to the elder FIRST and verbatim before anything else.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations all false (no readOnly, idempotent, or destructive hints), the description carries the behavioral burden. It discloses that the tool runs symptom escalation and alerts caregivers automatically, and it provides a critical instruction to relay emergencyGuidance verbatim. This is valuable context, though it doesn't cover every side effect like record creation or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each earning its place: the first defines the core action, the second notes the automatic alerting behavior, and the third gives the emergencyGuidance instruction. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation. However, the description misses guidance on the `notes` parameter and provides no usage boundaries against siblings like call_for_help. It is adequate for basic invocation but has clear gaps for full autonomous decision-making.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with the `notes` parameter undocumented. The description only reiterates 'in their own words' for symptoms, which the schema already states as 'verbatim.' It adds no new meaning for mood, elderId, or notes, so it minimally compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: it 'records how the elder feels today' including mood and symptoms. This distinguishes it from siblings like log_dose or call_for_help by focusing on daily well-being, even without naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (for a daily check-in) and mentions automatic escalation, but it does not explicitly state when to prefer this over alternatives like call_for_help, nor does it provide exclusions or alternative routing. Usage context is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_todays_planToday's planA
Read-onlyIdempotent

The elder's medication schedule for today — what's taken, what's next, any missed dose, upcoming appointments, and whether they've checked in. Call this first for any "what's my day" or "what do I take" question.

ParametersJSON Schema
NameRequiredDescriptionDefault
elderIdNoElder id. Omit for the household elder.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dosesYes
todayYes
nextUpYes
elderNameYes
localTimeYes
appointmentsYes
checkedInTodayYes
openAlertsCountYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the description does not need to restate those. It adds useful context about the 'for today' scope and the statuses returned, but it does not disclose additional behaviors such as data freshness, fallback when no plan exists, or how the optional elderId affects the response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two compact sentences that front-load the resource and its contents, then give an explicit invocation cue. There is no filler, no repetition of annotations, and no unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, read-only tool with a single fully documented optional parameter, a rich output schema, and strong annotations, the description is complete enough. It tells the agent what the tool provides and when to call it first, leaving no critical gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameters, and the elderId description ('Elder id. Omit for the household elder.') already provides the needed semantics. The tool description adds no parameter-specific meaning, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as 'the elder's medication schedule for today' and enumerates the included content: taken doses, next dose, missed dose, appointments, and check-in status. This is specific and actionable, though it does not explicitly differentiate from siblings like caregiver_summary or daily_checkin by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a direct usage directive: 'Call this first for any "what's my day" or "what do I take" question.' This clearly states when to use the tool, but it does not provide when-not-to-use guidance or name alternatives, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_doseLog a doseA

Record that the elder took a medication (name as spoken — brand names and "my blood pressure pill" work). A safety guard may REFUSE a duplicate or too-soon dose: when requiresConfirmation is true, do NOT retry silently — tell the elder what the guard said, ask them to confirm and say why, and only then call again with confirmOverride=true and overrideReason. If the name is ambiguous the result lists candidates; ask which one.

ParametersJSON Schema
NameRequiredDescriptionDefault
elderIdNoElder id. Omit for the household elder.
takenAtNoWhen it was taken (ISO 8601). Defaults to now.
medicationYesThe medication as the elder said it: "lisinopril", "my blood pressure pill", "Coumadin"
overrideReasonNoThe elder's own words for why the extra dose is needed. Required together with confirmOverride.
confirmOverrideNoSet true ONLY after the elder explicitly confirmed recording a dose the guard refused.

Output Schema

ParametersJSON Schema
NameRequiredDescription
reasonYes
statusYes
alertIdYes
verdictYes
recordedYes
candidatesNoPresent when the medication name was ambiguous.
medicationYes
lastTakenAtYes
nextAllowedAtYes
requiresConfirmationYestrue = do NOT retry silently; ask the elder to confirm and give a reason, then call again with confirmOverride.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (no read-only, no idempotency, no destruction hints), so the description carries the full burden of disclosing behavior. It transparently warns that a safety guard may refuse duplicate or too-soon doses, prohibits silent retries, requires explicit confirmation before using override flags, and describes ambiguity resolution via candidates. This is strong behavioral context beyond what any annotation or schema field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core action, and uses a clear, readable structure for the refusal and ambiguity scenarios. Every sentence earns its place: no filler, no restating of obvious schema fields, and no ambiguous pronouns. The contrast between 'do NOT retry silently' and 'then call again' makes the workflow easy to follow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (5 parameters, mutable state, safety guard, output schema), the description covers all the non-obvious context an agent needs: how to phrase medication input, how to handle guard refusals, when to use override fields, and what to do with ambiguous candidates. Return-value details for normal success are reasonably delegated to the output schema, and the safety-critical behaviors are fully disclosed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds practical meaning beyond the schema: medication should be recorded exactly as the elder spoke it ('my blood pressure pill' works), and the overrideReason should contain the elder's own words for why the extra dose is needed. It also ties confirmOverride/overrideReason to the refusal flow, which enriches the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Record that the elder took a medication' and immediately clarifies the input style as 'name as spoken', including brands and colloquial phrases. This clearly distinguishes it from siblings like skip_dose (which records omission) and add_medication (which adds a medication definition).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit procedural guidance for the two main edge cases: safety-guard refusal and ambiguous medication names. It states when to retry (only after elder confirmation with overrideReason), when to ask the elder, and how to handle candidates. It does not explicitly contrast with alternatives such as skip_dose, but the 'took a medication' framing makes the primary use case unmistakable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_alertAcknowledge or resolve an alertA
Idempotent

Caregiver acknowledges (seen, working on it) or resolves (done) an alert, with an optional resolution note for the audit trail. Safe to repeat.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
alertIdYes
resolutionNoWhat was done, for the audit trail.
caregiverIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
alertYes
changedYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds value beyond the annotations by clarifying that acknowledge means 'seen/working' and resolve means 'done', and by noting the optional resolution note is written to the audit trail. This is behavioral context not present in the annotations. The 'Safe to repeat' statement aligns with idempotentHint=true, so there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence. It front-loads the core purpose, then adds the optional note and the idempotency safety in a natural order. There is no redundancy or filler, and it is appropriately concise for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, which covers return values. The description explains the two actions, the audit trail note, and idempotency. It does not mention error conditions or prerequisites, but for a straightforward action tool with annotations and an output schema, it is reasonably complete. Minor gaps remain, such as what happens if the alert is already resolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 25% schema description coverage, the description compensates by explaining the meaning of the action parameter (acknowledge vs resolve) and the purpose of the resolution note. It does not explicitly define alertId or caregiverId, but those are self-explanatory identifiers. Overall it adds meaning to the most ambiguous parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pair: acknowledge or resolve an alert. It also explains the two action meanings (seen/working vs done) and the optional resolution note, which distinguishes it from sibling tools that handle medication, check-ins, or planning. Though it does not explicitly name an alternative, the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like call_for_help or daily_checkin. The description only explains what it does, not the conditions that select it. No exclusions, no context about appropriate scenarios, and no reference to other tools. This leaves an agent without routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

skip_doseSkip a doseA

Record that the elder is deliberately skipping the next due dose of a medication, with an optional reason. Skipping a critical medication alerts the caregiver.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
elderIdNoElder id. Omit for the household elder.
medicationYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
alertIdYes
verdictYes
recordedYes
candidatesNo
medicationYes
scheduledTimeYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Given readOnlyHint=false already signals a write, the description adds meaningful side-effect context: skipping a critical medication alerts the caregiver. It also scopes the action to a deliberate, next-dose skip, which helps avoid accidental or retroactive misuse. No contradiction with the annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver the core action, the optional reason, and a relevant side effect with no filler. The most decision-critical information is front-loaded, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists and the parameters are simple, so return values do not need description-level detail. However, the lack of explicit routing against log_dose and the vagueness around medication identification make the definition slightly incomplete for an agent choosing among these sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at only 33%, the description partially compensates by marking 'reason' as optional and grounding 'medication' as the next due dose. It does not specify how medication should be identified or address elderId beyond what the schema already says, leaving the compensation incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Record') and a clear resource ('skipping the next due dose of a medication'), while also noting the optional reason and caregiver alert. It distinguishes itself semantically from log_dose by emphasizing a deliberate skip, though it does not explicitly name or contrast the sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case—an intentional skip of the next due dose—but does not explicitly state when to prefer this over log_dose or when not to use it. There is no when-not-to-use guidance or alternative routing, so an agent must infer the boundary from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedadd_medication
    • First observedcall_for_help
    • First observedcaregiver_summary
    • First observeddaily_checkin
    • First observedget_todays_plan
    • First observedlog_dose
    • First observedresolve_alert
    • First observedskip_dose

TDQS

A4/5.0

Scored across 8 tools

Disambiguation4/5

Each tool maps to a distinct action (view plan, log/skip dose, check-in, call help, add med, resolve alert, summary), but get_todays_plan and caregiver_summary both surface daily doses and schedule, requiring the elder/caregiver framing to disambiguate.

Naming Consistency4/5

Mostly imperative snake_case verbs (log_dose, skip_dose, add_medication, resolve_alert), but get_todays_plan, daily_checkin, and caregiver_summary break the verb_noun pattern slightly.

Tool Count5/5

Eight tools cover the core elder-care workflows without bloat; each tool addresses a distinct need and no tool feels redundant.

Completeness4/5

Medication lifecycle is covered for daily use (plan, log, skip, add) and alert/check-in flows are complete, but there is no update/remove medication or appointment management, so minor gaps exist.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An AI-powered eldercare platform that enables AI agents to monitor passive sensors, generate personalized care plans, and access specialized healthcare knowledge bases. It provides tools for passive monitoring of senior activities, medical document OCR, and real-time alert management for caregivers.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables caregivers to log daily care activities and generate draft notification reports for the elderly, with automatic safety constraints to prevent AI-generated inaccuracies.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables AI-driven post-discharge patient monitoring and care coordination through tools for symptom triage, recovery tracking, exercise recommendation, and clinical reporting.
    -