apeiron-mcp
This server provides a unified MCP interface to retrieve and analyze client health data from the Apeiron backend, covering vitals, assessments, notes, and trends.
Get sleep data – Retrieve sleep stages, duration, efficiency, HRV during sleep, and sleep scores for a date range.
Get exercise data – Fetch workout logs (type, duration, calories, heart-rate zones, RPE) with optional activity-type filtering.
Get nutrition data – Access meals, macros, calories, hydration, and supplement information.
Get cardio metrics – Obtain resting HR, HRV, VO2max, blood pressure, and aerobic capacity, optionally filtered by metric.
Get fitness assessment – Pull periodic clinical measurements like bone density, body composition, balance, movement quality, and muscle strength.
Get cognitive data – Retrieve cognitive test scores (memory, reaction time, processing speed) with trends.
Get healthspan domain summary – Get cross-domain rolled-up healthspan scores.
Get lifestyle summary – View sleep/activity/nutrition adherence rollup for a period (week, month, quarter).
Get trends – Generic trend statistics (time series with deltas/direction) for any domain and metric, with configurable granularity.
Get notes – Fetch free-text clinical, coach, or client notes, optionally filtered by author.
Get chat history – Retrieve paginated coaching/chat conversation logs, with optional conversation ID and limit.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@apeiron-mcphow's my sleep trend been this month?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
apeiron-mcp
An MCP server exposing client health data from the Apeiron backend
(https://api.apeiron.life). Tool surface is intentionally consolidated by
data shape (time-series vitals, periodic assessments, unstructured text,
derived aggregates) rather than by raw domain, to keep LLM tool-selection
unambiguous.
Tools
# | Tool | Purpose | Apeiron API endpoint |
1 |
| Sleep analytics for a date window. |
|
2 |
| Workout analytics for a date window. |
|
3 |
| Nutrition score + per-question scores. |
|
4 |
| Cardio + aerobic (resting HR, HRV, VO2max, BP, aerobic capacity). |
|
5 |
| Body comp, bone density, balance, movement, muscle strength. |
|
6 |
| Cognitive assessment datasets and trends. |
|
7 |
| Cross-domain rolled-up healthy-signal scores. |
|
8 |
| Lifestyle/adherence rollup for a period. |
|
9 |
| Generic trend statistics for any (domain, metric). |
|
10 |
| Clinician/client/system free-text notes. |
|
11 |
| Coaching chat messages (most recent first). |
|
¹ Uses the …-day variant (GET /activity_feed/sleep-analytics-day,
GET /activity_feed/exercise-analytics-day) when the requested window is a
single day.
See Endpoint mapping for the parameters sent to each endpoint and the filters that are applied server-side vs. locally.
All tools return the envelope:
{
"client_id": "...",
"domain": "...",
"period": {"start": "...", "end": "..."},
"data": [],
"unit_system": "metric",
"last_synced_at": "...",
"source": "apeiron-api",
"notes": ["..."]
}sourceis"apeiron-api"for live data and"stub"when no credentials are configured.notesis only present when there is something to explain: a filter the server could not apply, a retried request, or an argument the API ignores.datacarries the API payload as returned by the backend — the server does not reshape it into an invented schema, so field names and units (weight_lbs,height_inch, …) are the backend's.client_idis the Apeironperson_idand must be numeric whenever credentials are configured.
Related MCP server: carechronicle-mcp
Install & run
pip install -e .
apeiron-mcp # stdio (default)
MCP_TRANSPORT=streamable-http apeiron-mcp # HTTP on http://127.0.0.1:8000/mcpThe transport is selected with MCP_TRANSPORT; see
Configuration for the bind address, port and other options.
Register with Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"apeiron": {
"command": "apeiron-mcp"
}
}
}Run with Docker
The image is host-agnostic: it defaults to the network-friendly
streamable-http transport so it can run as a long-lived service, and it can
also be driven over stdio by a local MCP client.
Build
docker build -t apeiron-mcp:0.1.0 .Run as an HTTP service (recommended for deployment)
# 8000 is taken by apeiron-ml's research-api on the shared EC2 -> publish on 8001.
docker run -d --name apeiron-mcp \
-p 8001:8000 \
-e APEIRON_ACCESS_TOKEN="$APEIRON_ACCESS_TOKEN" \
--restart unless-stopped \
apeiron-mcp:0.1.0The endpoint is then http://localhost:8001/mcp. Verify it:
curl -i -X POST http://localhost:8001/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1"}}}'Prefer a plain-JSON reply (easier to script against)? Add
-e MCP_JSON_RESPONSE=true.
Run over stdio (for a local MCP client)
-i keeps stdin attached and the image writes all logs to stderr, so
stdout stays a clean JSON-RPC stream:
docker run -i --rm -e MCP_TRANSPORT=stdio apeiron-mcp:0.1.0Point a client at it with:
{
"mcpServers": {
"apeiron": {
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT=stdio", "apeiron-mcp:0.1.0"],
"env": { "APEIRON_ACCESS_TOKEN": "..." }
}
}
}With Docker Compose
cp .env.example .env # then edit APEIRON_ACCESS_TOKEN
docker compose up -d --build
docker compose logs -fThe apeiron-mcp service is started by default; a stdio service lives behind a
profile (add -T when running from a script or CI, where no TTY is attached):
docker compose --profile stdio run --rm apeiron-mcp-stdioThe container runs as a non-root user (uid/gid 10001) with a read-only
root filesystem, all Linux capabilities dropped and a built-in HEALTHCHECK
(TCP probe on the configured port). Host port mapping is controlled by
MCP_HOST_PORT (host side, default 8001) and MCP_PORT (container side,
default 8000). MCP publishes on 8001 because apeiron-ml's research-api
already binds 8000 on the shared EC2 instance.
Configuration
Backend credentials:
Variable | Default | Purpose |
|
| Backend base URL. |
| (unset) | Bearer token for the API. |
| (unset) | Login email — used to mint a token when no bearer token is set. |
| (unset) | Login password. |
|
| Per-request timeout in seconds. |
With none of APEIRON_ACCESS_TOKEN / APEIRON_EMAIL+APEIRON_PASSWORD set, the
tools return labelled stub data ("source": "stub") instead of API results, so
the demos keep working offline.
Server behaviour:
Variable | Default | Purpose |
|
|
|
|
| HTTP bind address. |
|
| HTTP bind port. |
|
| Run without HTTP sessions (use behind a load balancer without sticky sessions). |
|
| Reply with plain JSON instead of SSE streams. |
| (unset) | Comma-separated |
| (unset) | Comma-separated |
|
| Compose only — host port published by |
Security note: binding to a non-loopback address (
0.0.0.0, required in a container) disables the SDK'sHost/Originvalidation so the service is reachable — the server logs a warning when this happens. When you expose the service beyond a trusted network, setMCP_ALLOWED_HOSTS(e.g.apeiron.example.com,apeiron.example.com:*— list both forms, because entries are matched against theHostheader including its port) and terminate TLS at a reverse proxy. The server itself ships no authentication layer, so do not publish the port directly to the internet.
CI/CD
Two GitHub Actions workflows mirror the apeiron-ml pipeline: they build the image in CI, push it to AWS ECR, and deploy a chosen tag to the staging/production host.
Workflow | Trigger | What it does |
push to any branch (or manual) | Builds | |
manual ( | Pulls the requested tag onto the target host and (re)starts the |
Image tags
Source branch | Image tag | Git tag |
|
| — |
| next semantic version, e.g. |
|
|
| — |
Every build also updates latest.
Required repository secrets
Secret | Used by | Purpose |
| build | AWS credentials for the ECR login. |
| build, deploy | Registry host, e.g. |
| deploy | SSH key for the target host. |
| deploy | EC2 hostname to deploy to. |
Deploying
Provide a
.env.dockerin the home directory of the target host (the workflow runsdocker run --env-file .env.docker). Start from the committed template and fill in the real values:cp .env.example ~/.env.docker # then edit APEIRON_ACCESS_TOKEN etc.See Configuration for every variable.
Run the Deploy MCP workflow and pick the image tag and the environment. MCP is published on host port
8001and the container always listens on8000;8000is left alone because apeiron-ml'sresearch-apiuses it.The workflow restarts the container and fails loudly (with the last 50 log lines) if it exits during startup.
The ECR_REPOSITORY (phs/apeiron-mcp), AWS_REGION (us-west-2),
CONTAINER_NAME (apeiron-mcp) and HOST_PORT (8001) values live in the
env: block at the top of each workflow, so a different registry, region,
container name or host port is a one-line change.
Exposing the service: the apeiron-ml EC2 security group only opens ports
8000and22, so8001is reachable on the host/VPC but not from the internet until an ingress rule for8001is added to the terraform security group (apeiron-ml/terraform_utils/research{,_staging}_env/ec2.tf). In the meantime reach it over an SSH/SSM tunnel or front it with a reverse proxy.
Authentication
Every read endpoint needs a bearer token; without one the API answers
401 {"message": "No authorization token found"}. Two ways to supply it:
Static token — set
APEIRON_ACCESS_TOKEN. Mint one with the login endpoint:curl -s https://api.apeiron.life/people/login \ -H 'Content-Type: application/json' \ -d '{"email":"you@example.com","password":"secret","app":"dashboard"}' # -> {"token": "...", "person": {...}}Email + password — set
APEIRON_EMAILandAPEIRON_PASSWORD; the client callsPOST /people/loginon first use, caches the token, and re-authenticates once if a call comes back401. To print a token forAPEIRON_ACCESS_TOKEN:python -c "from apeiron_mcp.apeiron_client import fetch_token; \ print(fetch_token('you@example.com', 'secret'))"
Endpoint mapping
Swagger spec: https://api.apeiron.life/_api.json (v13.0). The server only uses
read (GET) endpoints; the validated dataset types are bio_marker,
body_comp, cardiorespiratory, cognitive_health, lifestyle_assessment and
physical_assessment.
Tool | Request | Date handling |
|
| native (defaults to the last 7 days) |
| native, when | |
|
| native; |
|
| native |
|
| filtered locally on |
|
| filtered locally on |
|
| filtered locally on |
|
| native for healthy signals |
|
| native for the summary card |
|
| native; |
|
| native; |
|
| filtered and trimmed locally ( |
Filters applied locally
Some tool arguments have no counterpart in the API, so they are applied to the
response instead. When a filter matches nothing, the full response is returned
and an entry is added to the envelope's notes field, so results are never
silently emptied:
get_exercise_data.activity_type— matched against activity-name fields (activity_type_shown,name,workout, …).get_cardio_metrics.metric—resting_hr,hrv,vo2max,blood_pressureandaerobic_capacityare matched against metric/group names (sohrvalso matches "heart rate variability").get_notes.author— matched against author/role/source fields.Date bounds on
/datasets/{type}(which takes no date parameter) and on/message/.get_chat_history.limit— the most recent N messages are returned.get_trendsretries withoutmetric_keyswhen the API rejects"<domain>.<metric>", returning every trend metric for the window.get_chat_history.conversation_idis ignored: Apeiron keeps a single message thread per client.
Everything else (person_id, group, dates on range-aware endpoints) is sent to
the API as query parameters.
Development
pip install -e .
python examples/poc_walkthrough.py # all 11 tools, stub data
APEIRON_PERSON_ID=123 python examples/poc_walkthrough.py # same calls, live dataThe walkthrough prints the envelope's source and notes for every call, so it
doubles as a wiring check: without credentials everything reports
"source": "stub"; with credentials the same calls hit api.apeiron.life. Note
that the MCP stdio client only forwards a safe subset of the environment
(HOME, PATH, …), which is why the examples pass the APEIRON_* variables to
the server process explicitly.
LLM agent chat UI
examples/agent_demo.py puts a small browser chat UI in front of the server: an
LLM picks the MCP tools, every tool call is shown above the answer, and the reply
streams in as it is written. Start the server over streamable HTTP, set one LLM
key, then open http://127.0.0.1:8080:
MCP_TRANSPORT=streamable-http apeiron-mcp # terminal 1
.venv/bin/pip install anthropic openai
OPENAI_API_KEY=... .venv/bin/python examples/agent_demo.py # terminal 2 (or ANTHROPIC_API_KEY)Failures (a bad API key, the server not running) are reported inline in the page
instead of as a stack trace. MCP_URL overrides the MCP endpoint and
AGENT_HOST / AGENT_PORT the UI bind address.
API failures (bad token, unknown person_id, transport error) surface as MCP
tool errors carrying the failing method, path and HTTP status.
Available Tools
11 toolsget_cardio_metricsC
Cardio + aerobic metrics (same physiological domain).
Args: metric: If omitted, returns all supported metrics.
| Name | Required | Description | Default |
|---|---|---|---|
| metric | No | ||
| end_date | No | ||
| client_id | Yes | ||
| start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden, and it does not state whether this is read-only, whether client_id implies authorization scoping, what the default date window is, or whether data is paginated. The parenthetical about 'same physiological domain' hints at grouping but discloses no operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, with no filler sentences, but the brevity is under-specification rather than efficiency for a four-parameter tool. The parenthetical 'same physiological domain' is vague context rather than an actionable constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but the description leaves the required client_id and both date parameters unexplained and provides no behavioral notes for a data-retrieval tool with zero annotation coverage. An agent would be guessing about the time window and date format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, yet it only explains one of four parameters (metric and its omission default). start_date, end_date, and the required client_id are left entirely undefined, including formats and whether dates are inclusive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('Cardio + aerobic metrics') but is a bare noun phrase with no verb, so the agent infers retrieval rather than being told. Sibling tools are domain-named (get_sleep_data, get_nutrition_data), so the domain framing does implicitly differentiate, but the phrasing itself never states the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance given is that omitting metric returns all supported metrics, which is a parameter default rather than a when-to-use rule. There is no direction on when to call this versus get_fitness_assessment, get_healthspan_domain_summary, or get_trends, and no mention of the time-window parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_chat_historyC
Paginated coaching/chat conversation logs.
Args: conversation_id: Restrict to one conversation thread. limit: Max messages to return (default 50).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| end_date | No | ||
| client_id | Yes | ||
| start_date | No | ||
| conversation_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and does not discharge it. 'Paginated' is a useful trait, but there is no indication that access is scoped to a specific client_id, whether the logs are sensitive/audited, what permissions are required, or how paging beyond the first page is achieved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The prose line is front-loaded and tight, but the pseudo-args block only restates two schema fields and leaves the rest out, so the structure spends words without covering the surface it implies it covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the return shape need not be described, but for a 5-parameter tool with no annotations, three undocumented parameters (including the required one) and no usage routing make the definition insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it only documents 2 of 5 parameters (conversation_id, limit). The required client_id and the start_date/end_date date-range filters — the parameters most likely to be misused — are completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource — coaching/chat conversation logs — and adds the pagination characteristic, which clearly distinguishes it from the health-metric siblings (sleep, exercise, nutrition, trends). It lacks an explicit verb, but 'get_chat_history' plus 'conversation logs' leaves no ambiguity about what is returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. An agent gets no help deciding between this and the many other data-retrieval siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cognitive_dataC
Cognitive test scores (memory, reaction time, processing speed) with trends.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| client_id | Yes | ||
| start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not state that this is a read-only fetch, whether dates are inclusive/exclusive, what defaults apply when start_date/end_date are null, or how trends are computed. Only the return content (scores) is hinted at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded fragment with no wasted words, which is good. However, it is under-specified rather than genuinely concise for a three-parameter data tool, so it doesn't fully earn its brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the tool is a simple domain fetcher. Still, with no annotations and 0% parameter coverage, the description leaves date-range behavior and usage context unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters, so the description is the only possible source of meaning, yet it says nothing about client_id, start_date, or end_date. The names are largely self-explanatory, but date-range semantics and the null-default behavior are undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (cognitive test scores) and enumerates the specific sub-metrics returned (memory, reaction time, processing speed) plus trends. That clearly separates it from siblings like get_sleep_data or get_nutrition_data, though the verb is only implied by the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as get_trends, get_healthspan_domain_summary, or the other domain getters. No prerequisites, no mention of which client/date scenarios warrant it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exercise_dataB
Workout logs — type, duration, calories, heart-rate zones, RPE.
Args:
client_id: Client identifier.
start_date: ISO date, inclusive.
end_date: ISO date, inclusive.
activity_type: Optional filter (e.g. "run", "strength").
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| client_id | Yes | ||
| start_date | No | ||
| activity_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the returned data shape (type, duration, calories, HR zones, RPE), which is genuinely useful, but says nothing about read-only semantics, authentication/permission requirements, pagination, or what happens when start_date/end_date are omitted (both default to null).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource and its payload, followed by a compact Args block; nearly every line earns its place. Minor redundancy in restating field names already implied by the parameter list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description's parameter coverage is its main contribution and it omits which non-required params have null defaults and how a missing date range behaves. Adequate but with a visible gap for a 4-parameter, zero-coverage schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it documents all four parameters inline, marks activity_type as optional with concrete examples ('run', 'strength'), and specifies that start_date/end_date are ISO dates and inclusive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (workout logs) and enumerates the concrete fields returned — type, duration, calories, heart-rate zones, RPE — so an agent knows what data lands in the response. It does not, however, differentiate itself from plausible siblings such as get_cardio_metrics or get_fitness_assessment, which may surface overlapping metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: no statement of when to call this instead of get_cardio_metrics or get_fitness_assessment, no prerequisites, and no exclusions. The activity_type filter is described as 'Optional filter' but the description never explains the scenario that calls for it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fitness_assessmentB
Periodic clinical/assessment-style readings.
Covers bone density, body composition, balance, movement quality, and muscle strength — periodic (not continuous) measurements.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| client_id | Yes | ||
| start_date | No | ||
| assessment_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden; it discloses that data is periodic rather than continuous, which is genuinely useful context beyond structured fields. However, it says nothing about permissions, per-assessment granularity, units, or how results are returned, so the read-side behavioral profile is only partially covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short lines with no filler, and the domain list is front-loaded. Slightly terse for a schema with four undocumented parameters, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and this is a read tool with low risk. Still, the description omits that a client_id is mandatory and how the date-range parameters behave, which an agent needs in order to form a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 4 parameters, so the description must compensate — and it partially does by enumerating the five assessment types that correspond to the assessment_type enum. It says nothing about client_id (the sole required param) or how start_date/end_date scope the readings, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely (periodic clinical/assessment-style readings) and enumerates the covered domains — bone density, body composition, balance, movement quality, muscle strength — which maps directly onto the assessment_type enum. It also implicitly separates itself from continuous-metric siblings, though it never names them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(not continuous) measurements' gives an implied routing rule against continuous siblings like get_cardio_metrics or get_exercise_data, but no alternative tool is named and no when-not-to-use condition is stated. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_healthspan_domain_summaryC
Cross-domain rolled-up healthspan scores.
| Name | Required | Description | Default |
|---|---|---|---|
| client_id | Yes | ||
| as_of_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden, and it discloses almost nothing. 'Get' implies read-only retrieval and 'rolled-up' hints at aggregation, but there is no mention of permissions, whether a client_id is required, how as_of_date affects results, or freshness/rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single fragment rather than a sentence, and while short, it is under-specified rather than concise. There is no front-loaded action and no supporting detail, so brevity here reflects missing information rather than efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a two-parameter tool with zero annotation coverage and 0% schema description coverage the description is too thin. A consumer cannot tell whether this is a read-only aggregate, what client_id must be, or how as_of_date scopes the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across two parameters (client_id, as_of_date), and the description adds no meaning for either. It does not say what client_id refers to, whether as_of_date defaults to today, or what date format is expected, leaving both parameters fully undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Cross-domain rolled-up healthspan scores' names the resource and its aggregate scope, which does distinguish it from per-domain siblings like get_sleep_data and get_cognitive_data. However, it is a noun phrase with no verb, so the action (retrieve/read) is only inferred from the tool name. Purpose is understandable but not stated as a clear verb+resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus get_lifestyle_summary, get_trends, or the domain-specific getters. The only implicit signal is 'cross-domain', which hints at aggregation, but no conditions, prerequisites, or alternative choices are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_lifestyle_summaryB
Sleep/activity/nutrition adherence rollup for a period.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | week | |
| client_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It usefully signals that this is a computed 'adherence rollup' rather than raw data, which is meaningful context, but it does not disclose read-only status, permission requirements, or the meaning of 'adherence'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler words. It is efficient, though its terseness borders on under-specification rather than optimal density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and there are only two simple parameters. However, with no annotations and no parameter descriptions, the description leaves the client_id requirement and the notion of adherence undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, yet it only vaguely references 'a period' without tying it to the week/month/quarter enum or explaining the required client_id. The enum itself already conveys the allowed period values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: a multi-domain adherence rollup over sleep, activity, and nutrition for a given period. This distinguishes it from the single-domain siblings like get_sleep_data and get_nutrition_data, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this instead of get_sleep_data, get_nutrition_data, or get_healthspan_domain_summary. The phrase 'for a period' hints at the temporal scoping but offers no actionable selection guidance or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_notesD
Free-text clinical/coach/client notes.
| Name | Required | Description | Default |
|---|---|---|---|
| author | No | ||
| end_date | No | ||
| client_id | Yes | ||
| start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet 'free-text' is the only behavioral hint. It does not disclose that this is a read operation, whether results are paginated, what date-range semantics apply, or whose notes are returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but the brevity reflects under-specification rather than disciplined conciseness. A single noun fragment cannot front-load intent or scope for a tool with four parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the tool has four parameters at 0% coverage and no annotations. The description supplies none of the missing filtering, ordering, or safety context, leaving it wholly inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters (author, start_date, end_date, client_id), and the description names none of them. It adds zero meaning beyond the raw schema, so the author enum and date-range filtering are unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun fragment that essentially restates the tool name ('notes') and adds only a content-type qualifier (free-text, clinical/coach/client). It states no verb and no scope, and gives no basis for distinguishing it from the sibling get_chat_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the obvious alternative (get_chat_history) for retrieving conversational text. The agent is left to guess entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_nutrition_dataC
Meals logged, macros, calories, hydration, supplements.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| client_id | Yes | ||
| start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it says nothing about read-only vs mutating behavior, authorization requirements, pagination, or how much data is returned. The content list is useful scoping but not behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Seven words with no structure or front-loading; this is under-specification rather than conciseness, since the terse fragment leaves the core operation unstated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a parameterized, annotated-less query tool the description still omits the required parameter, the date-range semantics, and any usage context, leaving it well short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters (client_id, start_date, end_date) and the description adds no meaning: it never mentions the required client_id or that start/end dates bound the query, so it fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name gives a specific verb+resource (get_nutrition_data) and the description enumerates the content domains covered (meals, macros, calories, hydration, supplements), which implicitly separates it from the sleep/exercise/cognitive siblings. However, the description itself is a bare noun phrase with no verb, so the actual purpose is only inferred from the tool name rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when to prefer this over get_lifestyle_summary or get_trends, and no indication of prerequisites or date-range expectations despite the tool’s optional start/end date parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sleep_dataB
Sleep stages, duration, efficiency, HRV during sleep, sleep score.
Args: client_id: Client identifier. start_date: ISO date (YYYY-MM-DD), inclusive. Defaults to 7 days ago. end_date: ISO date (YYYY-MM-DD), inclusive. Defaults to today.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| client_id | Yes | ||
| start_date | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses default windows (start defaults to 7 days ago, end defaults to today) and that bounds are inclusive, which is genuine behavioral context. It says nothing about permissions, rate limits, or behavior on missing data, leaving meaningful gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is short and front-loads the metric scope before the Args block, with no filler sentences. The metric enumeration partly duplicates the output schema, but it is compact enough not to be wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the returned metrics need not be explained, and the parameter defaults are covered. But for a tool with zero annotation coverage and zero schema descriptions, the description omits behavioral essentials such as error/auth behavior and what happens when the range exceeds available data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it documents all three parameters, gives the ISO date format and inclusivity for both date bounds, and states their defaults. The only weak spot is client_id, described merely as 'Client identifier' with no format or sourcing detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource domain (sleep stages, duration, efficiency, HRV, sleep score), which makes it clearly distinguishable from the domain-partitioned siblings like get_cognitive_data or get_nutrition_data. However, it lists returned metrics rather than stating the retrieval action explicitly, so the verb is only implied by the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no conditions, and no mention of alternatives among the many sibling data tools. Usage for sleep analysis is only inferable from the metric list, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trendsC
Generic trend tool — time series with deltas/direction for any metric.
Args:
domain: e.g. "sleep", "cardio", "nutrition".
metric: metric name within that domain (e.g. "hrv", "sleep_score").
granularity: bucket size for the returned series.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | ||
| metric | Yes | ||
| end_date | Yes | ||
| client_id | Yes | ||
| start_date | Yes | ||
| granularity | No | day |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It notes the output contains deltas and direction, which is useful, but says nothing about permissions, whether it is read-only, rate limits, or how client_id scoping behaves. That is a large gap for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose statement is front-loaded and the Args block is compact, with no redundant prose. The 'Args:' framing is somewhat boilerplate but each line carries real content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be re-explained, but the tool has five required parameters and three of them (client_id, start_date, end_date) are undocumented in both schema and description. Combined with absent usage guidance and no annotations, the definition is under-specified for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It documents three of six parameters with concrete examples ('sleep'/'cardio', 'hrv'/'sleep_score') and explains granularity as bucket size, but leaves client_id, start_date, and end_date completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('time series with deltas/direction') and distinguishes itself as the 'generic' tool that works 'for any metric'. However, it never names the domain-specific siblings (get_sleep_data, get_cardio_metrics, etc.) that an agent must choose against, so the differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no conditions, and no mention of alternatives. The word 'generic' hints that this is the catch-all path versus the domain-specific siblings, but that inference is left entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
get_cardio_metrics - First observed
get_chat_history - First observed
get_cognitive_data - First observed
get_exercise_data - First observed
get_fitness_assessment - First observed
get_healthspan_domain_summary - First observed
get_lifestyle_summary - First observed
get_notes - First observed
get_nutrition_data - First observed
get_sleep_data - First observed
get_trends
TDQS
Scored across 11 tools
Most tools target distinct data domains (cognitive, sleep, exercise, nutrition, cardio, fitness assessment). Potential confusion exists between the two summary tools and between domain-specific get_*_data and get_trends, but descriptions clarify raw versus derived data.
All 11 tools use consistent snake_case with the get_ verb prefix followed by a descriptive noun phrase. The pattern is predictable and readable throughout, with only minor length variations.
11 tools is well-scoped for a read-only health data aggregation server. Each tool maps to a distinct data source or summary type, with no obviously redundant tools.
The surface covers core health domains (cognitive, sleep, exercise, nutrition, cardio, fitness assessments) plus summaries, trends, notes, and chat. Minor gaps include no client listing or metadata tool, but for a read-only aggregation API the coverage is strong.
Maintenance
Related MCP Connectors
Read wearables and lab health data — sleep, activity, workouts, timeseries, lab tests and orders.
- freddyOAuthcoach.freddy
Connect your wearables, rings and training apps, then ask your AI about your own health data.
Connect your health, fitness, nutrition, sleep, and wearable data to your AI assistant.
Gateway between LLM agents and world data through eight tools and a bundled endpoint catalog.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceExposes personal health data (recovery, sleep, strain, etc.) as MCP tools for AI agents to query and analyze.3MIT
- FlicenseNot gradedqualityBmaintenanceEnables LLMs to interact with clinical patient records using tools for document ingestion, structured conversion, patient profiling, record listing, search, and secure Q&A over patient documentation.-
- AlicenseNot gradedqualityBmaintenanceProvides MCP tools to read and update a local-first personal OS for goals, tasks, habits, food, workouts, and check-ins. Enables assistants to manage daily life data and interact with an evidence-grounded AI coach.MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to query health metrics such as activity, blood pressure, glucose, heart rate, sleep, and SpO2 from a PostgreSQL-backed wellness database, returning structured data and summaries through MCP tools.-