Skip to main content
Glama
kirollosatef

google-health-mcp

by kirollosatef
README.md
# google-health-mcp

**Read-only MCP server for the [Google Health API](https://developers.google.com/health).**
Give any AI agent — Claude Code, Claude Desktop, Cursor — access to your own
Fitbit, Pixel Watch and Health Connect data, and get advice grounded in *your*
baselines instead of generic wellness copy.

[![CI](https://github.com/kirollosatef/google-health-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/kirollosatef/google-health-mcp/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-informational.svg)](LICENSE)
[![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/downloads/)

Runs locally. Your data and tokens never leave your machine — the only network
calls are to Google.

> The legacy **Fitbit Web API was turned down in September 2026**. The Google
> Health API is its replacement, and this server targets that new API.

---

## Why this exists

The obvious tool surface — `get_steps()`, `get_sleep()` — produces bad advice.
Hand an agent 1,440 raw heart-rate points and it says *"try to sleep more."*

So the composed tools here return **your value, your baseline, the delta, and the
z-score**, plus an explicit confidence level and a list of any missing inputs:

```jsonc
// health_readiness("2026-03-14")   — illustrative values
{
  "score": 34,
  "verdict": "rest",
  "guidance": "Rest. Multiple recovery signals are off baseline together.",
  "confidence": "medium",
  "why": [
    "HRV well below your baseline (z=-1.72)",
    "resting HR +3.2 bpm over baseline",
    "8.7h of sleep debt over the last week"
  ],
  "missing_inputs": ["spo2"],
  "signals": {
    "hrv": { "value": 23.2, "baseline_mean": 30.73, "baseline_n": 9,
             "delta": -7.53, "z_score": -1.72, "status": "well below baseline" }
  }
}
```

An agent that can see *"HRV 23 ms against your 30-day baseline of 31, on 8.7 h of
sleep debt"* gives specific, checkable advice. One handed a bare score invents a
rationale for it.

## Read-only by construction

Google Health splits read and write into **separate OAuth scopes**. This server
requests only `.readonly` scopes, so the token it holds is *incapable* of
altering your health record. That is enforced by Google's token checks — not by a
flag in this code that a bug could flip.

## Tools

| Tool | What it gives you |
|---|---|
| `health_daily_brief(date)` | **Start here.** One day — sleep, resting HR, HRV, SpO₂, breathing rate, skin temperature, steps, active zone minutes — each against your trailing baseline |
| `health_readiness(date)` | Train / hold / rest, with the reasoning, a confidence level, and which inputs were missing |
| `health_trend(data_type, weeks)` | Week-over-week means and direction for one metric |
| `health_anomalies(days, threshold)` | Only the days where something moved >2 SD from your own norm |
| `health_daily(data_type, start, end)` | Daily series for any of the 40 data types |
| `health_points(data_type, start, end)` | Raw / intraday escape hatch, cross-source reconciled |
| `health_profile()` | Profile, units, timezone, paired devices, battery, last sync |
| `health_data_types()` | Every readable data type, with the wearable-backed ones flagged |
| `health_auth_status()` | Token health, and the exact command to fix it |

Then just ask: *"how did I sleep this week?"*, *"am I recovered?"*, *"what's been
off lately?"*

---

## Setup

Roughly 15 minutes, **$0**, and no security review.

### 1. Prerequisite

Your tracker must be syncing into the Google account you are about to authorize.
Open the Fitbit app and confirm it is signed in with that account — otherwise
everything below will authorize cleanly and return an empty dataset.

### 2. Google Cloud Console (one time)

1. Create or pick a project at [console.cloud.google.com](https://console.cloud.google.com).
2. **Enable the API** — [Google Health API](https://console.developers.google.com/apis/library/health.googleapis.com) → *Enable*.
3. **Consent screen** — [Audience](https://console.developers.google.com/auth/audience):
   User type **External**, publishing status **Testing**, and add your own Google
   account under **Test users**.
4. **Scopes** — [Data Access](https://console.developers.google.com/auth/scopes) →
   *Add or remove scopes* → search "Google Health API" → select these six:

   ```
   .../auth/googlehealth.activity_and_fitness.readonly
   .../auth/googlehealth.health_metrics_and_measurements.readonly
   .../auth/googlehealth.sleep.readonly
   .../auth/googlehealth.nutrition.readonly
   .../auth/googlehealth.profile.readonly
   .../auth/googlehealth.settings.readonly
   ```

5. **Credentials** — [Credentials](https://console.developers.google.com/apis/credentials) →
   *Create credentials* → *OAuth client ID* → type **Web application** → add this
   authorized redirect URI **exactly**:

   ```
   http://localhost:8787/oauth/callback
   ```

   Copy the client ID and secret.

> **Why Testing mode?** Every Google Health scope is *Restricted*. Publishing an
> app that uses them requires an annual [CASA](https://appdefensealliance.dev/casa)
> security assessment — $500–$4,500 and 2–6 weeks. Testing mode skips all of it,
> supports up to 100 users, and costs nothing. The trade-off is that refresh
> tokens lapse periodically and you re-run one command.

### 3. Install

```bash
git clone https://github.com/kirollosatef/google-health-mcp.git
cd google-health-mcp

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -e .

cp .env.example .env             # Windows: copy .env.example .env
```

Put your client ID and secret in `.env`.

### 4. Authorize

```bash
python scripts/authorize.py      # or: google-health-authorize
```

A browser opens. Google will warn **"Google hasn't verified this app"** — expected
in Testing mode, not an error. Choose *Advanced → Go to (your app)* and approve.
The script captures the callback, stores an encrypted refresh token, and makes a
live call to prove the chain works, printing your paired devices.

### 5. Connect your agent

```bash
python scripts/mcp_config.py > .mcp.json
```

That writes a config with the paths on your machine already resolved. Claude Code
picks up `.mcp.json` when started from this directory; for Claude Desktop or
Cursor, merge the printed block into that app's MCP config.

```bash
python tests/test_smoke.py       # 51 offline checks, no network or credentials
```

---

## Layout

```
src/google_health_mcp/
  config.py      settings, scopes, endpoints
  auth.py        OAuth flow, encrypted token store, auto-refresh
  client.py      REST client + polymorphic data-point parsing
  datatypes.py   40 data types: filter fields, value paths, units
  analytics.py   baselines, z-scores, sleep debt, readiness   (pure, no I/O)
  server.py      MCP tool surface
  cli.py         authorize / config entry points
```

`analytics.py` and the parsing in `client.py` are pure functions with no network
or auth dependency, so they lift unchanged into a Worker or Lambda if you later
want webhook subscriptions (`projects.subscribers.subscriptions`) or phone access.

## The published docs are wrong in five places

Building this against the live API turned up five errors in Google's reference,
each of which returns a 400 or silently produces wrong numbers. They are written
up in **[docs/API-CORRECTIONS.md](docs/API-CORRECTIONS.md)** — worth reading
before you write any Google Health code of your own:

1. `range.start` is a `CivilDateTime`, not a `Date` — the reference's own example is rejected.
2. Numeric fields arrive as **JSON strings** (`"8432"`), so naive parsers drop nearly every value.
3. Filter fields must be prefixed with the data type, with a different suffix per family. The full grammar exists only in the discovery document.
4. `dailyRollUp` does **not** support the daily data types (sleep, resting HR, HRV, SpO₂).
5. `pageSize` on rollups is a *duration* cap, not a row cap.

Plus units that are easy to get 1000× wrong: weight is in **grams**, distance in
**millimetres**.

`docs/discovery.json` is the authoritative schema; refresh it with
`python scripts/fetch_discovery.py`.

## Known limits

- **Re-authorization.** Refresh tokens lapse periodically in Testing mode. Tools
  then return `not_authorized` with the fix; re-run `scripts/authorize.py`.
- **Baselines need history.** Under 3 prior days yields `no_baseline_yet`. The
  default window is 30 days, set by `HEALTH_BASELINE_DAYS`.
- **SpO₂ can be empty** for a while after setup; it is reported under
  `missing_inputs` rather than invented.
- **Polling, not webhooks.** Subscriptions need a public HTTPS endpoint.
- Data availability varies by device — a screenless tracker has no GPS, and some
  models have no ECG.

## Privacy

Tokens are encrypted at rest under your OS user directory, never in the repo.
Health data is fetched on demand and never written to disk. See
[SECURITY.md](SECURITY.md).

## Not medical advice

This reads consumer wearable data with consumer-grade accuracy and applies simple
statistics to it. Readiness scoring is a prompt for a conversation, not a clinical
instrument, and every output carries a disclaimer. Persistent or severe changes
are a reason to see a clinician, not to ask an agent.

## Contributing

Issues and PRs welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). Especially
useful: data types this does not parse well yet, and devices whose payloads
differ from what is assumed here.

## License

MIT — see [LICENSE](LICENSE).

TDQS

A4.1/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: auth check, data type listing, profile, daily multi-metric brief, weekly trend, readiness, anomaly scan, raw points, and daily aggregate. No overlap between tools; health_daily_brief and health_daily differ by scope (single day vs range) and metric count.

Naming Consistency5/5

All tools follow a consistent health_ prefix with descriptive snake_case suffixes (auth_status, data_types, profile, daily_brief, trend, readiness, anomalies, points, daily). The pattern is uniform and predictable.

Tool Count5/5

9 tools is a well-scoped number for a health data server. Each tool covers a distinct aspect of reading and analyzing health metrics without redundancy or bloat.

Completeness4/5

The tool surface covers authentication, data discovery, profile, daily snapshots, trends, readiness, anomalies, raw data, and daily aggregates. Minor gap: no direct way to get multiple metrics across a date range in one call, but this can be composed from existing tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues