Skip to main content
Glama
onurerguden
by onurerguden
README.md
# GymRap: a personal AI strength coach built on Health Connect, Cloudflare and MCP

GymRap reads everything my watch and phone know about my body, works out what last night and yesterday's training actually mean, and emails me a coach report a few minutes after I wake up. A second, shorter report arrives after every gym session.

It is a full pipeline, with each layer doing one job:

- **Android (Kotlin):** a bridge that reads Health Connect.
- **Cloudflare Workers + D1:** a backend that does all of the analysis with deterministic, tested code.
- **MCP server:** OAuth-protected, so ChatGPT can read the data and act as the coach.
- **Gmail API:** delivers HTML emails that stay readable in Gmail.

**The backend makes no LLM API calls:** the "AI" part is the ChatGPT subscription I already pay for, so the system costs **$0 per month** to run.

> All screenshots and test data in this repository are synthetic. No personal health data, account IDs or addresses are included.

<p align="center">
  <img src="assets/screenshots/daily-email.png" width="31%" alt="Daily report email (synthetic data)">
  <img src="assets/screenshots/workout-email.png" width="31%" alt="Post-workout report email (synthetic data)">
  <img src="assets/screenshots/workout-heart-volume.png" width="31%" alt="Heart rate zones and muscle-group volume (synthetic data)">
</p>

---

## What it does

**Every morning, triggered by waking up instead of a fixed time**

- The phone notices that the main sleep has ended, that I have taken at least 30 steps and that HRV has arrived. It tells the backend: "today is ready."
- The backend builds the day's analysis:
  - a readiness score
  - 7/30/90-day baselines for every metric
  - sleep debt and a bedtime target for tonight
  - intensity minutes and the weight trend
  - a review of yesterday's session
  - today's workout, with exact kg and rep targets
- A coach writes the interpretation and the report goes out by email. A signed web panel adds real charts and the full history.

**After every Hevy workout**

- Hevy writes the session to Health Connect, the phone sees it end, and the backend matches it with the Hevy sets.
- The backend scores the session: plan vs. actual, personal records, progress per exercise, the heart-rate curve and zones, recovery, and weekly sets per muscle group.
- The result is a PT-style debrief email.

**Every number is computed, not generated.** The LLM never invents data. It receives a structured brief and writes the interpretation, and its JSON is validated against a schema before anything is stored.

## Architecture

```mermaid
flowchart LR
  subgraph Phone["Android phone"]
    W[Watch → Health Connect] --> A[GymRap bridge<br/>Kotlin · WorkManager<br/>change feed + watchdog]
  end
  A -- "HTTPS, device token<br/>changed days only" --> API

  subgraph CF["Cloudflare (free plan)"]
    API[Worker API] --> D1[(D1 SQLite)]
    API --> ENG[Analysis engine<br/>readiness · baselines ·<br/>double progression]
    MCP[MCP server<br/>OAuth 2.1 + PKCE] --> ENG
    CRON[Cron every 10 min<br/>guaranteed fallback] --> ENG
  end

  GPT[ChatGPT<br/>scheduled tasks + chat] <-- "13 tools" --> MCP
  HEVY[Hevy app] --> GPT
  ENG --> GMAIL[Gmail API<br/>owner's own inbox only]
  CAL[Google Calendar<br/>read-only, optional] --> ENG
```

| Layer | What it does |
|---|---|
| `android/` | Reads 23 Health Connect record types, including sleep stages, HRV (RMSSD), resting HR, per-minute heart rate, steps, energy, SpO₂, weight and exercise sessions. Uses the Health Connect **change feed** (`getChanges`), so it only reads and uploads days that actually changed. Waking up and a finished Hevy session are uploaded immediately. Includes a full-history backfill and a data inventory that explains why a metric is missing. |
| `backend/src/store.ts`, `insights.ts`, `analysis.ts` | Ingest with per-day, per-metric ordering, so a stale upload can never overwrite fresh data. A failed read keeps the last good value, and a confirmed "no data" clears it. On top of that: the readiness score, baselines, percentiles, streaks, trends and personal patterns. |
| `backend/src/training.ts`, `postworkout.ts` | Training analytics: double-progression targets, warm-up ramps, rest times, e1RM and records, stall detection with resets, weekly sets per muscle group, plan-vs-actual and heart-rate zones. |
| `backend/src/mcp.ts` | 13 MCP tools with honest read/write annotations, plus a single dispatcher, `get_next_task`, for scheduled runs. |
| `backend/src/report.ts`, `ledger.ts`, `gmail.ts` | Delivery: a report ledger, Gmail sending, retries and the fallback report. |
| `backend/src/render*.ts`, `charts.ts` | Email and panel rendering. The charts are pure HTML tables, because Gmail strips SVG and images. |

## The readiness score

`gymrap-readiness-1` combines four components, each measured against your own previous 30 days (at least 7 samples):

| Component | Weight | Signal |
|---|---|---|
| HRV | 35% | z-score of ln(RMSSD) |
| Sleep | 30% | duration vs. need, 3-night debt, efficiency |
| Resting heart rate | 20% | inverse z-score |
| Training load | 15% | 7:28-day acute:chronic ratio plus yesterday's spike |

If a component is missing, its weight is redistributed and the confidence drops. The score maps to four bands that drive the plan: **push ≥80**, **normal ≥65**, **easy ≥50**, **recover**. It is transparent by design: every email explains how many points each component cost. It is not Fitbit's Daily Readiness (which no public API exposes) and it is not medical advice.

## Training engine

- **Double progression per exercise:** add reps until every set hits the top of the rep range, then add load and start again at the bottom. After three sessions without beating the previous three, the engine prescribes a reset instead of grinding.
- **Today's session:** taken from the latest workout of the same split ("Push", "Upper"…), with target kg/reps, warm-up sets (e.g. `30×8, 47.5×5, 62.5×3`) and rest times.
  - On an **easy** day: same load, one set fewer, RPE ≤ 7.
  - On a **recover** day: no heavy work.
- **Muscle-group volume:** weekly sets per muscle group. The primary muscle counts as 1 set, a secondary muscle as 0.5. Each group is compared with a 10–20 sets/week target band.
- **Post-workout scoring:** reaching the targets, progress against the previous session of the same split, effort quality from RPE, volume adequacy and plan completion.
- **Optional Google Calendar override:** the server reads today's events (read-only). An event named `Antrenman - Pull` overrides the weekly split. Of the other events it keeps only busy time ranges, never titles.

## MCP server

The Worker is a remote MCP server that any MCP client can connect to. It uses `@cloudflare/workers-oauth-provider`: OAuth 2.1 with PKCE S256 only, Google sign-in, and access restricted to a single owner account.

| Tool | Kind | Purpose |
|---|---|---|
| `get_next_task` | read | Single entry point for scheduled runs. Returns `wait`, `sync_hevy`, `write_coach`, `write_post_workout` or `done`. Takes a 15-minute lease, so parallel runs never write the same report twice. |
| `get_morning_brief` | read | The full daily context: scores, analysis, the plan, coach instructions and the output schema. |
| `save_coach_report` / `save_post_workout_report` | write | Store the coach JSON (schema-validated, idempotent per report key). The server delivers the email itself. |
| `calculate_workout_metrics` | write | Turns Hevy workouts into stored summaries: working sets, e1RM, volume. Notes are never stored. |
| `resend_daily_report` | write | "Mail me today's report again", with a fresh coach text. At most 3 per day. |
| `query_health_history`, `get_training_history`, `get_post_workout`, `get_data_inventory`, `get_report_context` | read | History at any granularity, lifts, the last session, and data coverage. |
| `update_profile`, `log_feedback` | write | Goals and weekly split; feedback like "too long" or "bench felt heavy", which the coach must follow. |

Every tool call is logged to D1 (`mcp_calls`): tool, outcome code, size and duration, never the arguments. `/admin` shows the day's pipeline: wake time → ready → task calls → report → last phone upload.

## Reliability engineering

Most of the work went here, not into the happy path.

- **Report ledger:** each report moves through `claimed → generated → sending → sent | uncertain | failed`.
  - A Gmail 5xx or a timeout marks the report `uncertain`, and an uncertain report is **never** re-sent automatically. A duplicate email is worse than a missing one.
  - Hard failures retry with back-off (15/30/45 min) using the stored coach text, then surface an error code to the owner.
- **Guaranteed delivery:** a cron job checks every 10 minutes. If no coach report has arrived 90 minutes after waking (or by a deadline when there is no wake signal), the server sends the full numbers-only report itself. If a coach run is writing at that moment, the cron waits at most 15 minutes. A coach text that arrives late is attached to the web panel and never triggers a second email.
- **Phone watchdog:** Samsung's memory manager kills idle apps. When the kill lands right after a WorkManager run, the periodic job can silently disappear until the app is opened again. A self-rearming `AlarmManager` alarm re-registers the job and runs a sync if no run has happened for 30 minutes. This was verified on the device by deleting the job and killing the process, with `adb` diagnostics in `scripts/phone-diagnose.sh`.
- **Email that survives Gmail:**
  - Gmail strips `<svg>`, images, flexbox, grids and CSS variables. Every chart is a nested HTML table.
  - Gmail clips messages over ~102 KB, so a size budget drops optional charts first and never drops text.
  - Tests enforce both rules.
- **Security:**
  - Device tokens are stored only as hashes.
  - The Google refresh token is encrypted with AES-GCM using an HKDF-derived key.
  - The panel uses signed links and a CSP nonce.
  - Admin forms use CSRF tokens.
  - The recipient address can never come from tool input; it is fixed at Gmail-connect time.
  - "Delete everything" wipes all tables and revokes every MCP grant.

## The interesting failure: ChatGPT's safety wall

The original plan was fully autonomous: ChatGPT scheduled tasks poll the MCP server every 20 minutes, notice that I am awake, write the coach report and save it.

**What happened:**
- Reads always worked.
- In scheduled (unattended) runs, ChatGPT blocked the `save_coach_report` write in its own safety layer **before the request ever reached the server**. The task replied with "blocked by safety checks".
- The same call from an interactive chat went through.

**How I narrowed it down:**
- The D1 call log proved the write never arrived: `write_coach` was dispatched and no save followed.
- ChatGPT's conversation metadata showed which model and configuration each run used.

**Hypotheses tested, one change at a time:**

| Hypothesis | Change | Result |
|---|---|---|
| A write right after reading Google Calendar looks like data exfiltration | Moved the Calendar read to the server, so the task touches no Google app | still blocked |
| Unattended runs use the fast model, not the thinking one | Pinned the task's model and effort | the setting did not change the resolved model; still blocked |
| Hevy data in the same run | A run with no Hevy step | still blocked |

**Score: every unattended coach save was blocked, across 6 runs over three days.** Other write tools passed in the same tasks, and every interactive save passed.

**Resolution: design for the constraint instead of fighting it.**
- The server guarantees a deep numbers-only report every day.
- The coach interpretation is one sentence away in chat: *"mail me today's report with the coach's comments"* calls `resend_daily_report`.
- A free-tier option is ready if needed: Cloudflare Workers AI (`gpt-oss-120b`) would cost roughly 10–15% of the daily free allowance per report.

The lesson: when an agent platform puts an opaque safety layer between your agent and your API, the critical path must not depend on that agent acting unattended.

## How it was built

GymRap went from a first working pipeline to its current version in a few intense days, one tested increment at a time:

- **Tests with every change:** each behavior change landed with a test. The suite grew to **137 backend tests**, plus Android unit tests.
- **Debugging with production signals, not guesses:**
  - production D1 queries to trace a missing report
  - `dumpsys jobscheduler` and `exit-info` over `adb` to prove the Samsung process kills
  - ChatGPT run metadata to pin down the safety-layer block
- **Docs that keep up:** migrations, data contracts and operations notes stay in sync with the code.

## Tech stack

| Area | Stack |
|---|---|
| Android | Kotlin, Jetpack Compose, Health Connect, WorkManager, AlarmManager, Android Keystore |
| Backend | TypeScript, Cloudflare Workers, D1 (SQLite), KV, Cron Triggers, Workers Logs |
| Protocol | Model Context Protocol (`@modelcontextprotocol/server`), OAuth 2.1 (`@cloudflare/workers-oauth-provider`), Zod schemas |
| Delivery | Gmail API (`gmail.send`), Google Calendar API (read-only, optional) |
| Testing | Vitest with `@cloudflare/vitest-pool-workers` (tests run inside the Workers runtime against a local D1), Android unit tests, Prettier, GitHub Actions |

## Repository layout

```
android/              Health Connect bridge (Compose UI, sync engine, watchdog)
backend/
  src/                Worker: API, MCP, analysis, rendering, delivery
  migrations/         D1 schema (10 migrations)
  test/               Vitest suites (synthetic fixtures only)
docs/                 Setup, operations and data contracts (Turkish)
workflows/            ChatGPT task / project instructions, Hevy backfill prompt,
                      and an early PDF report prototype
scripts/              D1 backup + verified restore, phone diagnostics, preview screenshots
```

## Running it yourself

You need Node 24, JDK 21 and Android SDK 36.

```sh
npm ci
npm run check        # TypeScript
npm test             # 137 backend tests against a local D1
cd android && ./gradlew :app:assembleDebug :app:testDebugUnitTest :app:lintDebug
```

To deploy your own instance:
1. Create a D1 database, a KV namespace and a Google OAuth client.
2. Replace the `YOUR_…` placeholders in `backend/wrangler.jsonc` and `android/app/src/main/res/values/strings.xml`.
3. Set `GOOGLE_CLIENT_SECRET` and `SESSION_SECRET` with `wrangler secret put`.
4. Run the migrations and deploy.
5. Pair the phone from `/admin`, then connect `https://<your-worker>/mcp` in ChatGPT.

Step-by-step notes are in [`docs/SETUP.tr.md`](docs/SETUP.tr.md) and [`docs/OPERATIONS.tr.md`](docs/OPERATIONS.tr.md) (in Turkish). This is a single-user system by design: one owner, one active device, the Europe/Istanbul timezone and kilograms.

## Privacy

- **Stored:** daily summaries only. No raw sensor streams, no GPS, no workout notes and no calendar titles. The one exception is per-minute average heart rate during a Hevy session (at most 240 values) for the post-workout chart.
- **Who can access it:** data lives in the owner's own Cloudflare account and is reachable only by the owner's Google account and the paired phone.
- **Sending:** the only recipient is the owner's own verified address.
- **Deleting:** `/admin` can delete everything at any time.

## License

[MIT](LICENSE)

---

### Türkçe özet

GymRap, saat ve telefondaki Health Connect verilerini okuyup her sabah uyandıktan sonra, her antrenmandan sonra da kişisel bir koç raporu e-postalayan bir sistem.

- **Mimari:** Kotlin Android köprüsü, Cloudflare Worker + D1 backend ve OAuth korumalı bir MCP sunucusu.
- **Analiz:** Bütün sayısal analiz (hazırlık skoru, 7/30/90 gün tabanları, double progression, kas grubu hacmi) test edilmiş deterministik kodla sunucuda yapılıyor. ChatGPT yalnız yorumu yazıyor.
- **Maliyet:** LLM API'si çağrılmadığı için aylık maliyet 0.
- **En öğretici kısım:** ChatGPT'nin güvenlik katmanı, zamanlanmış görevlerin raporu sunucuya kaydetmesini engelledi. Kök nedeni kayıtlarla teşhis ettim ve sistemi buna göre yeniden kurdum. Sunucu her gün sayısal raporu garanti ediyor; koç yorumu ise sohbette tek cümleyle geliyor.
- **Yapım:** Birkaç yoğun günde, her değişiklik testiyle birlikte geliştirildi.