PersonalDB Memory MCP
by vitorcalvi
README.md
# PersonalDB Memory MCP
Production-oriented private memory and hybrid retrieval for **ChatGPT**, **Ava Mobile**, and **My Thoughs**, built on Cloudflare Workers + D1 FTS5 + Vectorize + weighted Reciprocal Rank Fusion (RRF).
## Product goal
End-user setup is intentionally minimal:
```text
Install / Set up PersonalDB Memory
↓
Sign in with Cloudflare
↓
Allow access
↓
Done
```
There are **no PersonalDB API keys, no GitHub login, no OAuth client secrets, and no end-user configuration**.
A Cloudflare account is the user identity. Cloudflare Access Managed OAuth handles OAuth 2.0/2.1 for ChatGPT and other non-browser clients, and the Worker reads the authenticated identity from `ctx.access`.
## Architecture
```text
ChatGPT / mobile OAuth client
│
▼
Cloudflare Access Managed OAuth
Cloudflare identity provider
│
▼
PersonalDB Worker
├─ ctx.access.getIdentity()
├─ opaque PersonalDB user id derived from Access user_uuid
├─ D1 records/chunks/tombstones (canonical)
├─ D1 FTS5 BM25 (lexical)
└─ Vectorize (semantic; namespace per PersonalDB user)
│
└─ parallel retrieval → weighted RRF → top_k
```
The cloud never requires embedding inference. Mobile clients may generate embeddings locally. The default Vectorize index is 384 dimensions; all clients writing vectors to one index must use the same embedding model/dimension.
## Authentication and isolation
Cloudflare Access is the only production authentication boundary.
- ChatGPT and other MCP clients authenticate through **Access Managed OAuth**.
- Mobile/REST clients use the same OAuth flow; they do not receive a PersonalDB API key.
- The Worker fails closed when `ctx.access` is absent.
- The Access `user_uuid` is hashed into an opaque `usr_...` PersonalDB tenant id; email is not used as the database key.
- Request bodies may never supply `user_id`.
- Every D1 query is scoped by the authenticated PersonalDB user id.
- Every Vectorize query uses a per-user namespace and results are joined back through D1 as a second isolation boundary.
## MCP tools
Personal memory:
- `memory_add`
- `memory_get`
- `memory_search`
- `memory_list`
- `memory_update`
- `memory_delete`
- `memory_health`
Business knowledge:
- `knowledge_search`
- `knowledge_list`
Lower-level:
- `personaldb_search`
- `personaldb_sync`
## REST contract
The same Cloudflare Access identity protects the REST endpoints:
- `POST /v1/records/upsert`
- `GET /v1/records/:id`
- `DELETE /v1/records/:id`
- `POST /v1/search`
- `GET /v1/sync?cursor=...`
- `POST /v1/sync/ack`
Writes are versioned/idempotent. Deletion writes a tombstone and the same id cannot be resurrected by a later sync. Sync is cursor-based and incremental; whole SQLite files are never uploaded.
## Local development
Local development simulates a Cloudflare Access user through `wrangler.jsonc -> access.dev`.
```bash
npm install
npm run check
npm run local:smoke
npm run dev
```
`local:smoke` verifies authenticated MCP startup, cross-user data isolation using two simulated Access identities, incremental sync, FTS fallback, deletion/tombstone behavior, and no-resurrection semantics.
## Deploy
Authenticate Wrangler once as the product owner:
```bash
npx wrangler login
npm run check
bash scripts/provision.sh
```
The provisioning script creates or reuses D1 and Vectorize, applies migrations, and deploys the Worker. It does **not** create API keys or application OAuth secrets.
### One-time Cloudflare Access setup
After deployment, configure the Worker/hostname in Cloudflare Zero Trust:
1. Protect the Worker or production hostname with **Cloudflare Access**.
2. Select **Cloudflare** as the identity provider.
3. For a public end-user product, turn **off** `Restrict to account members` so users can sign in with their own Cloudflare accounts rather than needing membership in your Cloudflare account.
4. Add an **Allow / Everyone** policy for authenticated users.
5. Enable **Managed OAuth** on the Access application.
This is product-owner infrastructure setup, not end-user configuration. See [`docs/cloudflare-access.md`](docs/cloudflare-access.md).
## Connect ChatGPT
The production MCP URL is simply:
```text
https://<your-production-host>/mcp
```
ChatGPT follows the OAuth challenge exposed by Cloudflare Access Managed OAuth, opens Cloudflare sign-in, and reconnects with the issued OAuth token. PersonalDB itself does not run an OAuth authorization server.
See [`docs/chatgpt.md`](docs/chatgpt.md).
## Mobile sync guidance
Keep existing SQLite local-first storage intact:
1. Local transaction commits first.
2. Enqueue changed record id/version + on-device embedding.
3. Upload incremental changed records only.
4. Save server cursor per device.
5. Pull `/v1/sync` from the cursor and apply newer versions/tombstones locally.
Mobile clients authenticate to the same Access-protected origin using OAuth; never ship a shared PersonalDB secret in the app.
## Benchmark
```bash
PERSONALDB_URL=https://... \
PERSONALDB_ACCESS_TOKEN='<oauth-access-token>' \
BENCH_QUERIES='[{"query":"refund policy","expected":["record-id"]}]' \
npm run benchmark
```
`PERSONALDB_ACCESS_TOKEN` is an OAuth access token for operator testing, not a long-lived API key. It is unnecessary when benchmarking a local `access.dev` instance.
## Security
See [`SECURITY.md`](SECURITY.md). PersonalDB does not log API keys because it does not issue them. Do not log Access assertions, OAuth bearer tokens, raw private memory, or vectors.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues