Skip to main content
Glama
rsjb195
by rsjb195
README.md
# care-agent

A personal care-management agent, medical module first. Built as a reusable
MCP + LangGraph skeleton — the medical domain is the first instance, not the
only one. Finance, licensing, or other domains can reuse the same graph
shape later with different tools and a different retrieval source plugged
in, without touching the graph itself.

Deliberately a **separate repo from J-Bishop-PA**, not a merge into it.
J-Bishop-PA sprawled across four phases and became hard to track. This repo
is scoped to one domain on purpose. MCP's whole design point is that a
client can connect to several servers at once — there's no need to share a
codebase to use both alongside each other.

## Why these choices

**LangGraph over a linear pipeline.** The domain has two real branches:
does this query need literature research or just local history, and does
the answer require an external action that needs sign-off first. A fixed
chain can't express that; a state graph with conditional edges can.
`interrupt()` plus a checkpointer gives human-in-the-loop for free instead
of hand-rolling a pause/resume mechanism.

**Chroma, embedded, local default embeddings.** The corpus is small — tens
to low hundreds of abstracts, not millions of documents — and cost matters.
No hosted vector DB to stand up or pay for, no per-call embeddings API
charge for a static, small corpus. Revisit this if the corpus grows past a
few hundred documents or retrieval quality turns out to need more than a
small local model can give.

**PubMed plus a short allowlist, not general web search.** RAG output is
only as trustworthy as what's indexed. General web search pulls in forums
and content-farm health content right alongside real research — the source
restriction is a correctness decision, not a shortcut.

**No send-capable tool anywhere in this server.** Every provider-facing
action is draft-only (`draft_provider_email`), mirroring J-Bishop-PA's
`draft_reply` pattern. Enforced by `tests/test_no_send_tool.py`, not just a
convention in a docstring — see the note below on why that distinction
matters.

**Synthesis never states a diagnostic or causal conclusion.** The model has
no clinical grounding to draw that line itself. Output is always framed as
"here's what's published, here's what to ask your doctor," never a
standalone conclusion — enforced in the system prompt for the synthesize
node, checked again by the guardrail node.

**Same model as judge, not a separate one.** Cheaper and simpler at this
scale — a second Claude call with a distinct system prompt. Known
limitation, worth being upfront about: this is weaker separation than a
genuinely independent judge model or human review. Fine here; wouldn't be
at higher stakes.

**On the guardrail-as-code test:** an earlier project (Baysa Analytics) had
a leakage-prevention test that looked real but mostly asserted on database
rows rather than actual function output — it would pass even if the thing
it was meant to catch happened. `test_no_send_tool.py` here is written
against the actual registered tool list for that reason, not against
something adjacent to it.

## Status

Scaffolded, not built. Everything in `graph/nodes.py`, `tools/*.py`, and
`rag/*.py` is a docstring and a `NotImplementedError` describing intent —
see the build plan for what gets implemented in what order.

## Setup

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env  # fill in ANTHROPIC_API_KEY and PUBMED_EMAIL
```

Maintenance

ActivitySlowing
ResponsivenessNo issues