Skip to main content
Glama
README.md
<div align="center">

# Todoie

**The to-do list that does the filing.**

Text it like a friend — an AI agent files your tasks, catches duplicates,
and answers questions about your list with citations.

[**Try the live demo →**](https://todoie.net) *(no signup — every visitor gets an isolated sandbox)*

[![tests](https://github.com/DavidStor/todoie/actions/workflows/ci.yml/badge.svg)](https://github.com/DavidStor/todoie/actions/workflows/ci.yml)
![routing eval](https://img.shields.io/badge/routing_eval-77_labeled_cases-58a6ff)
![injection safety](https://img.shields.io/badge/prompt--injection_safety-12%2F12-3fb950)
![python](https://img.shields.io/badge/python-3.11+-4584b6)
![license](https://img.shields.io/badge/license-MIT-8A98A8)

<img src=".github/assets/board-desktop.png" alt="The Todoie board — Today view with due tasks and routines" width="85%">

</div>

## What it does

- **Capture by talking.** *"renew passport before october"* → filed under **Life**, due **Oct 1**. The model extracts the time phrase *verbatim*; the calendar date is resolved **deterministically in code** — that one design choice took due-date accuracy from 54.5% (model does the math) to **96.6%, zero hallucinated dates**.
- **One model call routes everything.** A single Cohere **Command A+** tool-use call classifies intent and extracts fields across **8 typed actions** — capture, routine, complete, edit, search, ask, list, reply — including compound edits (*"move the report to Sunday and the rest to July 31st"* → one edit per task).
- **Grounded answers, not vibes.** Questions run a **plan-execute RAG** loop: the planner picks exact filters, semantic search, or the routine corpus; the answer cites the actual tasks it read — and refuses when retrieval comes back empty.
- **Duplicate detection that asks first.** New captures are compared by cosine similarity over **Embed v4** embeddings (cached in SQLite as float32 blobs). The 0.65 threshold is the **F0.5-optimal point on a labeled precision/recall sweep** — and a merge never happens without your tap.
- **Routines with streaks** — and one freeze that absorbs a missed day.
- **Every decision on the record.** Each model call lands in an `ai_decisions` audit table (input, output, similarity, latency, action) and surfaces in the UI as a per-answer **"Steps"** trace — a view over the log, never faked.
- **MCP server.** The same pipeline, exposed as typed tools (`add_task`, `ask_tasks`, `complete_task`, …) for Claude or any MCP-capable agent — confirmation-gating included.

## Watch it think

The live demo plays the real app beside its pipeline — route → embed → decide,
plan → retrieve → cite — as each decision lands:

<div align="center">
<a href="https://todoie.net"><img src=".github/assets/landing-hero.png" alt="The real app running beside its decision pipeline" width="90%"></a>
</div>

## Measured, not guessed

Every prompt and threshold is backed by an eval you can re-run — **six suites**, one command:

<div align="center">
<img src=".github/assets/landing-eval.png" alt="Eval numbers and the dedup threshold sweep with the F0.5-optimal operating point" width="90%">
</div>

| Suite | What it measures | Latest run (2026-07-21) |
|---|---|---|
| Routing | intent accuracy over labeled messages, incl. history-dependent follow-ups | **92% · 77 cases** (the original 56 stable at ~93% across runs; the 21 added are deliberately hard) |
| Due dates | verbatim-phrase extraction + deterministic resolution | **96.6% · 29 cases · 0 hallucinated** |
| Dedup threshold | precision/recall sweep over labeled pairs with hard negatives | operating point **0.65 = max F0.5** |
| Search floor | relevance-cut sweep on a labeled query set | **max(0.20, top×0.88)** — earned, not hand-tuned |
| Safety | adversarial prompt-injection routing | **12/12 safe** |
| Vectors | exact NumPy vs HNSW benchmark | NumPy wins at this scale — the measured case for **no vector DB** |

```bash
python eval/run_eval.py            # all suites (needs a Cohere key)
python eval/run_eval.py --router   # or one at a time: --dates --dedupe --search --safety --vectors
```

Both labeled sets grow from production: any misroute noticed in the wild becomes
tomorrow's eval case. **170 offline tests** run with Cohere fully mocked — no
network, ~2 seconds — on every push.

## How it works

```
message ──► router (Command A+, one tool-use call)
              │
              ├─ file_todo(title, note, category, due_phrase)
              │     └─ Embed v4 → cosine vs ALL open items
              │          ├─ no near-match → saved silently
              │          └─ near-match → recommend merge / keep — USER confirms
              ├─ file_routine(title, rule)    → recurring habits
              ├─ complete_todo(query)         → semantic match; confident → done + Undo
              ├─ edit_todo(query, changes)    → reschedule / rename / reprioritize
              ├─ search_todos(query)          → semantic existence check
              ├─ ask_todos(query)             → plan (filters | search | routines)
              │                                 → retrieve → grounded, cited answer
              ├─ list_todos(category | null)  → grouped list in chat
              └─ reply(text)                  → just answers
```

The pipeline is channel-agnostic — the web app, the Telegram bot, and the MCP
server are thin clients over the same `route → validate → act → log` core.
[docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) explains every moving part in plain language.

## Quickstart

```bash
git clone https://github.com/DavidStor/todoie && cd todoie
cp .env.example .env            # paste your Cohere key (free at dashboard.cohere.com)
pip install -r requirements.txt
python -m app.main              # → http://localhost:8080/demo/
```

**Use it from Claude (or any MCP client):**

```json
{ "todoie": { "command": "python3", "args": ["-m", "app.mcp_server"], "cwd": "/path/to/todoie" } }
```

**Optional Telegram surface:** set `TELEGRAM_BOT_TOKEN` in `.env` (token from
@BotFather) and the same process runs the bot + Mini App board. Without it,
Todoie runs web-only.

## Docs

| Doc | What's inside |
|---|---|
| [ARCHITECTURE.md](docs/ARCHITECTURE.md) | how it's wired — the router pattern, embeddings, confidence gating, why no vector DB |
| [MCP.md](docs/MCP.md) | the MCP server: tools, client config, the Cohere North seam |
| [OPERATIONS.md](docs/OPERATIONS.md) | running it in public — per-message costs, abuse bounds, the daily budget breaker |
| [ROADMAP.md](docs/ROADMAP.md) | known issues (honest ones) and what's next |
| [STATE.md](docs/STATE.md) | the current feature inventory, measured |

*The todoie.net landing page ships from a separate deploy repo — this one is the agent.*

## Security model

Keys live in `.env` (gitignored). The public demo binds every request to a
per-session SQLite sandbox that can never touch the owner's data — a property
the test suite asserts in both directions. Nothing is hard-deleted: merged,
superseded, and deleted are soft states, so the original text of every capture
survives every operation.

## License

MIT — built by [David Storozhenko](https://www.linkedin.com/in/david-storozhenko/),
powered by [Cohere Command A+ and Embed v4](https://cohere.com).

Maintenance

ActivityMaintained
ResponsivenessSyncing