Skip to main content
Glama
README.md
# Tech Recall

**Save a technology now. Recall it when your work needs it.**

English · [简体中文](README.zh-CN.md)

[![CI](https://github.com/WItaZhang/tech-recall/actions/workflows/ci.yml/badge.svg)](https://github.com/WItaZhang/tech-recall/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

Tech Recall is a **local Codex agent plugin that turns technology bookmarks into suggestions for the task at hand**. Save a framework, method, paper, or repository through a conversation. The agent researches it, prepares a sourced technology card, and asks you to confirm it. In a later task, the plugin retrieves relevant cards and lets the current Codex agent decide whether one deserves a brief suggestion.

The problem is familiar: you discover a useful tool, save its link, then forget about it when it could help. Tech Recall connects **what you collected** with **what you are building**, while keeping adoption under your control. Accepting a suggestion produces a trial plan; implementation remains a separate decision.

**Stack:** Python 3.12 · uv · SQLite FTS5 · local stdio MCP · Codex skills and hooks

[Quick start](#quick-start) · [Usage](#usage) · [Components](#components-and-runtime) · [Engineering decisions](#engineering-decisions) · [Validation](#validation-and-evaluation)

## What it looks like

An illustrative interaction—not a recorded model run:

```text
You:    $new-stuff Pydantic
Codex:  Opens primary sources and prepares a card with uses, exclusions,
        integration cost, repositories, and checked versions.
Form:   Review the card → Save / Cancel

Later, in a different project:
You:    Design input validation for my Python agent's tools.
Codex:  Continues the task and, if relevant, adds one short hint:
        “Your saved Pydantic card may fit this tool boundary. Expand …”

You:    Expand the suggestion.
Codex:  Checks current sources and explains the fit and constraints.
Form:   Accept and plan / Not in this task / Delete / Update
You:    Accept and plan.
Codex:  Proposes an integration point, a minimal experiment,
        acceptance criteria, dependencies, and a rollback path.
```

Ignoring a hint leaves the main task running. Each turn allows at most one proactive suggestion; each saved technology is suggested at most once per task. You can explicitly revisit it later.

**Current status — v0.1:** the core workflow, native-confirmation protocol, installed runtime, and cross-project discovery have been tested. Actual Codex desktop form rendering and autonomous recommendation behavior still require a fresh-task acceptance run. See [validation](#validation-and-evaluation) for the evidence and remaining checks.

## Quick start

### 1. Prepare the repository

You need Git and [uv](https://docs.astral.sh/uv/). The repository selects Python 3.12 and pins dependencies in `uv.lock`.

```sh
git clone https://github.com/WItaZhang/tech-recall.git
cd tech-recall
uv sync --frozen
```

To inspect the protocol workflow before installing the plugin:

```sh
uv run tech-recall demo --record artifacts/demo.cast
```

The demo starts a real stdio MCP subprocess and uses a temporary library with **simulated user answers**. It needs no Codex session or model API key and does not modify your collection. A checked-in [terminal recording](docs/demo.cast) and [playback HTML](docs/demo.html) are available; download/open the HTML locally to play it. This is a protocol recording, not a desktop UI recording.

### 2. Install globally in Codex

The host must support plugins, `SessionStart` / `UserPromptSubmit` hooks, and native MCP form elicitation. Local integration was checked on **Windows with Codex `0.155.0-alpha.16`**; this is a tested version, not a minimum-version claim. The installer also requires Codex's bundled `plugin-creator` helpers.

```sh
uv run python scripts/install.py
```

The installer builds an allowlisted bundle, registers it in your personal marketplace, installs it globally, prepares its uv environment, and adds the managed `$new-stuff` shortcut. It preserves unrelated skills and marketplace entries. The collection is stored outside the plugin cache.

In Codex, open **`/hooks`**, review and trust Tech Recall's two hooks, then **start a fresh task**. If the desktop build has no hook browser, use the Codex CLI `/hooks` with the same configuration. Hook trust is a host requirement and is never enabled by the installer.

### 3. Save your first technology

Enter this in the **Codex conversation**, not the terminal:

```text
$new-stuff Pydantic
```

Review the native confirmation form and save the card. The library starts empty; automatic suggestions become useful after you collect technologies. Daily research and judgment use your existing Codex session and browsing capabilities; no separate OpenAI API key is needed for the plugin.

Missing helpers, unsupported forms, updates, and removal are covered in the [installation guide](docs/installation.md).

## Usage

### Collect and inspect

Supply a name or a source URL. Ambiguous names are resolved with you before a card is prepared.

```text
$new-stuff https://github.com/pydantic/pydantic
$tech-recall:tech-recall Search my collection for Python input validation.
$tech-recall:tech-recall Show deleted cards so I can restore one.
```

`$new-stuff` is the installed shortcut; `$tech-recall:new-stuff` is the canonical plugin skill. `$tech-recall:tech-recall` handles search and collection management. These are conversational skill invocations, not shell subcommands.

Every **technology card** records:

| Information | Why it is stored |
| --- | --- |
| Name, canonical identity URL, bilingual aliases, tags | Resolve identity and retrieve it across different task wording |
| Problem, applicable conditions, excluded uses | Help the agent assess fit rather than rely on a shared keyword |
| Integration cost and constraints | Make adoption effort part of the decision |
| Source claims, URLs, and checked timestamps | Make the research inspectable and freshness explicit |
| Related repositories and optional version | Support a concrete trial; methods and papers can leave these empty |

### Act on a suggestion

Only expanding a hint opens the action flow. The agent checks current primary sources first, reusing a task/revision-scoped check for up to one hour by default. Failed checks are marked `unknown`; checking alone does not overwrite the saved card.

| Choice | Result |
| --- | --- |
| **Accept and plan** | Generate a trial plan with the project's integration point, smallest experiment, checked dependencies/version, measurable criteria, rollback, and citations |
| **Not in this task** | Suppress this technology for the current task; other tasks remain eligible |
| **Delete** | Remove it from active retrieval globally; keep a restorable snapshot |
| **Update** | Research a new snapshot, show the revision diff, and save only after native confirmation |

For example, a validation trial could target the agent's tool-input boundary and define valid, malformed, and unexpected-field cases as acceptance criteria. The plan must explain how to remove the experiment. Accepting it does not install dependencies, edit the project, or execute a trial.

## Components and runtime

The **current Codex agent handles research and semantic judgment**. The local Python service handles retrieval and persistent state. A separate model adapter exists only for explicitly enabled evaluation. Hooks use the current task's context and never scan other conversations.

MCP (Model Context Protocol) lets Codex call the local service's tools over standard input/output. Skills describe how the agent should use those tools; hooks supply context when a session starts or a user prompt arrives.

```mermaid
flowchart TD
    U[Current task input] --> H[Local hook: retrieve up to 5 candidates]
    H --> A[Codex agent: assess fit using skill instructions]
    S[Explicit new-stuff request] --> A
    A -->|No useful match| Q[Continue without a hint]
    A -->|Read, propose, or manage| M[Local stdio MCP server]
    M -->|Operations requiring consent| F[Codex native confirmation form]
    F -->|User response| M
    M --> C[Core service: validate state and revisions]
    C --> D[(SQLite: cards, history, task state, traces)]
    H -->|Read-only FTS5 search| D
```

| Component | Responsibility | Start reading |
| --- | --- | --- |
| **Skills** | Identity resolution, source research, fit assessment, and trial-plan instructions | [`skills/`](skills) |
| **Hooks** | `SessionStart` supplies the workflow convention; `UserPromptSubmit` retrieves candidates from the current prompt and this task's saved summary | [`hooks.py`](src/tech_recall/hooks.py), [`hook.py`](scripts/hook.py) |
| **MCP adapter** | Expose typed tools over stdio and receive native user confirmation | [`server.py`](src/tech_recall/server.py) |
| **Core service and contracts** | Manage drafts, cards, revisions, feedback, verification caches, and task isolation | [`service.py`](src/tech_recall/service.py), [`models.py`](src/tech_recall/models.py) |
| **Storage and retrieval** | SQLite transactions, FTS5/BM25 ranking, bilingual aliases, and Chinese character bigrams | [`storage.py`](src/tech_recall/storage.py), [`search.py`](src/tech_recall/search.py) |
| **Evaluation** | Run fixed-case replay or an opt-in OpenAI Responses judge; report decisions, latency, and available usage | [`evaluation/`](src/tech_recall/evaluation), [`evals/`](evals) |

The core does not depend on Codex UI APIs. A future host adapter can reuse its state machine. Full tool schemas and behavior are summarized in [MCP contracts](docs/interfaces.md).

## Engineering decisions

**1. Separate retrieval from recommendation.** FTS5 returns up to five candidates; Codex then checks prerequisites, exclusions, the existing stack, and adoption cost. For example, a resource-authorization tool should not be suggested as a fraud classifier simply because both mention “user risk.” Lexical retrieval keeps the hook inexpensive, but can miss paraphrases; bilingual aliases and agent query expansion help without guaranteeing recall.

**2. Make confirmation a state transition.** Cards first become immutable, expiring drafts. Native MCP elicitation supplies the user's answer outside the model-visible arguments; there is no `confirmed=true` tool parameter. The commit transaction rechecks ownership, expiry, and revision after the form returns. Concurrent updates produce a conflict, and a repeated completed save returns its previous result. Canonical source identity prevents same-name technologies from being silently merged.

**3. Bound interruptions and failures.** SQLite uniqueness constraints reserve one hint per task/card and per task/turn. The hook performs local reads only, with no model call, network research, or dependency synchronization. A one-second host timeout and process watchdog bound execution. A hook failure releases the main task; missing native confirmation leaves a draft uncommitted.

**4. Track freshness without silently rewriting memory.** Source checks are cached by task, card, and revision. A changed card invalidates the cache. Updating the collection requires a newly confirmed diff, while delete immediately removes the FTS entry and restore rebuilds it. The service validates evidence structure and timestamps; the host agent remains responsible for actually reading and interpreting the sources.

**5. Keep evidence reproducible.** Metadata traces record run IDs, stage durations, record/source references, and outcomes. Full prompts and source code are not logged; short task summaries are explicitly stored through a bounded tool. Unavailable token/cost values remain `null`. Evaluation inputs exclude answer labels, and config snapshots, predictions, and metrics are saved per run.

See [architecture and tradeoffs](docs/architecture.md) for the state machine, trust boundaries, and storage behavior.

## Validation and evaluation

The repository distinguishes **protocol correctness**, **local performance**, and **recommendation quality**.

| Area | Evidence | What it establishes |
| --- | --- | --- |
| Automated regression | **33 passing tests**, [Windows and Linux CI](https://github.com/WItaZhang/tech-recall/actions/workflows/ci.yml) | Confirmation/cancellation, expiry, deduplication, concurrent writes, task isolation, restore, and failure handling |
| Installed integration | [Two-project Codex registry check](docs/benchmarks/codex-registry.json) plus installed MCP/hook smoke checks | Global discovery and executable local integration; hooks still require user trust |
| Hook performance | **1,000 synthetic cards, 30 process invocations; P95 332.5 ms, max 452.3 ms** on Windows / Python 3.12.4 | End-to-end local hook timing including uv and interpreter startup; below the 500 ms target in this run |
| Recommendation regression | **50 fixed synthetic scenarios, 10 sourced cards** | Coverage of suitable uses, homonyms, ambiguity, constraint conflicts, routine edits, and stale snapshots |
| Desktop experience | [Manual acceptance procedure](docs/acceptance.md) | Still pending: actual popup rendering and relevance/quietness in real Codex tasks |

Raw hook samples are [checked in](docs/benchmarks/hooks-windows.json). They are a development-machine measurement, not a cross-platform latency guarantee. The scenarios were reviewed by the implementing agent; they are not an independently human-annotated benchmark.

**What the quality numbers mean.** On this deliberately adversarial dataset, the keyword-only baseline made 49 suggestions, of which 20 were correct: **40.8% precision**, **100% recall**, and **29/30 negative tasks interrupted**. These numbers describe this fixture only. The checked-in semantic replay was authored from expected answers, so its perfect scores validate regression plumbing, not model quality. **No live semantic-accuracy improvement or API-cost result is claimed.** [Metrics and methodology](docs/evaluation.md)

Run the checks from the repository root:

```sh
uv run python -m pytest -q
uv run ruff check .
uv run ruff format --check .
uv run tech-recall eval
uv run python scripts/benchmark_hooks.py
```

For an **explicit paid model evaluation**, set `[model].model` in [`configs/evaluation.toml`](configs/evaluation.toml), supply `OPENAI_API_KEY` outside the repository, and run:

```sh
uv run tech-recall eval --live
```

This sends fixture tasks and retrieved candidates to the configured OpenAI Responses model, without expected labels or your personal collection. Each run writes config, predictions, results, timings, and available token/cost data under `logs/`. No live API evaluation has been run for v0.1. Independent API evaluation also does not replace real Codex acceptance testing.

## Repository map

```text
tech-recall/
├── .codex-plugin/       Plugin manifest
├── .mcp.json            Local MCP launch configuration
├── skills/              Collection and recall workflows
├── hooks/               Codex hook registrations
├── src/tech_recall/     Core service, storage, retrieval, MCP, CLI, evaluation
├── configs/             Runtime policy example and evaluation/benchmark settings
├── evals/               Fixed cards, 50 scenarios, and authored replay fixtures
├── tests/               State, protocol, hook, evaluation, and packaging tests
├── scripts/             Bundle builder, installer, integration checks, benchmark
├── docs/                Architecture, interfaces, evidence, and demo recording
└── uv.lock              Pinned dependency environment
```

For a code review, start with [`server.py`](src/tech_recall/server.py) → [`service.py`](src/tech_recall/service.py) → [`tests/test_mcp.py`](tests/test_mcp.py) to follow the confirmation boundary, or [`search.py`](src/tech_recall/search.py) → [`evaluation/runner.py`](src/tech_recall/evaluation/runner.py) to follow retrieval and measurement.

## Configuration and maintenance

Collections live in `%LOCALAPPDATA%/tech-recall` on Windows, or `${XDG_DATA_HOME:-~/.local/share}/tech-recall` elsewhere. Set `TECH_RECALL_HOME` to override the location. Storage is local plaintext; cards, revisions, and feedback survive plugin upgrades. Metadata traces default to 30-day retention, pruned on subsequent trace writes.

```sh
uv run tech-recall doctor                 # Runtime capabilities and data location
uv run tech-recall traces --limit 20      # Recent workflow metadata
```

Copy [`configs/policy.example.toml`](configs/policy.example.toml) to the data directory as `config.toml` to adjust policy. Defaults include five candidates, 24-hour draft expiry, one-hour verification reuse, and a one-second maximum hook deadline. Source checkout, generated plugin bundle, installed environment, and collection are separate; the bundle excludes Git history, virtual environments, logs, and user data.

To update an existing installation:

```sh
git pull
uv sync --frozen
uv run python scripts/install.py --refresh
```

Review changed hooks and start a fresh task. If suggestions are missing, first check hook trust, that the library contains relevant cards, and once-per-task suppression. More diagnostics and uninstall instructions are in the [installation guide](docs/installation.md).

## Scope

v0.1 targets **local, single-user Codex**. Claude integration, cloud synchronization, background scheduled research, a standalone management UI, and automatic trial execution are outside this release. The next validation priorities are a real Codex UI recording, independently reviewed holdout cases, and measured live recommendation quality.

[MIT License](LICENSE)