tech-recall
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@tech-recallsave a technology bookmark for Pydantic"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tech Recall
Save a technology now. Recall it when your work needs it.
English · 简体中文
Tech Recall is a local Codex agent plugin that turns technology bookmarks into suggestions for the task at hand. Save a framework, method, paper, or repository through a conversation. The agent researches it, prepares a sourced technology card, and asks you to confirm it. In a later task, the plugin retrieves relevant cards and lets the current Codex agent decide whether one deserves a brief suggestion.
The problem is familiar: you discover a useful tool, save its link, then forget about it when it could help. Tech Recall connects what you collected with what you are building, while keeping adoption under your control. Accepting a suggestion produces a trial plan; implementation remains a separate decision.
Stack: Python 3.12 · uv · SQLite FTS5 · local stdio MCP · Codex skills and hooks
Quick start · Usage · Components · Engineering decisions · Validation
What it looks like
An illustrative interaction—not a recorded model run:
You: $new-stuff Pydantic
Codex: Opens primary sources and prepares a card with uses, exclusions,
integration cost, repositories, and checked versions.
Form: Review the card → Save / Cancel
Later, in a different project:
You: Design input validation for my Python agent's tools.
Codex: Continues the task and, if relevant, adds one short hint:
“Your saved Pydantic card may fit this tool boundary. Expand …”
You: Expand the suggestion.
Codex: Checks current sources and explains the fit and constraints.
Form: Accept and plan / Not in this task / Delete / Update
You: Accept and plan.
Codex: Proposes an integration point, a minimal experiment,
acceptance criteria, dependencies, and a rollback path.Ignoring a hint leaves the main task running. Each turn allows at most one proactive suggestion; each saved technology is suggested at most once per task. You can explicitly revisit it later.
Current status — v0.1: the core workflow, native-confirmation protocol, installed runtime, and cross-project discovery have been tested. Actual Codex desktop form rendering and autonomous recommendation behavior still require a fresh-task acceptance run. See validation for the evidence and remaining checks.
Related MCP server: mcp-bookmark-server
Quick start
1. Prepare the repository
You need Git and uv. The repository selects Python 3.12 and pins dependencies in uv.lock.
git clone https://github.com/WItaZhang/tech-recall.git
cd tech-recall
uv sync --frozenTo inspect the protocol workflow before installing the plugin:
uv run tech-recall demo --record artifacts/demo.castThe demo starts a real stdio MCP subprocess and uses a temporary library with simulated user answers. It needs no Codex session or model API key and does not modify your collection. A checked-in terminal recording and playback HTML are available; download/open the HTML locally to play it. This is a protocol recording, not a desktop UI recording.
2. Install globally in Codex
The host must support plugins, SessionStart / UserPromptSubmit hooks, and native MCP form elicitation. Local integration was checked on Windows with Codex 0.155.0-alpha.16; this is a tested version, not a minimum-version claim. The installer also requires Codex's bundled plugin-creator helpers.
uv run python scripts/install.pyThe installer builds an allowlisted bundle, registers it in your personal marketplace, installs it globally, prepares its uv environment, and adds the managed $new-stuff shortcut. It preserves unrelated skills and marketplace entries. The collection is stored outside the plugin cache.
In Codex, open /hooks, review and trust Tech Recall's two hooks, then start a fresh task. If the desktop build has no hook browser, use the Codex CLI /hooks with the same configuration. Hook trust is a host requirement and is never enabled by the installer.
3. Save your first technology
Enter this in the Codex conversation, not the terminal:
$new-stuff PydanticReview the native confirmation form and save the card. The library starts empty; automatic suggestions become useful after you collect technologies. Daily research and judgment use your existing Codex session and browsing capabilities; no separate OpenAI API key is needed for the plugin.
Missing helpers, unsupported forms, updates, and removal are covered in the installation guide.
Usage
Collect and inspect
Supply a name or a source URL. Ambiguous names are resolved with you before a card is prepared.
$new-stuff https://github.com/pydantic/pydantic
$tech-recall:tech-recall Search my collection for Python input validation.
$tech-recall:tech-recall Show deleted cards so I can restore one.$new-stuff is the installed shortcut; $tech-recall:new-stuff is the canonical plugin skill. $tech-recall:tech-recall handles search and collection management. These are conversational skill invocations, not shell subcommands.
Every technology card records:
Information | Why it is stored |
Name, canonical identity URL, bilingual aliases, tags | Resolve identity and retrieve it across different task wording |
Problem, applicable conditions, excluded uses | Help the agent assess fit rather than rely on a shared keyword |
Integration cost and constraints | Make adoption effort part of the decision |
Source claims, URLs, and checked timestamps | Make the research inspectable and freshness explicit |
Related repositories and optional version | Support a concrete trial; methods and papers can leave these empty |
Act on a suggestion
Only expanding a hint opens the action flow. The agent checks current primary sources first, reusing a task/revision-scoped check for up to one hour by default. Failed checks are marked unknown; checking alone does not overwrite the saved card.
Choice | Result |
Accept and plan | Generate a trial plan with the project's integration point, smallest experiment, checked dependencies/version, measurable criteria, rollback, and citations |
Not in this task | Suppress this technology for the current task; other tasks remain eligible |
Delete | Remove it from active retrieval globally; keep a restorable snapshot |
Update | Research a new snapshot, show the revision diff, and save only after native confirmation |
For example, a validation trial could target the agent's tool-input boundary and define valid, malformed, and unexpected-field cases as acceptance criteria. The plan must explain how to remove the experiment. Accepting it does not install dependencies, edit the project, or execute a trial.
Components and runtime
The current Codex agent handles research and semantic judgment. The local Python service handles retrieval and persistent state. A separate model adapter exists only for explicitly enabled evaluation. Hooks use the current task's context and never scan other conversations.
MCP (Model Context Protocol) lets Codex call the local service's tools over standard input/output. Skills describe how the agent should use those tools; hooks supply context when a session starts or a user prompt arrives.
flowchart TD
U[Current task input] --> H[Local hook: retrieve up to 5 candidates]
H --> A[Codex agent: assess fit using skill instructions]
S[Explicit new-stuff request] --> A
A -->|No useful match| Q[Continue without a hint]
A -->|Read, propose, or manage| M[Local stdio MCP server]
M -->|Operations requiring consent| F[Codex native confirmation form]
F -->|User response| M
M --> C[Core service: validate state and revisions]
C --> D[(SQLite: cards, history, task state, traces)]
H -->|Read-only FTS5 search| DComponent | Responsibility | Start reading |
Skills | Identity resolution, source research, fit assessment, and trial-plan instructions | |
Hooks |
| |
MCP adapter | Expose typed tools over stdio and receive native user confirmation | |
Core service and contracts | Manage drafts, cards, revisions, feedback, verification caches, and task isolation | |
Storage and retrieval | SQLite transactions, FTS5/BM25 ranking, bilingual aliases, and Chinese character bigrams | |
Evaluation | Run fixed-case replay or an opt-in OpenAI Responses judge; report decisions, latency, and available usage |
The core does not depend on Codex UI APIs. A future host adapter can reuse its state machine. Full tool schemas and behavior are summarized in MCP contracts.
Engineering decisions
1. Separate retrieval from recommendation. FTS5 returns up to five candidates; Codex then checks prerequisites, exclusions, the existing stack, and adoption cost. For example, a resource-authorization tool should not be suggested as a fraud classifier simply because both mention “user risk.” Lexical retrieval keeps the hook inexpensive, but can miss paraphrases; bilingual aliases and agent query expansion help without guaranteeing recall.
2. Make confirmation a state transition. Cards first become immutable, expiring drafts. Native MCP elicitation supplies the user's answer outside the model-visible arguments; there is no confirmed=true tool parameter. The commit transaction rechecks ownership, expiry, and revision after the form returns. Concurrent updates produce a conflict, and a repeated completed save returns its previous result. Canonical source identity prevents same-name technologies from being silently merged.
3. Bound interruptions and failures. SQLite uniqueness constraints reserve one hint per task/card and per task/turn. The hook performs local reads only, with no model call, network research, or dependency synchronization. A one-second host timeout and process watchdog bound execution. A hook failure releases the main task; missing native confirmation leaves a draft uncommitted.
4. Track freshness without silently rewriting memory. Source checks are cached by task, card, and revision. A changed card invalidates the cache. Updating the collection requires a newly confirmed diff, while delete immediately removes the FTS entry and restore rebuilds it. The service validates evidence structure and timestamps; the host agent remains responsible for actually reading and interpreting the sources.
5. Keep evidence reproducible. Metadata traces record run IDs, stage durations, record/source references, and outcomes. Full prompts and source code are not logged; short task summaries are explicitly stored through a bounded tool. Unavailable token/cost values remain null. Evaluation inputs exclude answer labels, and config snapshots, predictions, and metrics are saved per run.
See architecture and tradeoffs for the state machine, trust boundaries, and storage behavior.
Validation and evaluation
The repository distinguishes protocol correctness, local performance, and recommendation quality.
Area | Evidence | What it establishes |
Automated regression | 33 passing tests, Windows and Linux CI | Confirmation/cancellation, expiry, deduplication, concurrent writes, task isolation, restore, and failure handling |
Installed integration | Two-project Codex registry check plus installed MCP/hook smoke checks | Global discovery and executable local integration; hooks still require user trust |
Hook performance | 1,000 synthetic cards, 30 process invocations; P95 332.5 ms, max 452.3 ms on Windows / Python 3.12.4 | End-to-end local hook timing including uv and interpreter startup; below the 500 ms target in this run |
Recommendation regression | 50 fixed synthetic scenarios, 10 sourced cards | Coverage of suitable uses, homonyms, ambiguity, constraint conflicts, routine edits, and stale snapshots |
Desktop experience | Still pending: actual popup rendering and relevance/quietness in real Codex tasks |
Raw hook samples are checked in. They are a development-machine measurement, not a cross-platform latency guarantee. The scenarios were reviewed by the implementing agent; they are not an independently human-annotated benchmark.
What the quality numbers mean. On this deliberately adversarial dataset, the keyword-only baseline made 49 suggestions, of which 20 were correct: 40.8% precision, 100% recall, and 29/30 negative tasks interrupted. These numbers describe this fixture only. The checked-in semantic replay was authored from expected answers, so its perfect scores validate regression plumbing, not model quality. No live semantic-accuracy improvement or API-cost result is claimed. Metrics and methodology
Run the checks from the repository root:
uv run python -m pytest -q
uv run ruff check .
uv run ruff format --check .
uv run tech-recall eval
uv run python scripts/benchmark_hooks.pyFor an explicit paid model evaluation, set [model].model in configs/evaluation.toml, supply OPENAI_API_KEY outside the repository, and run:
uv run tech-recall eval --liveThis sends fixture tasks and retrieved candidates to the configured OpenAI Responses model, without expected labels or your personal collection. Each run writes config, predictions, results, timings, and available token/cost data under logs/. No live API evaluation has been run for v0.1. Independent API evaluation also does not replace real Codex acceptance testing.
Repository map
tech-recall/
├── .codex-plugin/ Plugin manifest
├── .mcp.json Local MCP launch configuration
├── skills/ Collection and recall workflows
├── hooks/ Codex hook registrations
├── src/tech_recall/ Core service, storage, retrieval, MCP, CLI, evaluation
├── configs/ Runtime policy example and evaluation/benchmark settings
├── evals/ Fixed cards, 50 scenarios, and authored replay fixtures
├── tests/ State, protocol, hook, evaluation, and packaging tests
├── scripts/ Bundle builder, installer, integration checks, benchmark
├── docs/ Architecture, interfaces, evidence, and demo recording
└── uv.lock Pinned dependency environmentFor a code review, start with server.py → service.py → tests/test_mcp.py to follow the confirmation boundary, or search.py → evaluation/runner.py to follow retrieval and measurement.
Configuration and maintenance
Collections live in %LOCALAPPDATA%/tech-recall on Windows, or ${XDG_DATA_HOME:-~/.local/share}/tech-recall elsewhere. Set TECH_RECALL_HOME to override the location. Storage is local plaintext; cards, revisions, and feedback survive plugin upgrades. Metadata traces default to 30-day retention, pruned on subsequent trace writes.
uv run tech-recall doctor # Runtime capabilities and data location
uv run tech-recall traces --limit 20 # Recent workflow metadataCopy configs/policy.example.toml to the data directory as config.toml to adjust policy. Defaults include five candidates, 24-hour draft expiry, one-hour verification reuse, and a one-second maximum hook deadline. Source checkout, generated plugin bundle, installed environment, and collection are separate; the bundle excludes Git history, virtual environments, logs, and user data.
To update an existing installation:
git pull
uv sync --frozen
uv run python scripts/install.py --refreshReview changed hooks and start a fresh task. If suggestions are missing, first check hook trust, that the library contains relevant cards, and once-per-task suppression. More diagnostics and uninstall instructions are in the installation guide.
Scope
v0.1 targets local, single-user Codex. Claude integration, cloud synchronization, background scheduled research, a standalone management UI, and automatic trial execution are outside this release. The next validation priorities are a real Codex UI recording, independently reviewed holdout cases, and measured live recommendation quality.
Related MCP Connectors
Shared, permission-aware company context for AI agents, with provenance, approvals and audit.
Save and organize web finds in persistent, user-controlled collections for AI assistants.
Governed memory and workspace for any AI: tasks, calendar, mail and pages, with per-action consent.
Cited, versioned knowledge for agents: retrieve sourced passages and propose owner-approved fixes.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables LLM agents to create annotated bookmarks and narrative trails to document code exploration and decision-making. It organizes these bookmarks into a navigable knowledge graph stored in a local JSON file, helping users track the context and reasoning behind code changes.10 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to save and search bookmarks using OpenAI's RAG capabilities for intelligent bookmark management and retrieval.MIT
- AlicenseBqualityDmaintenanceEnables AI assistants to save and search bookmarks with semantic search using OpenAI, allowing storage of URLs with metadata and intelligent retrieval across collections.219 PyPIMIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to organize and manage Chrome/Edge bookmarks through natural language, with user-controlled permissions, privacy filters, and safety mechanisms.Apache 2.0