rekall
Rekall is an MCP server that manages Codex VS Code conversation compaction with optional verified handoffs and automatic continuation.
probe_compaction(threadId): read-only preflight checking owner/thread, extension compatibility, runtime layout, and IPC access before scheduling.schedule_compaction(threadId, handoff?): queue one compaction after the current turn, optionally saving a handoff (summary, preserve/discard lists, nextStep, resume flag) for a single authorized continuation.compaction_status(threadId): inspect the current job, state, metrics, and completion/continuation outcome.cancel_compaction(threadId, jobId): cancel a pending dispatch for the exact job before it is sent.It also provides CLI equivalents and enforces safeguards: never guesses threads, no automatic retries, and continuation only with sufficient context headroom.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@rekallCompact this thread with a handoff, then continue the remaining work once."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Rekall
Long Codex tasks accumulate logs, research, and intermediate decisions. Rekall lets the agent clear that accumulated context at a useful checkpoint while keeping a written handoff of the task, constraints, and next step.
The agent saves the handoff, finishes its turn, and asks the Codex VS Code extension to compact the conversation. Rekall can then resume the authorized work once, carrying the verified handoff into the next turn.
Works with the Codex extension in VS Code on Windows, macOS, and Linux. Standalone Codex CLI sessions are not supported: Rekall depends on the extension's IPC owner and conversation lifecycle, which its current adapter cannot access for a standalone CLI session.
Install
Requires Windows, macOS, or Linux, Node.js 20 or newer on PATH, the Codex VS Code extension, and a Codex CLI with the plugin commands. The plugin installation commands below work in PowerShell, zsh, and bash.
codex plugin marketplace add DitriXNew/rekall
codex plugin add rekall@rekallThis installs the MCP server and the bundled skill together. Start a new chat in the Codex VS Code extension, then ask:
Compact this thread with a handoff, then continue the remaining work once.
To check access without compacting, ask Codex to run probe_compaction for the current thread. The repository includes its marketplace entry, plugin manifest, and MCP declaration; Codex resolves ${PLUGIN_ROOT} to the installed plugin directory.
Clone the repository and register the MCP server with an absolute path:
git clone https://github.com/DitriXNew/rekall.git
Set-Location rekall
npm ci
codex mcp add rekall -- node "$PWD/bridge.mjs" mcpManual MCP registration installs only the server. Copy skills/rekall into $CODEX_HOME/skills/rekall (default: ~/.codex/skills/rekall) to install the skill, then start a new extension chat.
On macOS or Linux, use cd rekall instead of Set-Location rekall; the other manual installation commands work in zsh or bash.
If migrating an existing manual installation to the plugin, remove the old manual MCP registration and the manually copied skill to avoid duplicate tool/skill discovery. The retired server name was context-compact; current manual installations use rekall. Keep the job directory so existing jobs and locks remain available.
Related MCP server: compaction-mcp
Measured results
Observed context-token reductions: 78–90%, with automatic continuation about one second after compaction.
Run | Context tokens before → after | Reduction | Compaction | Continuation delay |
Release verification | 102,826 → 10,537 | 89.8% | 87.5 s | 0.908 s |
Installed-package verification | 55,856 → 12,128 | 78.3% | 88.5 s | 1.065 s |
macOS ARM verification | 80,137 → 9,216 | 88.5% | 98.4 s | 0.812 s |
Earlier user-reported run | 91,384 → 10,798 | 88.2% | ~2 min | ~0.9 s |
The resumed agent read the saved handoff in all three measured verification runs. The macOS run also verified its SHA-256 and exact thread/job binding. Linux operation was additionally reported as tested by the maintainer; no Linux token or timing measurements were supplied. See the sanitized verification record. Measurement limits and compatibility details are below.
Scope
Rekall operates on chats owned by the Codex VS Code extension on Windows, macOS, or Linux. Standalone Codex CLI sessions, the Codex desktop app, and Claude Code are not supported. The CLI commands below are another way to address an extension-owned chat; they do not add support for standalone CLI conversations. Node.js is required; there is no standalone executable.
Windows uses the extension's named pipe. macOS and Linux use $CODEX_HOME/ipc/ipc.sock, defaulting to ~/.codex/ipc/ipc.sock. Before connecting, Rekall requires the IPC directory and socket to belong to the current user and have no group/other permissions; symlinks at those two paths are rejected. Rekall does not create or change the socket or its permissions, and does not fall back to a shared temporary socket.
MCP tools
Tool | Purpose |
| Read-only compatibility, owner, and thread-state check. |
| Queue one compaction after the current response becomes idle. |
| Read the current job and recorded metrics. |
| Cancel a dispatch that has not already been sent. |
Read compaction_status and probe_compaction before scheduling. Use only the exact current CODEX_THREAD_ID, and do not queue a second unfinished job. Rekall does not impose a context-use threshold on a user-requested compaction.
CLI reference
Run these commands from the repository or installed package directory:
Command | Purpose |
| Check the current owner, state, and compatibility. |
| Read the current job's status. |
| Compact without automatic continuation. |
| Compact with the saved handoff and its explicit resume setting. |
| Cancel pending dispatches for the exact job. |
| Run the MCP stdio server. |
Square brackets denote an optional argument, not literal command text. Commands with an optional threadId use CODEX_THREAD_ID when it is omitted. cancel requires both IDs explicitly; copy jobId from status. Outside the extension session, pass its exact known thread ID. Rekall never guesses a thread or chooses the most recently updated conversation. worker is an internal subprocess entry point, not a command to launch manually.
Write handoff JSON as UTF-8 outside the repository and quote its absolute path. Its format is described next.
Handoffs and continuation
An optional handoff has this shape:
{
"summary": "The refactor is complete and its tests pass; release review remains.",
"preserve": ["User constraints", "Changed files and test results"],
"discard": ["Repeated command output", "Superseded investigation notes"],
"nextStep": "Review the package contents, report the result, and stop.",
"resume": true
}The complete handoff is limited to 32,000 UTF-8 bytes and each list to 40 entries. discard identifies conversation history that may be summarized; it never authorizes file deletion. The handoff is stored separately, bound to the thread and job, and checked by SHA-256 before continuation.
Set resume explicitly. When it is true, nextStep must identify concrete work the user has already authorized and include a stopping condition. When the task is finished, the user asked to stop, or further work needs an answer, set resume to false. An automatic continuation does not authorize another compaction.
After scheduling, finish the current response: the worker waits for idle. Do not wait for compaction within that same active turn. On continuation, read the saved handoff, verify the exact job and its result, and perform only the authorized next step.
Automatic continuation requires fresh telemetry showing reduced context tokens and no more than 60% of the context window in use. Otherwise, including when telemetry is missing or stale, Rekall preserves the compaction result and skips continuation with resumeSkipped: "insufficient_headroom". It rechecks this immediately before resuming. This guard limits automatic continuation; it never blocks compaction itself.
Scheduling does not mean compaction completed. scheduled, waiting_for_idle, requesting, and accepted are intermediate states. completed requires a newly observed completed compaction record. resumed means the owner returned a follow-up turn ID; it does not mean that turn's work succeeded. Rekall records the compaction ID, completion time, resume turn ID, and up to 20 per-thread measurements, including job number, compaction duration, resume delay, reclaimed tokens, and reclaimed fraction.
Before dispatch, Rekall requires stable idle state, no pending permission request, and no unconfirmed submission. New user input or a stopped or failed turn cancels a pending dispatch. A request already sent cannot be recalled. The worker deadline is 15 minutes from worker start, including idle waiting. Timeouts and unknown outcomes are terminal and are never retried automatically.
Compatibility and extension updates
The full live compaction/continuation cycle has been verified with openai.chatgpt-26.901.22334-win32-x64 and, on macOS 26.6.2 with Node.js 26.5.0, openai.chatgpt-26.901.22334-darwin-arm64. The macOS run passed IPC, exact-thread ownership, runtime layout, and public-schema checks without an override, observed completed compaction, and resumed with a verified saved handoff. The maintainer also reports successful Linux testing; its distribution, architecture, extension version, and measurements have not been recorded here. Intel Mac live verification is still outstanding. Platform paths are covered by isolated tests. Rekall uses an internal extension IPC protocol, which is not a stable public API. Run probe_compaction after extension updates.
Rekall accepts any extension version number and checks compatibility through extension identity, public schema, runtime layout, owner/thread binding, and the IPC protocol. A version update alone does not block compaction. Protocol changes can still require an adapter update; accepting a version number does not mean every past or future protocol is supported.
probe_compaction reports compatibility.extensionVersion for diagnostics. The former REKALL_ALLOW_UNVERIFIED setting is no longer needed and has no effect.
Reloaded public history may omit the private completed field. Rekall accepts those historical records but does not count them as completion signals. Automatic continuation still requires a newly observed item with completed: true. A read-only probe also passed on Windows with extension 26.903.61454; this is not a live compaction/continuation measurement.
Report update-related failures through the compatibility issue template, including the extension version and redacted probe output or error. You do not need to attempt compaction to report a failed probe.
Compatibility checks compare the public App Server schema with the internal completion signal Rekall observes. Before a thread exposes a compaction item, the probe reports layout_compatible: the public lifecycle and state layout are compatible, while the private completion field has not yet been observed. Rekall locates a unique extension-bundled executable from the extension's PATH entries: codex.exe on Windows or codex on macOS or Linux. When that is not possible, set REKALL_CODEX_BINARY to its absolute path. Schema generation exports files and exits; it does not start another App Server. Test transports using REKALL_PIPE deliberately skip the schema subprocess and production endpoint discovery/validation; do not use this test override for a live socket.
Security
Rekall is designed for a single-user workstation. Any local process able to connect to the extension's pipe or socket can interact with its IPC protocol, subject to the extension's own checks. Rekall does not add a separate authentication boundary. Owner/thread checks prevent accidental misrouting; they do not protect against an untrusted process with access to the same account and IPC endpoint.
Linux uses the same private per-user socket location and ownership/permission checks as macOS. Rekall does not authenticate the peer process separately. Globally shared temporary sockets are not supported. See the historical upstream socket-isolation report #8965. Linux CI uses isolated test sockets and does not establish live extension compatibility.
Read SECURITY.md for the trust model, handoff-integrity limits, and private vulnerability reporting.
Data and privacy
Rekall does not export the transcript. It keeps the active thread snapshot in memory while processing state updates. Local jobs/ files contain thread and job identifiers, timestamps, state, errors, token counts, paths, checksums, and the handoff text the user explicitly asked it to preserve. These files are excluded from Git and npm packages. Do not publish them or include them in bug reports.
Jobs default to $CODEX_HOME/tools/rekall/jobs, or ~/.codex/tools/rekall/jobs when CODEX_HOME is unset. Current environment variables use the REKALL_ prefix. Legacy CONTEXT_COMPACT_* names remain aliases, and an existing $CODEX_HOME/tools/context-compact/jobs directory is reused so active locks and history are not lost during migration. REKALL_JOBS_DIR can select a different local journal directory for isolated use.
Verification and limits
The measured 78–90% reduction describes context tokens reclaimed in four individual runs, including one earlier user-reported run. It is not a percentage of the full model window, a latency distribution, or a guarantee for another task. The 15-minute deadline has not been validated against a representative workload distribution or very large threads.
The extension's compaction request does not accept custom instructions. preserve and discard guide the resumed model; they do not override the native compaction prompt or guarantee selective retention. Rekall does not change global Codex permissions or configuration.
Development
npm ci
npm test
npm pack --dry-runGitHub Actions runs the suite on Windows, macOS, and Linux with Node.js 20 and 22, plus package validation. Tests use dedicated pipes/sockets, temporary job directories, and child processes. They must never target a live conversation. Passing isolated tests does not establish live extension IPC support. A sandbox that prohibits Unix-socket listeners can cause listen EPERM; run the isolated suite in an environment that permits local test sockets.
The npm package uses an explicit file allowlist. Inspect npm pack --dry-run before publishing. Plugin and marketplace manifests live in .codex-plugin/plugin.json and .agents/plugins/marketplace.json; the MCP declaration is .mcp.json. The marketplace points to the plugin at the repository root.
The HOL scanner workflow uses a SHA-pinned action with reviewed scanner version 3.0.103, requires a score of at least 80 and no critical/high findings, and uploads SARIF to GitHub code scanning. It installs Cisco's skill analyzer and requires that analysis to complete. The Cisco Skill Scan badge reflects this combined workflow, including its mandatory Cisco analysis; it is not a separate workflow. Network analyzers and automatic catalog submissions are disabled. For the same local gate in an isolated scanner installation, run:
pipx install "plugin-scanner[cisco]==3.0.103"
plugin-scanner scan . --format text --cisco-skill-scan on --min-score 80 --fail-on-severity highScanner findings and optional analyzer availability are separate signals; a passing score does not establish runtime safety. Dependency updates are tracked by Dependabot, and .codexignore excludes local runtime and build artifacts without excluding source code from review.
Each successful scanner run publishes a JSON report and skill evidence artifact bound to its Git commit and the SHA-256 of skills/rekall/SKILL.md. The skill's metadata.commit identifies an immutable revision containing the same instruction body; metadata-only changes may differ. CI compares that body before recording a match. Skill tags and language use Codex's supported metadata field. Rekall does not add unsupported top-level fields or a self-declared verified flag to increase its separate Skill Trust score. Read the report's analyzer status and findings alongside any numeric rating.
This project is licensed under the MIT License.
Releases
Pushing a new version tag such as v0.3.2 triggers the Release workflow. The tag must match the versions in package.json, package-lock.json, the plugin manifest, and the skill metadata. The packaged MCP must report that same version.
The workflow runs the shared Windows/macOS/Linux CI matrix and the HOL scanner, then builds the npm archive. It checks the package allowlist, verifies a fresh installation and MCP startup, and generates SHA256SUMS.txt. GitHub Release publication happens only after those checks pass. Prerelease versions produce prereleases. npm registry publication is separate and is not enabled by this workflow.
After updating and committing all version fields, validate the intended tag and push it:
npm run release:check -- v0.3.2
git tag v0.3.2
git push origin v0.3.2Replace v0.3.2 with the new package version. Never move a published tag. The workflow must exist in the tagged commit, so it does not retroactively build older tags such as v0.3.1.
Use Run workflow on the Release workflow for a build-only check without creating a tag or publishing anything. Download its rekall-release-<version> artifact to inspect the archive, checksum, release notes, and source manifest. For a local build from a clean committed checkout, run npm run release:build; files are written to the ignored dist/ directory.
On a repeated tag run, matching published assets are left unchanged. Missing or mismatched published assets cause a failure instead of replacement. An interrupted draft can resume its missing uploads; the release becomes public only after both expected assets have been verified.
Uninstall
Remove a plugin installation with codex plugin remove rekall@rekall. For a manual installation, run codex mcp remove rekall and remove the manually copied skills/rekall directory from your Codex home.
A worker that has already started continues until it records a result or reaches its deadline. Inspect its journal before handling a stale lock; never remove a lock while its recorded process is still running.
References
Available Tools
4 toolscancel_compactionA
Cancel pending compaction or continuation dispatches only for the exact job returned by schedule_compaction or compaction_status. Already sent requests cannot be undone. Use compaction_status for observation and inspect it after cancellation.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | Exact jobId returned by schedule_compaction or compaction_status for this thread. Never infer it. | |
| threadId | Yes | Current CODEX_THREAD_ID; obtain from the shell environment. Never guess or select a different thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that only pending dispatches are cancellable, that already sent requests cannot be undone, and that the caller should inspect compaction_status afterward. These are meaningful behavioral constraints that help an agent predict the tool's effect and avoid false assumptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences deliver the core action, the key precondition, and the post-cancellation verification step. There is no filler or repetition; every sentence contributes information needed for correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two required parameters fully documented in the schema and clear usage guidance in the description, the definition is nearly complete. The absence of an output schema is partially mitigated by instructing the agent to use compaction_status for verification after cancellation, though the direct return value of cancel_compaction itself is not described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% parameter coverage with explicit descriptions for jobId and threadId, including the 'never infer it' and 'never guess' guidance. The tool description reinforces the exactness requirement but does not add new semantic meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('cancel') and resource ('pending compaction or continuation dispatches') and clearly distinguishes the action from observation tools. It also references the exact source of the job ID, leaving no ambiguity about what the tool operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use the tool only for jobs returned by schedule_compaction or compaction_status, and directs agents to use compaction_status for observation and verification after cancellation. It also flags the limitation that already sent requests cannot be undone, which is important for deciding whether cancellation is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compaction_statusARead-only
Read-only local job inspection. Run before probe_compaction or schedule_compaction to detect an unfinished job, and after scheduling or cancellation to inspect its exact outcome and metrics. It does not probe extension compatibility. scheduled, waiting_for_idle, requesting, and accepted are intermediate; completed requires a newly observed completed contextCompaction event. resumed identifies a returned continuation turn, not successful work. Missing or stale post-compaction telemetry can skip automatic continuation with insufficient_headroom.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | Yes | Current CODEX_THREAD_ID; obtain from the shell environment. Never guess or select a different thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Read-only'. It adds substantial behavioral context beyond annotations: intermediate status semantics, 'completed requires a newly observed completed contextCompaction event', 'resumed identifies a returned continuation turn, not successful work', and the telemetry-staleness caveat about insufficient_headroom. No contradiction detected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the tool's nature, the usage timing, the status semantics, and the telemetry caveat. The most important usage guidance is front-loaded, and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and lack of output schema, the description is complete: it explains when to call it, what it inspects, how to interpret statuses, and a critical failure mode. An agent has enough information to invoke it correctly and interpret its results in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, threadId, is fully documented in the schema with clear instructions about obtaining it from CODEX_THREAD_ID and not guessing. The description adds no parameter-specific meaning, but since schema coverage is 100%, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Read-only local job inspection') and clearly distinguishes itself from siblings by naming them and noting what it does not do ('does not probe extension compatibility'). An agent can tell this tool from probe_compaction and schedule_compaction without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names when to use the tool: before probe_compaction or schedule_compaction to detect an unfinished job, and after scheduling or cancellation to inspect outcome and metrics. It also names the sibling alternatives and an exclusion, making selection unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probe_compactionARead-only
Read-only preflight for an extension-owned Codex VS Code chat on Windows, macOS, or Linux. Run after compaction_status and before schedule_compaction to verify the exact owner/thread, extension version, runtime layout, public schema, and IPC access. It does not compact or prove live compaction compatibility; layout_compatible means only that the required layout was found. Standalone Codex CLI sessions are unsupported because they have no VS Code extension owner. Use compaction_status instead to inspect an existing job.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | Yes | Current CODEX_THREAD_ID; obtain from the shell environment. Never guess or select a different thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, and the description reinforces and extends this by explaining that the tool does not compact and does not prove live compaction compatibility. The layout_compatible clarification and unsupported-standalone-session note add meaningful behavioral context beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the first states the tool's purpose, the second gives usage sequence, the third clarifies limitations, and the fourth handles unsupported cases and alternatives. The description is front-loaded with the most important information and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, sequencing, limitations, and unsupported environments, and even explains the meaning of layout_compatible despite there being no output schema. The only minor gap is that it does not explicitly describe the return value or result shape beyond that one field, though the preflight framing makes this less critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already fully explains threadId including the instruction to obtain it from the environment and never guess. The description adds no meaningful parameter-level detail beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only preflight for an extension-owned Codex VS Code chat and lists the exact things it verifies: owner/thread, extension version, runtime layout, public schema, and IPC access. It also distinguishes itself by stating what it does not do and by referencing its position relative to sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit sequencing guidance: run after compaction_status and before schedule_compaction. It also states when not to use it, telling users that standalone Codex CLI sessions are unsupported and directing them to use compaction_status instead to inspect an existing job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schedule_compactionA
Schedule ONE compaction after this answer. Only when the user authorizes compaction. Optional handoff describes what to keep/summarize and is saved exactly in a per-job file. resume:true explicitly requests ONE subsequent turn with that handoff and nextStep; use only for authorized continuation. Cancels on observed new user input or a stopped turn. Return the handoff in context before finishing. Compaction prompt itself is not overridden. Check status later; no automatic retries.
| Name | Required | Description | Default |
|---|---|---|---|
| handoff | No | Optional verified handoff saved for this job. It guides the resumed agent but does not alter the extension compaction prompt. | |
| threadId | Yes | Current CODEX_THREAD_ID; obtain from the shell environment. Never guess or select a different thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses key behaviors that annotations do not convey: scheduling is one-shot, the job cancels on new user input or a stopped turn, there are no automatic retries, and the handoff is persisted to a per-job file. It also tells the agent to return handoff in context before finishing. No contradiction with readOnlyHint=false or destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Each sentence carries a distinct operational constraint: timing, authorization, handoff persistence, resume semantics, cancellation, context handoff, prompt override, and retry behavior. The most critical fact (scheduling one authorized compaction) is front-loaded, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and a complex conditional job, the description covers the full lifecycle an agent needs: when to call, what handoff means, how resume works, when it cancels, how follow-up should happen (check status), and the no-retry policy. The schema already documents parameter details and threadId sourcing, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value by explaining the runtime semantics of handoff and resume ('resume:true explicitly requests ONE subsequent turn with that handoff and nextStep') and noting the handoff is saved exactly per job, which the schema alone does not state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Schedule ONE compaction after this answer' names a specific action, target, and timing. It also states the authorization condition ('Only when the user authorizes compaction'), which separates scheduling from the sibling probe/status/cancel operations without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when the tool is appropriate: only when the user explicitly authorizes compaction and only after the answer. It also gives a precise condition for resume:true ('use only for authorized continuation') and recommends monitoring via status rather than retrying, implicitly distinguishing it from compaction_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.3.1- Changed
cancel_compaction1 field changed- added
Input schema / properties / jobId / descriptionAdded value: +"Exact jobId returned by schedule_compaction or compaction_status for this thread. Never infer it."
- Changed
schedule_compaction3 fields changed- added
Input schema / properties / handoff / descriptionAdded value: +"Optional verified handoff saved for this job. It guides the resumed agent but does not alter the extension compaction prompt." - changed
Input schema / properties / handoff / properties / nextStep / descriptionPrevious value: -"Exact authorized next action, including where to stop."New value: +"Exact authorized next action, including where to stop. Required even when resume is false." - changed
Input schema / properties / handoff / properties / resume / descriptionPrevious value: -"Explicit opt-in to ONE automatic next turn. Must be false if the user asked to stop."New value: +"Explicit opt-in to ONE automatic next turn. Use false when work is complete, the user asked to stop, or the next step needs user input."
4 tool updates
v0.3.0- First observed
cancel_compaction - First observed
compaction_status - First observed
probe_compaction - First observed
schedule_compaction
TDQS
Scored across 4 tools
Each tool covers a distinct part of the compaction lifecycle: probe checks preflight compatibility, schedule creates a job, status inspects job state, and cancel aborts pending jobs. There is no meaningful overlap, and the descriptions reinforce the correct ordering and separation of responsibilities.
Three tools follow a clear verb_noun pattern: probe_compaction, schedule_compaction, and cancel_compaction. compaction_status deviates slightly by using noun_noun instead of a verb-led form like get_compaction_status, but the overall snake_case convention and recognizable pattern keep the set highly readable.
Four tools is exactly the right scope for a compaction lifecycle management server. Each tool has a distinct role and none feel redundant or missing, making the surface compact and focused.
The tool set covers the full job lifecycle: compatibility probing, scheduling, status inspection, and cancellation. There are no obvious dead ends for the stated purpose, and the explicit sequencing guidance reduces gaps in workflow coverage.
Maintenance
Related MCP Connectors
Adaptive plan/build/review cycles for AI coding assistants, persisted across sessions.
Agent checkpoints. Resume after context resets and handoffs with retry-safe, versioned saves.
Stop re-explaining yourself to Agents. Give it the right context, right when needed.
Research-backed linting + generation for agent context files (CLAUDE.md, AGENTS.md, Cursor rules).
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables Claude agents to checkpoint their context state and reset back to saved points with handoff messages, maintaining clean context windows during complex tasks by avoiding fragmented summaries from compactions.52-
- FlicenseBqualityDmaintenanceBrings Claude Code's context compaction to any MCP host, enabling agents to gauge context pressure, summarize history, re-hydrate files, and persist rules across session boundaries.19-
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to maintain continuous project memory by automatically relaying compacted context from any MCP compactor into a persistent context log.1Apache 2.0
- FlicenseNot gradedqualityBmaintenanceEnables coding agents to preserve and recover structured task state across context loss, recording objectives, constraints, decisions, and plan, and reconstructing a bounded deterministic context block on session resume.-