OpenReflex
This server gives AI agents ambient, privacy-preserving project memory for coding and reasoning tasks: it plans tasks from past experience, monitors progress, explains decisions, and records outcomes.
Plan tasks:
get_execution_contextretrieves similar past tasks, suggests a strategy with alternatives, estimates budget, likely files, and lessons.Check progress:
check_progressrecommends continue, pivot, or stop based on success estimates, budget use, and detected loops/stalls.Explain decisions:
explain_decision,get_execution_trace,get_candidate_paths, andget_reflex_scoreshow why a recommendation was made, the decision timeline, alternative paths considered, and machine-readable scores.Choose or record paths:
choose_pathlets the agent declare the strategy it is actually following;record_outcomefinalizes a task with verified success/failure evidence and runs a path check.Search memory:
search_experiencefinds past tasks by description similarity and returns lessons;explain_nodeinspects any Experience Graph node and its relationships.Get project insights:
get_project_insightssummarizes capture volume, reuse, outcomes, efficiency comparisons, routing agreement, live alerts, and lesson count.Manage privacy/capture:
approve_projectenables capture only when explicitly requested;forget_experiencedeletes one past task's memory when explicitly asked.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@OpenReflexget execution context for refactoring the auth module"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Website · Docs · Walkthrough · PyPI · Issues · Buy me a coffee
What is OpenReflex
OpenReflex gives AI coding agents ambient project muscle memory across coding, investigation, and reasoning work. Lifecycle hooks quietly record how work actually goes, compile repeated project behaviour into reusable Reflexes, and inject only the relevant project procedure into new tasks. The model does not need to remember to call OpenReflex, and MCP diagnostics are opt-in rather than loaded into normal agent context.
You install it once and keep working normally. Core memory stays on your machine in a local SQLite database. There is no OpenReflex account or hosted service. Optional Claude token accounting uses Claude Code's local OpenTelemetry export to a loopback-only OpenReflex receiver and stores counts only.
Related MCP server: LumenCore
Why OpenReflex
It learns how the project behaves. OpenReflex compiles successful work into project-scoped Reflexes such as authentication changes, Helm configuration, database migrations or CI work. Resolution combines semantic intent, task family and module/file locality, so a novel task can reuse a project procedure without repeating an earlier task.
It catches loops while they happen. Repeated failing commands, identical retries, stalled progress, and runaway context growth raise one alert that says whether to continue, pivot to another approach, or stop and check in with you, never a stream of nags.
It learns after every task. Build work closes against tests/lint/build evidence; investigations can close against cross-checked sources; reasoning-only work can complete without external tools. A simple Path check only names a better option when comparable past tasks actually support one.
It is private by design. Only coarse, project-relative metadata is stored. File contents, commands, tool output, and transcripts never are, and capture is off until you approve a project.
Quick start
Requires Python 3.11 or newer.
Claude Code — recommended: install the plugin, reload it, inspect the current state, then explicitly enable project memory.
/plugin marketplace add vishnu-77/openreflex
/plugin install openreflex@openreflex
/reload-plugins
/openreflex
/openreflex approveClaude may display the canonical plugin namespace as /openreflex:openreflex; both refer to the same user-invoked
OpenReflex control surface when the bare alias is available. The no-argument command is read-only and compact.
approve explicitly enables local memory for the current project. The plugin prepares an exact-version runtime
under ~/.openreflex/runtime/ and normal agent work remains ambient after approval.
The Claude plugin does not execute a global openreflex from PATH. An older pipx/uv/pip installation can coexist
without taking over the plugin runtime.
CLI / other agents: use a persistent package install.
pipx install openreflex # or: uv tool install openreflex
cd your-project
openreflex install codex # or cursor / opencode / claude-codeAfter a few tasks:
openreflex context "fix the login redirect bug" # preview the context a task would receive
openreflex status # what has been captured and learned
openreflex doctor # installation checks and recent hook errors
openreflex update --check # check the latest stable releasePackage updates preserve local memory and project configuration. Managed pipx and uv tool installs can use
openreflex update, openreflex update --reinstall, or openreflex self reinstall. The Claude plugin manages
its own pinned runtime automatically; these package-manager commands are only needed for standalone CLI installs.
Use it with your agent
OpenReflex installs per project with openreflex install <agent>, or as a plugin.
Agent | Connects through | Install | Verified |
Claude Code | Plugin hooks, or project hooks |
| Live sessions |
Codex | Plugin hooks, or project hooks |
| Live sessions |
Cursor | Lifecycle hooks |
| Protocol and fuzz tests |
OpenCode | Local lifecycle plugin |
| Protocol and fuzz tests |
With the Claude plugin, /openreflex shows a compact read-only state view. Use /openreflex approve to enable
local memory for the current project, then work normally. Other explicit controls include status, reflexes,
memory, doctor, why, trace, and revoke. OpenReflex control turns are not learned as project
experiences and do not trigger normal task verification. The control skill is user-invoked only and does not enable
MCP diagnostics or participate in normal model routing. Claude may show the canonical namespaced form
/openreflex:openreflex. Per-agent guides:
Claude Code, Codex, Cursor,
OpenCode.
How it works
Hooks send lifecycle events (prompt, tool start, tool end, compaction, stop) to the OpenReflex engine, which stores them in an Experience Graph:
Task -caused-> Execution -used-> Context -used-> Experience
CandidatePath -recommended_for-> Task Lesson -recommended_for-> Task
Execution -failed_with-> ToolCall -resolved_by-> ToolCall
Execution -caused-> Outcome -caused-> Experience -caused-> LessonBefore a task: similar experiences are retrieved and OpenReflex first identifies the work mode. BUILD uses
inspect-first,test-first, andincremental; INVESTIGATE usessource-first,cross-check, andbroad-then-deep; THINK usesreason-first,compare-options, andevidence-first. Paths are ranked using past evidence, estimated cost, risk, uncertainty and reversibility. A compact context is injected only when relevant experience exists.During a task: when a detector finds a failure loop, repeated calls, stalled progress, context growth, or work past the budget, OpenReflex estimates the marginal value of more work. The current path's success estimate is updated with each call that makes no progress or fails, and compared with the cost of the work left and with the best untried alternative. The alert ends with a recommendation to continue, pivot to another strategy, or stop and ask the user. Each problem, pivot, or stop is raised once, with a cooldown between messages.
After a task: lifecycle hooks close the execution automatically from work-mode-appropriate evidence; the model does not need to call OpenReflex. Explicit MCP outcome/path operations remain available only for manual or diagnostic workflows. The Path check says either
better option: <path>when comparable completed tasks support it orbetter option: none proven. Recommendations are never presented as paths that were actually executed.
Every recommendation is stored as a decision snapshot with a Reflex Score (0-100): how strong the
recommendation is, which is separate from the estimated chance that the task succeeds. openreflex why explains
the latest decision and openreflex trace shows the timeline. In Claude Code, a short recap appears when a task
starts, when OpenReflex recommends a pivot or stop or detects trouble, and when the task completes.
Set optional limits for every task with OPENREFLEX_BUDGET, for example calls=40,minutes=20,tokens=60000.
Routing, scoring, budget and recap settings come from a versioned policy: the packaged defaults, overridden by
~/.openreflex/config.toml and then by .openreflex.toml in the project.
Normal work does not load OpenReflex MCP tools. If you explicitly want model-facing introspection for a project,
enable it with openreflex diagnostics enable <agent>. Disable it again with
openreflex diagnostics disable <agent>. The standalone MCP Registry package remains available for users who
deliberately install OpenReflex as an MCP server.
Tool | What it does |
| Plan a task: similar past tasks, the suggested strategy with alternatives, a budget, likely files, lessons. Optional |
| Whether more work on the current path is worth it: continue, pivot or stop. |
| Declare the strategy being followed when it differs from the suggestion. |
| Record a confirmed outcome (tests/builds, cross-checked research, user confirmation, or failure) and learn from it. |
| Why the latest recommendation was made: Reflex Score, signals, confidence, next-best route. |
| The decision timeline of the most recent task. |
| The latest Reflex Score and its components as JSON. |
| Past tasks in the project by description similarity, with their lessons. |
| One Experience Graph node and its relations. |
| What has been recorded, reused and learned in the project. |
| Enable capture, only when the user explicitly asks. |
| Delete one past task's memory, only when the user explicitly asks. |
CLI
Command | Purpose |
| Install the hooks-only ambient runtime and enable the project |
| Explicitly opt in/out of model-facing MCP diagnostics |
| Remove OpenReflex's integration entries; captured memory is kept |
| Check or update a managed pipx/uv-tool installation; optionally force a reinstall |
| Repair/reinstall the managed OpenReflex package while preserving memory/config |
| Remove only the managed OpenReflex package; memory/config remain on disk |
| Enable or disable capture for the current project |
| Preview the Execution Context a task would receive |
| What has been captured, reused, learned, and how model-token usage compares |
| Explain the latest recommendation, or show the decision timeline |
| Installation, project resolution, and recent hook activity checks |
| Delete the project's data |
| Opt in to local Claude Code token accounting, inspect it, or remove OpenReflex-owned telemetry settings |
| Run the simulated benchmark |
| Used by ambient agent integrations |
| Run the optional diagnostics/admin MCP server explicitly |
Privacy
Stored | Never stored |
Tool name and a coarse category ( | File contents |
A fingerprint of the arguments | Command text |
Project-relative file paths | Tool output |
Pass or fail, duration, output size | Transcripts and model output |
Optional model-token counts, model/source label, estimated cost | Prompt/response/tool contents from Claude telemetry |
A masked one-line error signature | Paths outside the project |
The prompt as a task description (up to 1,000 characters, secrets redacted) | Anything sent to an OpenReflex-hosted service: there is none |
Data lives in ~/.openreflex/projects/<hash>/experience.sqlite3. Set OPENREFLEX_HOME to move it,
OPENREFLEX_DISABLE=1 to turn capture off everywhere, openreflex forget --yes to delete a project's data, or ask
your agent to forget_experience a single task.
Research
OpenReflex is also a research project in budget-aware execution: instead of treating success as a yes or no, it studies how agents choose execution paths, spend tool calls, model tokens and context, respond to uncertainty, and whether past work makes similar future work cheaper without reducing outcome quality. The loop is experience retrieval, evidence-aware path selection, budget-aware execution, work-mode-specific completion and plain-language path comparison. The Researcher view on the site explains the idea, lets you step through one reflex forming in the graph, and places it next to related work.
The reproducible evaluation protocol, pinned real-agent corpus, exact prompts, and claim boundaries are documented in docs/evaluation.md.
Community & Contributing
Issues and ideas: GitHub Issues
Support the project: Buy me a coffee
Develop locally:
git clone https://github.com/vishnu-77/openreflex && cd openreflex pip install -e ".[dev]" pytest && ruff check src tests scripts
OpenReflex is listed in the MCP Registry as io.github.vishnu-77/openreflex and on Glama:
License
OpenReflex is released under the MIT License.
Available Tools
14 toolsapprove_projectEnable projectAIdempotent
Enable OpenReflex capture for this project.
Returns: one line confirming the project is enabled, or 'Not approved' when confirm is false.
Use when: the user explicitly asks to enable OpenReflex here. Never call it on your own initiative; the user
can also run openreflex approve, and openreflex revoke disables capture again.
Not for: anything else; every other tool answers with a 'not enabled' message until this has happened.
Side effects: writes the project's approval to the local OpenReflex home; enabling an already enabled
project changes nothing. Touches no project files.
Errors: none.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | true only when the user has explicitly asked to enable OpenReflex for this project; with false (the default) nothing changes. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses side effects ('writes the project's approval to the local OpenReflex home'), idempotence ('enabling an already enabled project changes nothing'), non-destructiveness ('Touches no project files'), and errors ('none'). These details go well beyond the annotations, which only state idempotentHint and destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with labeled sections (Returns, Use when, Never, Not for, Side effects, Errors) and every sentence carries useful information. The most important usage constraint is front-loaded, making it easy for an agent to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single optional parameter with an output schema, the description covers the return value, invocation conditions, exclusions, side effects, idempotence, and error behavior. Nothing an agent needs to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents `confirm` clearly with its default and meaning. The description adds the return behavior tied to `confirm=false`, but does not add substantial parameter semantics beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Enable OpenReflex capture for this project.' This clearly differentiates it from sibling tools like get_execution_context or record_outcome, so an agent can tell what this tool is for at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('the user explicitly asks to enable OpenReflex here'), when not to call it ('Never call it on your own initiative'), and names the CLI alternatives `openreflex approve` and `openreflex revoke`. This is exemplary routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_progressCheck progressARead-onlyIdempotent
Estimate whether more work on the current task's path is still worth it. Returns: plain text starting with 'Recommendation: continue', 'pivot' or 'stop', then the success estimates for the current path and the best alternative, the marginal value of each option, the execution budget and how much of it is used, and any detected problems (failure loop, repeated calls, stalled progress, context growth, budget overrun). Use when: unsure mid-task whether to keep going. Follow 'pivot' by switching to the named strategy; follow 'stop' by summarizing what was tried and asking the user. Not for: planning a task (use get_execution_context). Side effects: none; read-only on the most recent task. Errors: asks for get_execution_context first when no task exists; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description aligns with annotations by stating 'Side effects: none; read-only on the most recent task.' It adds valuable behavioral context beyond annotations: error behavior when no task exists, the 'not enabled' state until project approval, and the form of the returned recommendation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then organized into labeled sections: return format, usage conditions, follow-up actions, exclusions, side effects, and errors. Every sentence adds necessary information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-parameter tool with an output schema, the description covers everything an agent needs: when to call it, what the response looks like, how to react to each recommendation, and failure/error modes. This is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing for the description to document; schema coverage is vacuously 100%. The description instead explains what will happen when the tool is invoked, which is the appropriate contribution at this level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Estimate whether more work on the current task's path is still worth it.' It also identifies the unique output ('Recommendation: continue', 'pivot' or 'stop'), which clearly distinguishes it from planning tools like get_execution_context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it ('unsure mid-task whether to keep going'), what to do after 'pivot' and 'stop', and what it is not for, naming the alternative tool get_execution_context. This is complete usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
choose_pathChoose pathAIdempotent
Declare the strategy you are following for the most recent task. Returns: one line, 'Recorded chosen path: '. Use when: deliberately departing from the suggested path, so OpenReflex remembers what was actually done rather than confusing the recommendation with the observed work. Not for: tasks that follow the suggestion; the path is then inferred from tool activity. Side effects: writes the choice to the local Experience Graph; calling it again replaces the earlier choice. Touches no project files. Errors: 'No current execution' before any task was planned (call get_execution_context first); a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | No | For a custom strategy only: 2-4 short steps, e.g. ['Prototype the toml loader', 'Swap the callers']. Ignored for suggested strategies. | |
| strategy | Yes | The strategy being followed: one of the suggested names 'inspect-first', 'test-first' or 'incremental', or a short kebab-case name for your own, e.g. 'spike-then-rewrite'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that the tool writes to the local Experience Graph, replaces earlier choices on repeated calls, and touches no project files. It also lists two specific error conditions, giving the agent a clear behavioral model. The idempotentHint annotation is consistent with replacing an earlier choice with the same choice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into labeled segments — Returns, Use when, Not for, Side effects, Errors — that are easy to scan. Every sentence carries operational value, and the core purpose is front-loaded before details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers when to use the tool, what it returns, side effects, prerequisites, and failure modes. Given the rich schema and output schema, nothing essential is missing for an agent to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both strategy and steps already explained in the input schema. The description adds context about the strategy's role but no additional parameter-level meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Declare the strategy you are following for the most recent task.' It clearly distinguishes the tool from siblings like record_outcome or check_progress by framing this as the act of recording a deliberately chosen path rather than observing activity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' and 'Not for' guidance, stating it is for departing from the suggested path and not for tasks that follow the suggestion. It also references the prerequisite get_execution_context in the error section, giving the agent actionable routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_decisionExplain decisionARead-onlyIdempotent
Explain OpenReflex's latest recommendation for the most recent task. Returns: plain text with the recommended strategy, the Reflex Score (0-100: how strong the recommendation is, not the chance of success), the estimated success probability, confidence, the signals behind the score, the next-best strategy with its route advantage, and the policy version. Use when: the user asks why OpenReflex suggested a path, pivot or stop. Not for: the same numbers as structured data (use get_reflex_score) or every decision in order (use get_execution_trace). Side effects: none; read-only. Errors: 'No decision snapshot recorded yet' before any task was planned; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond annotations: it explicitly states 'Side effects: none; read-only' and discloses specific error messages ('No decision snapshot recorded yet', 'not enabled'). It also clarifies that the Reflex Score is a recommendation strength, not success probability, which is a valuable interpretive detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with labeled sections ('Returns', 'Use when', 'Not for', 'Side effects', 'Errors') that front-load the most important information. Every sentence adds distinct value, and the length is justified by the rich behavioral and error context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter input and the presence of an output schema, the description fully covers what an agent needs: what is returned, when to use it, exclusions, side effects, and error conditions. No gaps are apparent for safe and correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description need not explain parameter semantics. The schema coverage is 100% and there is nothing to add; the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Explain') and resource ('OpenReflex's latest recommendation for the most recent task'), clearly stating the tool's scope. It also distinguishes itself from siblings by explicitly naming get_reflex_score and get_execution_trace as alternatives for structured numbers or chronological decisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' and 'Not for' sections, giving clear conditions for invocation and excluding the sibling tools. It leaves no ambiguity about when to choose this tool over the alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_nodeExplain graph nodeARead-onlyIdempotent
Show one Experience Graph node with its direct relations. Returns: structured 'root', 'nodes' (Task, Execution, ToolCall, Outcome, Experience, Lesson, Context or CandidatePath and their stored fields) and typed 'edges' (used, caused, failed_with, resolved_by, recommended_for). Embeddings are omitted. Use when: tracing why a lesson or recommendation exists, after finding an id with search_experience or get_execution_context. Not for: searching by topic (use search_experience). Side effects: none; read-only. Errors: an MCP error reports 'Unknown graph node' for an id that does not exist, and an MCP error is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| node_id | Yes | Id of an Experience Graph node: an experience id from search_experience (starts with 'exp-'), the execution id from get_execution_context, or any node id from an earlier explain_node result. |
Output Schema
| Name | Required | Description |
|---|---|---|
| root | Yes | Id of the graph node that was requested. |
| edges | Yes | Direct typed relationships to or from the root node. |
| nodes | Yes | The root node and directly related nodes; embeddings are omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, and the description reinforces with 'Side effects: none; read-only.' It also adds non-redundant behavioral details: embeddings are omitted, and error behavior (unknown id yields 'Unknown graph node', and errors until project approval) is disclosed. This exceeds the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with labeled sections (Returns:, Use when:, Not for:, Side effects:, Errors:). The main purpose is front-loaded, and every section contributes either to usage or safety. Slightly verbose but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (per context signals), the description needn't enumerate return fields, but it summarizes types (root, nodes, edges) and omissions (embeddings). It also covers error conditions and usage context, making it complete for a single-parameter read-only tool. Minor gap: output schema description is not in the text, but that is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter description for node_id already explains valid id sources (experience id, execution id, or previous explain_node result). The description adds little beyond reusing this guidance, so it stays at the baseline for well-covered schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Show one Experience Graph node') and explicitly differentiates from the sibling search_experience in 'Not for: searching by topic'. The return structure (root, nodes, edges) is also outlined, making the tool's exact scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' guidance ('tracing why a lesson or recommendation exists, after finding an id with search_experience or get_execution_context') and a clear exclusion ('Not for: searching by topic (use search_experience)'), giving agents a definitive decision rule among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forget_experienceForget experienceADestructiveIdempotent
Delete one past task from this project's memory.
Returns: one line listing what was removed (task, executions, tool calls, outcomes, experience, lessons), or
'Not forgotten' when confirm is false.
Use when: the user explicitly asks to forget a specific past task, for example one that involved sensitive
work or taught a wrong lesson. Never call it on your own initiative.
Not for: deleting everything; the user runs openreflex forget --yes for that.
Side effects: permanently removes that task's nodes from the local Experience Graph; other tasks are
untouched. Touches no project files.
Errors: 'Unknown experience id' for an id that does not exist, is not an experience, or was already
forgotten; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | true only when the user has explicitly asked to forget this task; with false (the default) nothing is deleted. | |
| experience_id | Yes | The experience id to delete, as returned by search_experience; starts with 'exp-'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, and the description meaningfully enriches both: it specifies the exact destruction scope ('that task's nodes from the local Experience Graph; other tasks are untouched'), asserts side-effect boundaries ('Touches no project files'), states permanence ('permanently removes'), enumerates error modes ('Unknown experience id' for non-existent/already-forgotten; 'not enabled' until project approval), and discloses the confirm=false no-op path via the 'Not forgotten' return. All of this goes well beyond the structured annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured into clearly labeled sections (Returns, Use when, Not for, Side effects, Errors) with the core action front-loaded. Every section earns its place for a permanently destructive tool, and the labels make it highly scannable. Slight redundancy exists between 'Never call it on your own initiative' and confirm's schema description, keeping it just shy of maximal conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a permanently destructive tool, the description covers the deletion target, return format in both confirm states, every error condition including the approval prerequisite ('not enabled' until project approved), the side-effect boundary, and the user-consent requirement. With an output schema also present to carry return-structure detail, nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — experience_id is documented with its 'exp-' pattern and source (returned by search_experience), and confirm with its guardrail semantics ('true only when the user has explicitly asked'). The description reinforces confirm's default-false behavior through the 'Not forgotten' return case but does not materially extend the schema's parameter explanations. The 100%-coverage baseline of 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with 'Delete one past task from this project's memory' — a specific verb, a specific resource, and an explicit scope ('one'). The 'Not for: deleting everything' clause further carves out the boundary versus bulk deletion, and the resource ('past task', 'experience graph nodes') distinguishes it from record_outcome, search_experience, and other siblings. No tautology; the title 'Forget experience' is expanded into an operational statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit triggering condition ('Use when: the user explicitly asks to forget a specific past task') with a concrete example ('sensitive work or taught a wrong lesson'), an absolute exclusion ('Never call it on your own initiative'), and a named alternative for the excluded case (user runs `openreflex forget --yes`). This is complete routing guidance with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_candidate_pathsGet candidate pathsARead-onlyIdempotent
Show every candidate path OpenReflex considered for the most recent task side by side. Returns: plain text with each strategy's score, success probability, cost/risk numbers, and whether it was dominated; which path was recommended; and which one was actually followed - explicit (via choose_path), inferred from tool-call evidence, or not yet determined while the execution is still running. Use when: the user asks what other approaches were considered, or whether the agent followed the suggested path. Not for: a single recommendation's reasoning (use explain_decision) or the decision timeline (use get_execution_trace). Side effects: none; read-only. Errors: 'No execution recorded yet' before any task was planned; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations already declare readOnlyHint and destructiveHint false, the description adds substantial context: it details the exact content returned (score, success probability, cost/risk, domination status), explains how the followed path is determined (explicit, inferred, or undetermined), lists specific error messages ('No execution recorded yet', 'not enabled'), and states side effects. This goes far beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections: main purpose, returns, use-when, not-for, side effects, and errors. Each sentence adds unique value with no redundancy. It is appropriately detailed without being verbose, and the primary purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the tool takes no parameters, the description covers all necessary aspects: purpose, usage conditions, alternatives, errors, side effects, and return details. Nothing an agent needs to decide whether to invoke this tool and interpret its result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema fully covers parameter semantics (trivially). The baseline for 0 params is 4, and the description adds no parameter-specific information because none is needed. It correctly focuses on output and behavior instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Show'), a clear resource ('every candidate path OpenReflex considered'), and a scope ('for the most recent task'). It explicitly differentiates from siblings by naming what it is not for and pointing to alternatives, leaving no ambiguity about its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when' and 'Not for' conditions, naming the exact user intents that should trigger this tool and the sibling tools (explain_decision, get_execution_trace) that should be used instead. This is the gold standard for usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_execution_contextGet execution contextA
Plan a task from this project's past experience. Returns: plain text with the learned project Reflex procedure when one applies; otherwise a neutral working approach, relevant past experience, budget, likely files/sources and lessons. Internal strategy names and candidate scores are deliberately omitted from this normal surface; use get_candidate_paths for diagnostics. Use when: starting any substantial coding, investigation, research, analysis, review, planning or reasoning task and no [OpenReflex] block was injected. Not for: looking up history (use search_experience) or checking progress mid-task (use check_progress). Side effects: starts or re-plans the current task in the local Experience Graph; touches no project files, runs no commands, sends nothing over the network. Calling it again for the same task returns the same plan unless new limits are given. Errors: invalid non-positive limits are rejected by the input schema; a 'not enabled' message is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | The task in one or two plain sentences, e.g. 'Fix the login redirect loop after logout'. Used to find similar past tasks; secrets are redacted before it is stored. | |
| max_minutes | No | Optional cap on active working time in minutes; a positive number. Default: no cap. | |
| max_tool_calls | No | Optional cap on tool calls for this task; a positive integer. Default: no cap. | |
| max_context_tokens | No | Optional cap on tokens of tool output added to context; a positive integer. Default: no cap. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only mark non-read-only/non-destructive, which is thin; the description carries the behavioral burden. It discloses that the call starts/re-plans a task in the local Experience Graph, touches no files/commands/network, is repeatable (same plan unless new limits are given), and can return a 'not enabled' message. This significantly exceeds annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though longer than typical tool descriptions, it is broken into labeled sections (Returns, Use when, Not for, Side effects, Errors) that each add distinct value. The core purpose is front-loaded in the first sentence, and no sentence is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers every decision an agent needs: what it returns, when to use it, which siblings to prefer, side effects, repeat-call behavior, and error conditions. Combined with the output schema, nothing important is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (task, max_minutes, max_tool_calls, max_context_tokens) already has a clear schema description. The tool description adds only minor context ('unless new limits are given'), so it meets the baseline but does not meaningfully compensate beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb-object ('Plan a task from this project's past experience') and clarifies the return format. It explicitly distinguishes itself from sibling tools like get_candidate_paths, search_experience, and check_progress, so an agent can select it confidently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains an explicit 'Use when' condition (starting substantial work with no [OpenReflex] block injected) and a 'Not for' section naming alternatives (search_experience for history, check_progress for mid-task checks). This is exactly the when/when-not/alternatives guidance requested.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_execution_traceGet execution traceARead-onlyIdempotent
Return the decision timeline of the most recent task. Returns: plain text, one line per decision in order, with elapsed time, phase (start, runtime, complete), action, strategy, Reflex Score and the event that triggered it. Never includes prompts, commands or tool output. Use when: reviewing how a task unfolded. Not for: only the latest decision (use explain_decision). Side effects: none; read-only. Errors: 'No decision snapshots recorded yet' before any task was planned; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only and idempotent behavior, and the description adds valuable behavioral context: plain text output format, one line per decision, content exclusions, and specific error messages. It also confirms no side effects, complementing the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact yet well-structured with labeled sections: Returns, Use when, Not for, Side effects, and Errors. Every sentence adds useful information, and the main purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a read-only, parameterless operation, and the description covers purpose, output format, usage boundaries, side effects, and error conditions. An output schema exists, so return-value details need not be repeated. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete and the description does not need to explain parameter meaning. The baseline of 4 applies; the description adds useful output details even though parameters are absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Return the decision timeline of the most recent task.' It also distinguishes itself from explain_decision by explicitly noting it is not for only the latest decision. This makes it easy for an agent to select among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance: 'Use when: reviewing how a task unfolded' and 'Not for: only the latest decision (use explain_decision).' This clearly tells the agent when to call this tool and when to choose an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_insightsGet project insightsARead-onlyIdempotent
Summarize what OpenReflex has recorded and learned in this project. Returns: structured activation, engagement, experience reuse, outcomes, observational efficiency with and without prior experience, path-comparison trends, routing agreement, live alerts, execution-control metrics and lesson count. Efficiency comparisons are observational and are not presented as causal evidence. Use when: the user asks how OpenReflex is doing in this project or a program needs project-level metrics. Not for: individual past tasks (use search_experience). Side effects: none; read-only. Errors: an MCP error is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| lessons | Yes | Number of extracted lessons retained in this project. |
| project | Yes | Local project path represented by this Experience Graph. |
| routing | Yes | Recommendation agreement with retrospective realised performance. |
| outcomes | Yes | Known and verified outcome coverage. |
| activation | Yes | Activation and first-use metrics. |
| engagement | Yes | Capture and usage volume. |
| path_check | Yes | Simple evidence-backed path comparison counts. |
| live_alerts | Yes | Counts of loop, repetition, stagnation, context and budget alerts. |
| experience_reuse | Yes | How often prior experience was reused. |
| project_reflexes | Yes | Project-specific reusable procedures learned from execution evidence. |
| execution_control | Yes | Runtime verdict and budget-adherence metrics. |
| efficiency_observational | Yes | Observed efficiency with versus without reused experience; this is not a controlled comparison. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior, and the description reinforces this with 'Side effects: none; read-only.' It adds genuinely useful context beyond annotations: the error condition until project approval and the caveat that efficiency comparisons are observational, not causal evidence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with labeled sections, front-loads the core purpose, and every section earns its place: purpose, return content, usage conditions, exclusions, side effects, and errors. The length is justified given the breadth of insights returned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This tool has no parameters and a rich output schema, and the description covers everything needed for correct invocation: when to use it, when not to, the read-only guarantee, the project-approval error condition, and the causal-evidence caveat. No critical context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is no schema gap for the description to compensate for. The baseline for a zero-parameter tool is 4, and the description appropriately focuses on what the tool returns and when to use it rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Summarize what OpenReflex has recorded and learned in this project.' It also distinguishes itself from search_experience by clarifying it covers project-level insights, not individual past tasks. The detailed Returns list further disambiguates the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Use when' and 'Not for' sections give clear invocation conditions and explicitly name search_experience as the alternative for individual past tasks. This gives an agent direct routing guidance with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_reflexesGet project ReflexesARead-onlyIdempotent
Show the project Reflex and specialised procedures OpenReflex is learning from repeated work. Returns: the root Project Reflex after the first captured experience, plus specialised Reflexes with learning/learned/proven/stale maturity, support/verification counts and procedure steps. Use when: the user asks what OpenReflex has learned specifically about this project, which reusable procedures exist, or how a named Reflex works. Not for: backend routing internals or generic strategies (use get_candidate_paths for those diagnostics). Side effects: may refresh derived local Reflex data from already-captured evidence; touches no project files. Errors: a 'not enabled' message is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds valuable context: a potential refresh side effect (non-destructive) and an error condition ('not enabled' until approval). This goes beyond the structured annotations and clarifies behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (Returns, Use when, Not for, Side effects, Errors) and every sentence adds value. It is front-loaded with the primary purpose and remains concise despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter tool and the presence of an output schema, the description provides sufficient context: it describes what is returned, when to use, exclusions, side effects, and error conditions. Nothing critical is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline for parameter semantics is 4. The description correctly avoids any parameter explanations since none exist, and the schema coverage is 100% (trivially).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool shows project Reflexes and specialised procedures, with a specific verb and resource. It explicitly differentiates from sibling get_candidate_paths by stating what it is not for, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit 'Use when' scenarios and a 'Not for' section that names the alternative tool (get_candidate_paths). This fully guides an agent on when to select this tool versus others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reflex_scoreGet Reflex ScoreARead-onlyIdempotent
Return the latest decision's Reflex Score and machine-readable components. Returns: structured fields for availability, Reflex Score (0-100 recommendation strength, not success probability), success probability, confidence, strategy, next-best strategy, route advantage, evidence, context cost, budget use, component signals and policy version. If no decision exists, available=false and reason explains why. Use when: a program or agent needs decision numbers it can inspect without parsing prose. Not for: a readable explanation (use explain_decision) or the full decision timeline (use get_execution_trace). Side effects: none; read-only. Errors: an MCP error is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| reason | No | Why no score is available, when available is false. |
| signals | No | Normalised component signals used to compute the Reflex Score. |
| strategy | No | Recommended strategy for the latest decision. |
| available | Yes | Whether a decision snapshot is available for the most recent decided task. |
| budget_used | No | Largest fraction of the task budget consumed across time, calls and context. |
| reflex_score | No | Strength of the recommendation on a 0-100 scale; not success probability. |
| context_tokens | No | Estimated tokens injected from prior experience for this task. |
| evidence_count | No | Number of relevant prior experiences supporting the decision. |
| policy_version | No | OpenReflex decision-policy version used for the score. |
| route_advantage | No | Normalised advantage of the recommended route over the next best route. |
| next_best_strategy | No | Highest-ranked alternative strategy, if any. |
| decision_confidence | No | Confidence in the recommendation from the available evidence. |
| success_probability | No | Estimated probability of success for the recommended strategy. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and non-destructive, and the description adds value by stating side effects are none, errors occur until project approval, and availability behavior when no decision exists. It also clarifies that the Reflex Score is recommendation strength, not success probability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into purpose, return contents, usage guidance, exclusions, side effects, and errors. Each section earns its place, and the core purpose is front-loaded before the detailed field list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters and an output schema present, the description still covers availability, error behavior, side effects, and sibling-tool distinctions. Nothing needed to invoke or interpret the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter burden for the description to carry. The description instead clarifies the output contract, which is the relevant semantic content for this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb 'Return', the resource 'latest decision's Reflex Score', and the machine-readable component fields. It explicitly distinguishes itself from explain_decision and get_execution_trace, so an agent can identify its unique purpose without inspecting siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'Use when' guidance for programmatic inspection of decision numbers, and 'Not for' guidance with named alternatives for prose explanations and full timelines. This leaves no ambiguity about when to select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_outcomeRecord outcomeAIdempotent
Record the verified outcome of the most recent task and learn from it. Returns: the recorded status plus a plain-language Path check. A better path is only named when comparable completed tasks provide evidence; otherwise the result says that no better option is proven. Use when: the result is confirmed: tests/lint/build passed, research was cross-checked against relevant evidence, the user confirmed the answer, or the task failed/was abandoned. Not for: declaring the strategy (use choose_path). Side effects: finalizes the task in the local Experience Graph and updates its experience, lessons and path comparison; calling it again for the same task replaces the recorded outcome. Touches no project files. Errors: 'No execution to record an outcome for' before any task was planned; a 'not enabled' message until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | 'success' when the result is confirmed (tests/builds passed, research was cross-checked, or the user confirmed it); 'failure' when the task failed or was abandoned. | |
| evidence | Yes | Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Secrets are redacted. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already note idempotentHint=true and destructiveHint=false, and the description adds meaningful context without contradicting them: it finalizes the task in the local Experience Graph, updates experience/lessons/path comparison, replaces prior recorded outcomes on repeat calls, and touches no project files. It also discloses error conditions like 'No execution to record an outcome for' and the 'not enabled' project approval state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then organized into clear labeled sections: Returns, Use when, Not for, Side effects, and Errors. Despite being detailed, every sentence earns its place and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return details are partially covered, and the description adds the necessary decision context, side effects, error conditions, and prerequisites. It tells an agent when to call it, what will happen, what will not happen, and what errors to expect, making it fully actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already fully explains status and evidence, including the status enum semantics. The description reinforces the status meaning in the 'Use when' section but adds no new parameter-level syntax or format details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Record the verified outcome of the most recent task and learn from it.' The 'Not for' line explicitly contrasts it with choose_path, making the tool's scope unambiguous and differentiating it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when' conditions with concrete examples such as tests/lint/build passing, research cross-checked, user confirmation, or task failure/abandonment. It also gives an explicit exclusion, 'Not for: declaring the strategy (use choose_path),' which is clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_experienceSearch experienceARead-onlyIdempotent
Search this project's past tasks by description similarity. Returns: structured 'experiences' (id, score, description, class, agent, strategy, status, tool_calls, minutes, files), best match first, plus structured 'lessons' (text, confidence, support). Both lists are empty when nothing is similar enough. Use when: the user asks what was learned or what worked before, or to find an experience id for explain_node or forget_experience. Not for: planning a new task (use get_execution_context). Side effects: none; read-only. Matching is lexical, on task descriptions only. Errors: invalid limit values are rejected by the input schema; an MCP error is returned until the project is approved.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of past tasks to return. Default 5; allowed 1-20. | |
| query | Yes | Words describing the topic, e.g. 'expired token login bug'. Matched lexically against past task descriptions. |
Output Schema
| Name | Required | Description |
|---|---|---|
| lessons | Yes | Lessons derived from the returned experiences. |
| experiences | Yes | Matching past tasks ordered from best to worst match. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds meaningful behavior beyond the annotations: describes return ordering, empty-list behavior, lexical matching scope, side-effect-free read-only nature, and error conditions including the approval requirement. This is rich contextual information that helps the agent predict tool behavior accurately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with labeled sections: Returns, Use when, Not for, Side effects, Errors. Every sentence provides distinct value, and the most important usage guidance is front-loaded near the top.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with a fully documented input schema, rich annotations, and an output schema, the description covers purpose, usage scenarios, exclusions, side effects, return behavior, and error cases. Nothing critical for an agent to select and invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema descriptions cover 100% of parameters, so the description is not required to re-document them. It adds minor context about lexical matching and validation errors, but the schema already explains query and limit adequately, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search this project's past tasks by description similarity.' It clearly identifies what the tool does and distinguishes itself from siblings by naming get_execution_context as the alternative for planning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'Use when' conditions: asking what was learned or worked before, or finding an experience id for explain_node or forget_experience. It also states what it is not for and names the alternative tool, leaving no ambiguity about when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.9.6- Changed
get_project_insights3 fields changed- added
Output schema / $defs / ProjectReflexMetricsAdded value: +{ + "properties": { + "credit_signals": { + "description": "Action-level execution-credit signals stored across surfaced Reflexes.", + "minimum": 0, + "title": "Credit Signals", + "type": "integer" + }, + "credit_spine_actions": { + "description": "Action signals with enough repeated evidence to participate in compressed Reflex procedures.", + "minimum": 0, + "title": "Credit Spine Actions", + "type": "integer" + }, + "learned": { + "description": "Specialised Reflexes promoted after repeated successful executions.", + "minimum": 0, + "title": "Learned", + "type": "integer" + }, + "learning": { + "description": "Specialised Reflexes still accumulating evidence.", + "minimum": 0, + "title": "Learning", + "type": "integer" + }, + "project_state": { + "description": "Current maturity of the root Project Reflex.", + "enum": [ + "cold", + "learning", + "learned", + "proven", + "stale" + ], + "title": "Project State", + "type": "string" + }, + "project_support": { + "description": "Captured experiences supporting the root Project Reflex.", + "minimum": 0, + "title": "Project Support", + "type": "integer" + }, + "proven": { + "description": "Specialised Reflexes backed by repeated explicitly verified evidence.", + "minimum": 0, + "title": "Proven", + "type": "integer" + }, + "specialised_total": { + "description": "Number of specialised area/procedure Reflexes.", + "minimum": 0, + "title": "Specialised Total", + "type": "integer" + }, + "stale": { + "description": "Previously learned Reflexes contradicted by newer evidence.", + "minimum": 0, + "title": "Stale", + "type": "integer" + }, + "visible": { + "description": "Project and specialised Reflexes currently surfaced.", + "minimum": 0, + "title": "Visible", + "type": "integer" + } + }, + "required": [ + "visible", + "project_state", + "project_support", + "specialised_total", + "learning", + "learned", + "proven", + "stale", + "credit_signals", + "credit_spine_actions" + ], + "title": "ProjectReflexMetrics", + "type": "object" +} - added
Output schema / properties / project_reflexesAdded value: +{ + "$ref": "#/$defs/ProjectReflexMetrics", + "description": "Project-specific reusable procedures learned from execution evidence." +} - changed
Output schema / requiredPrevious value: -[ - "project", - "activation", - "engagement", - "experience_reuse", - "outcomes", - "efficiency_observational", - "path_check", - "routing", - "live_alerts", - "execution_control", - "lessons" -]New value: +[ + "project", + "activation", + "engagement", + "experience_reuse", + "outcomes", + "efficiency_observational", + "path_check", + "routing", + "live_alerts", + "execution_control", + "project_reflexes", + "lessons" +]
- Added
get_project_reflexes
1 tool update
v0.5.3- Added
get_candidate_paths
8 tool updates
v0.5.1- Changed
choose_path2 fields changed- added
Input schema / properties / strategy / maxLengthAdded value: +60 - added
Input schema / properties / strategy / minLengthAdded value: +1
- Changed
explain_node9 fields changed- added
Input schema / properties / node_id / maxLengthAdded value: +160 - added
Input schema / properties / node_id / minLengthAdded value: +1 - added
Output schema / $defsAdded value: +{ + "GraphEdge": { + "properties": { + "relation": { + "description": "Typed relationship between the source and target nodes.", + "enum": [ + "used", + "caused", + "failed_with", + "resolved_by", + "recommended_for" + ], + "title": "Relation", + "type": "string" + }, + "source": { + "description": "Source node id.", + "title": "Source", + "type": "string" + }, + "target": { + "description": "Target node id.", + "title": "Target", + "type": "string" + } + }, + "required": [ + "source", + "relation", + "target" + ], + "title": "GraphEdge", + "type": "object" + }, + "GraphNode": { + "additionalProperties": true, + "properties": { + "id": { + "description": "Stable node id within this project's local Experience Graph.", + "title": "Id", + "type": "string" + }, + "kind": { + "description": "Experience Graph node type.", + "enum": [ + "Task", + "Context", + "CandidatePath", + "Execution", + "ToolCall", + "Outcome", + "Experience", + "Lesson" + ], + "title": "Kind", + "type": "string" + } + }, + "required": [ + "kind", + "id" + ], + "title": "GraphNode", + "type": "object" + } +} - added
Output schema / properties / edgesAdded value: +{ + "description": "Direct typed relationships to or from the root node.", + "items": { + "$ref": "#/$defs/GraphEdge" + }, + "title": "Edges", + "type": "array" +} - added
Output schema / properties / nodesAdded value: +{ + "description": "The root node and directly related nodes; embeddings are omitted.", + "items": { + "$ref": "#/$defs/GraphNode" + }, + "title": "Nodes", + "type": "array" +} - removed
Output schema / properties / resultRemoved value: -{ - "title": "Result", - "type": "string" -} - added
Output schema / properties / rootAdded value: +{ + "description": "Id of the graph node that was requested.", + "title": "Root", + "type": "string" +} - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "root", + "nodes", + "edges" +] - changed
Output schema / titlePrevious value: -"explain_nodeOutput"New value: +"ExplainNodeResult"
- Changed
forget_experience3 fields changed- added
Input schema / properties / experience_id / maxLengthAdded value: +160 - added
Input schema / properties / experience_id / minLengthAdded value: +5 - added
Input schema / properties / experience_id / patternAdded value: +"^exp-"
- Changed
get_execution_context5 fields changed- changed
Input schema / properties / max_context_tokens / anyOfPrevious value: -[ - { - "type": "integer" - }, - { - "type": "null" - } -]New value: +[ + { + "exclusiveMinimum": 0, + "type": "integer" + }, + { + "type": "null" + } +] - changed
Input schema / properties / max_minutes / anyOfPrevious value: -[ - { - "type": "number" - }, - { - "type": "null" - } -]New value: +[ + { + "exclusiveMinimum": 0, + "type": "number" + }, + { + "type": "null" + } +] - changed
Input schema / properties / max_tool_calls / anyOfPrevious value: -[ - { - "type": "integer" - }, - { - "type": "null" - } -]New value: +[ + { + "exclusiveMinimum": 0, + "type": "integer" + }, + { + "type": "null" + } +] - added
Input schema / properties / task / maxLengthAdded value: +1000 - added
Input schema / properties / task / minLengthAdded value: +1
- Changed
get_project_insights15 fields changed- added
Output schema / $defsAdded value: +{ + "ActivationMetrics": { + "properties": { + "approved_at": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Unix timestamp when capture was approved for this project.", + "title": "Approved At" + }, + "first_session_captured": { + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "description": "Whether the first observed session produced an experience.", + "title": "First Session Captured" + }, + "seconds_to_first_task": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Seconds from approval to the first captured task.", + "title": "Seconds To First Task" + } + }, + "required": [ + "approved_at", + "seconds_to_first_task", + "first_session_captured" + ], + "title": "ActivationMetrics", + "type": "object" + }, + "EfficiencyObservational": { + "properties": { + "model_token_change": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Relative mean model-token change when real telemetry exists; observational, not causal.", + "title": "Model Token Change" + }, + "time_change": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Relative mean active-time change; observational, not causal.", + "title": "Time Change" + }, + "tool_call_change": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Relative mean tool-call change; observational, not causal.", + "title": "Tool Call Change" + }, + "with_prior_experience": { + "$ref": "#/$defs/EfficiencyProfile", + "description": "Observed metrics for tasks that reused prior experience." + }, + "without_prior_experience": { + "$ref": "#/$defs/EfficiencyProfile", + "description": "Observed metrics for tasks without prior experience." + } + }, + "required": [ + "with_prior_experience", + "without_prior_experience", + "tool_call_change", + "model_token_change", + "time_change" + ], + "title": "EfficiencyObservational", + "type": "object" + }, + "EfficiencyProfile": { + "properties": { + "known_outcomes": { + "description": "Number of known outcomes in the group.", + "minimum": 0, + "title": "Known Outcomes", + "type": "integer" + }, + "minutes": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Mean active execution time in minutes.", + "title": "Minutes" + }, + "model_tokens": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Mean real model-token usage when Claude telemetry is available.", + "title": "Model Tokens" + }, + "n": { + "description": "Number of experiences in this observational group.", + "minimum": 0, + "title": "N", + "type": "integer" + }, + "success_rate": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Success rate among known outcomes in the group.", + "title": "Success Rate" + }, + "tool_calls": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Mean tool calls in the group.", + "title": "Tool Calls" + }, + "tool_output_tokens_estimate": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Mean legacy estimate of tool-output tokens.", + "title": "Tool Output Tokens Estimate" + } + }, + "required": [ + "n", + "tool_calls", + "model_tokens", + "tool_output_tokens_estimate", + "minutes", + "success_rate", + "known_outcomes" + ], + "title": "EfficiencyProfile", + "type": "object" + }, + "EngagementMetrics": { + "properties": { + "active_week_ratio": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Share of elapsed weeks containing activity.", + "title": "Active Week Ratio" + }, + "active_weeks": { + "description": "Distinct weeks containing captured executions.", + "minimum": 0, + "title": "Active Weeks", + "type": "integer" + }, + "agents": { + "description": "Agents observed in this project.", + "items": { + "type": "string" + }, + "title": "Agents", + "type": "array" + }, + "executions": { + "description": "Total execution records.", + "minimum": 0, + "title": "Executions", + "type": "integer" + }, + "experiences": { + "description": "Completed executions retained as reusable experiences.", + "minimum": 0, + "title": "Experiences", + "type": "integer" + }, + "substantial_tasks": { + "description": "Tasks eligible for planning and experience retrieval.", + "minimum": 0, + "title": "Substantial Tasks", + "type": "integer" + }, + "tasks": { + "description": "Total captured tasks, including non-substantial tasks.", + "minimum": 0, + "title": "Tasks", + "type": "integer" + }, + "weeks_since_start": { + "description": "Weeks elapsed since activation or first task.", + "minimum": 1, + "title": "Weeks Since Start", + "type": "integer" + } + }, + "required": [ + "tasks", + "substantial_tasks", + "executions", + "experiences", + "active_weeks", + "weeks_since_start", + "active_week_ratio", + "agents" + ], + "title": "EngagementMetrics", + "type": "object" + }, + "ExecutionControlMetrics": { + "properties": { + "success_after_pivot_or_stop": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Observed success rate after pivot/stop advice; not causal.", + "title": "Success After Pivot Or Stop" + }, + "tasks_within_tool_call_budget": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Share of measured tasks finishing within the tool-call budget.", + "title": "Tasks Within Tool Call Budget" + }, + "verdicts": { + "additionalProperties": { + "type": "integer" + }, + "description": "Counts of runtime continue, pivot and stop verdicts.", + "title": "Verdicts", + "type": "object" + } + }, + "required": [ + "verdicts", + "success_after_pivot_or_stop", + "tasks_within_tool_call_budget" + ], + "title": "ExecutionControlMetrics", + "type": "object" + }, + "ExperienceReuseMetrics": { + "properties": { + "benefit_rate": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Deprecated compatibility alias of reuse_rate; it does not measure causal benefit.", + "title": "Benefit Rate" + }, + "reuse_rate": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Share of substantial tasks that received relevant prior experience; this measures coverage, not benefit.", + "title": "Reuse Rate" + }, + "tasks_with_prior_experience": { + "description": "Number of substantial tasks that received prior experience.", + "minimum": 0, + "title": "Tasks With Prior Experience", + "type": "integer" + } + }, + "required": [ + "reuse_rate", + "benefit_rate", + "tasks_with_prior_experience" + ], + "title": "ExperienceReuseMetrics", + "type": "object" + }, + "OutcomeMetrics": { + "properties": { + "known": { + "description": "Experiences whose outcome is success or failure rather than unknown.", + "minimum": 0, + "title": "Known", + "type": "integer" + }, + "success_rate": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Success rate among experiences with known outcomes.", + "title": "Success Rate" + }, + "verified": { + "description": "Outcomes explicitly verified by the agent or user.", + "minimum": 0, + "title": "Verified", + "type": "integer" + } + }, + "required": [ + "known", + "verified", + "success_rate" + ], + "title": "OutcomeMetrics", + "type": "object" + }, + "PathCheckMetrics": { + "properties": { + "better_option_found": { + "description": "Tasks where comparable past evidence indicated a better option.", + "minimum": 0, + "title": "Better Option Found", + "type": "integer" + }, + "by_class": { + "additionalProperties": { + "type": "integer" + }, + "description": "Number of evidence-backed path checks by task class.", + "title": "By Class", + "type": "object" + }, + "comparisons": { + "description": "Completed tasks with enough comparable past evidence to check another path.", + "minimum": 0, + "title": "Comparisons", + "type": "integer" + } + }, + "required": [ + "comparisons", + "better_option_found", + "by_class" + ], + "title": "PathCheckMetrics", + "type": "object" + }, + "RoutingMetrics": { + "properties": { + "agreement": { + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "description": "Share of comparable executions where the recommendation matched retrospective best.", + "title": "Agreement" + }, + "compared_executions": { + "description": "Number of executions eligible for routing-agreement comparison.", + "minimum": 0, + "title": "Compared Executions", + "type": "integer" + }, + "retrospective_best": { + "additionalProperties": { + "type": "string" + }, + "description": "Best realised strategy per task class when enough evidence exists.", + "title": "Retrospective Best", + "type": "object" + } + }, + "required": [ + "retrospective_best", + "agreement", + "compared_executions" + ], + "title": "RoutingMetrics", + "type": "object" + } +} - added
Output schema / properties / activationAdded value: +{ + "$ref": "#/$defs/ActivationMetrics", + "description": "Activation and first-use metrics." +} - added
Output schema / properties / efficiency_observationalAdded value: +{ + "$ref": "#/$defs/EfficiencyObservational", + "description": "Observed efficiency with versus without reused experience; this is not a controlled comparison." +} - added
Output schema / properties / engagementAdded value: +{ + "$ref": "#/$defs/EngagementMetrics", + "description": "Capture and usage volume." +} - added
Output schema / properties / execution_controlAdded value: +{ + "$ref": "#/$defs/ExecutionControlMetrics", + "description": "Runtime verdict and budget-adherence metrics." +} - added
Output schema / properties / experience_reuseAdded value: +{ + "$ref": "#/$defs/ExperienceReuseMetrics", + "description": "How often prior experience was reused." +} - added
Output schema / properties / lessonsAdded value: +{ + "description": "Number of extracted lessons retained in this project.", + "minimum": 0, + "title": "Lessons", + "type": "integer" +} - added
Output schema / properties / live_alertsAdded value: +{ + "additionalProperties": { + "type": "integer" + }, + "description": "Counts of loop, repetition, stagnation, context and budget alerts.", + "title": "Live Alerts", + "type": "object" +} - added
Output schema / properties / outcomesAdded value: +{ + "$ref": "#/$defs/OutcomeMetrics", + "description": "Known and verified outcome coverage." +} - added
Output schema / properties / path_checkAdded value: +{ + "$ref": "#/$defs/PathCheckMetrics", + "description": "Simple evidence-backed path comparison counts." +} - added
Output schema / properties / projectAdded value: +{ + "description": "Local project path represented by this Experience Graph.", + "title": "Project", + "type": "string" +} - removed
Output schema / properties / resultRemoved value: -{ - "title": "Result", - "type": "string" -} - added
Output schema / properties / routingAdded value: +{ + "$ref": "#/$defs/RoutingMetrics", + "description": "Recommendation agreement with retrospective realised performance." +} - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "project", + "activation", + "engagement", + "experience_reuse", + "outcomes", + "efficiency_observational", + "path_check", + "routing", + "live_alerts", + "execution_control", + "lessons" +] - changed
Output schema / titlePrevious value: -"get_project_insightsOutput"New value: +"ProjectInsightsResult"
- Changed
get_reflex_score16 fields changed- added
Output schema / properties / availableAdded value: +{ + "description": "Whether a decision snapshot is available for the most recent decided task.", + "title": "Available", + "type": "boolean" +} - added
Output schema / properties / budget_usedAdded value: +{ + "anyOf": [ + { + "minimum": 0, + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Largest fraction of the task budget consumed across time, calls and context.", + "title": "Budget Used" +} - added
Output schema / properties / context_tokensAdded value: +{ + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Estimated tokens injected from prior experience for this task.", + "title": "Context Tokens" +} - added
Output schema / properties / decision_confidenceAdded value: +{ + "anyOf": [ + { + "maximum": 1, + "minimum": 0, + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Confidence in the recommendation from the available evidence.", + "title": "Decision Confidence" +} - added
Output schema / properties / evidence_countAdded value: +{ + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Number of relevant prior experiences supporting the decision.", + "title": "Evidence Count" +} - added
Output schema / properties / next_best_strategyAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Highest-ranked alternative strategy, if any.", + "title": "Next Best Strategy" +} - added
Output schema / properties / policy_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "OpenReflex decision-policy version used for the score.", + "title": "Policy Version" +} - added
Output schema / properties / reasonAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Why no score is available, when available is false.", + "title": "Reason" +} - added
Output schema / properties / reflex_scoreAdded value: +{ + "anyOf": [ + { + "maximum": 100, + "minimum": 0, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Strength of the recommendation on a 0-100 scale; not success probability.", + "title": "Reflex Score" +} - removed
Output schema / properties / resultRemoved value: -{ - "title": "Result", - "type": "string" -} - added
Output schema / properties / route_advantageAdded value: +{ + "anyOf": [ + { + "maximum": 1, + "minimum": 0, + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Normalised advantage of the recommended route over the next best route.", + "title": "Route Advantage" +} - added
Output schema / properties / signalsAdded value: +{ + "additionalProperties": { + "type": "number" + }, + "description": "Normalised component signals used to compute the Reflex Score.", + "title": "Signals", + "type": "object" +} - added
Output schema / properties / strategyAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Recommended strategy for the latest decision.", + "title": "Strategy" +} - added
Output schema / properties / success_probabilityAdded value: +{ + "anyOf": [ + { + "maximum": 1, + "minimum": 0, + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Estimated probability of success for the recommended strategy.", + "title": "Success Probability" +} - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "available" +] - changed
Output schema / titlePrevious value: -"get_reflex_scoreOutput"New value: +"ReflexScoreResult"
- Changed
record_outcome4 fields changed- changed
Input schema / properties / evidence / descriptionPrevious value: -"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Up to 500 characters; secrets are redacted."New value: +"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Secrets are redacted." - added
Input schema / properties / evidence / maxLengthAdded value: +500 - added
Input schema / properties / evidence / minLengthAdded value: +1 - changed
Input schema / properties / status / descriptionPrevious value: -"'success' when tests, lint or build passed or the user confirmed the result; 'failure' when the task failed or was abandoned."New value: +"'success' when the result is confirmed (tests/builds passed, research was cross-checked, or the user confirmed it); 'failure' when the task failed or was abandoned."
- Changed
search_experience11 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum number of past tasks to return, 1-20; values outside the range are clamped. Default 5."New value: +"Maximum number of past tasks to return. Default 5; allowed 1-20." - added
Input schema / properties / limit / maximumAdded value: +20 - added
Input schema / properties / limit / minimumAdded value: +1 - added
Input schema / properties / query / maxLengthAdded value: +1000 - added
Input schema / properties / query / minLengthAdded value: +1 - added
Output schema / $defsAdded value: +{ + "ExperienceSummary": { + "properties": { + "agent": { + "description": "Agent that produced the experience.", + "title": "Agent", + "type": "string" + }, + "class": { + "description": "OpenReflex task class.", + "title": "Class", + "type": "string" + }, + "description": { + "description": "Redacted task description stored for the past execution.", + "title": "Description", + "type": "string" + }, + "files": { + "description": "Project-relative files associated with the execution, capped for compactness.", + "items": { + "type": "string" + }, + "title": "Files", + "type": "array" + }, + "id": { + "description": "Experience Graph id for the past task.", + "title": "Id", + "type": "string" + }, + "minutes": { + "description": "Active execution time in minutes.", + "minimum": 0, + "title": "Minutes", + "type": "number" + }, + "score": { + "description": "Similarity/retrieval score for this query; higher ranks first.", + "title": "Score", + "type": "number" + }, + "status": { + "description": "Observed outcome status: success, failure or unknown.", + "title": "Status", + "type": "string" + }, + "strategy": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Strategy used for the past task when known.", + "title": "Strategy" + }, + "tool_calls": { + "description": "Number of captured tool calls in the execution.", + "minimum": 0, + "title": "Tool Calls", + "type": "integer" + } + }, + "required": [ + "id", + "score", + "description", + "class", + "agent", + "strategy", + "status", + "tool_calls", + "minutes", + "files" + ], + "title": "ExperienceSummary", + "type": "object" + }, + "LessonSummary": { + "properties": { + "confidence": { + "description": "Combined confidence in this lesson.", + "maximum": 1, + "minimum": 0, + "title": "Confidence", + "type": "number" + }, + "support": { + "description": "Number of experiences supporting the lesson.", + "minimum": 1, + "title": "Support", + "type": "integer" + }, + "text": { + "description": "A compact lesson extracted from one or more past executions.", + "title": "Text", + "type": "string" + } + }, + "required": [ + "text", + "confidence", + "support" + ], + "title": "LessonSummary", + "type": "object" + } +} - added
Output schema / properties / experiencesAdded value: +{ + "description": "Matching past tasks ordered from best to worst match.", + "items": { + "$ref": "#/$defs/ExperienceSummary" + }, + "title": "Experiences", + "type": "array" +} - added
Output schema / properties / lessonsAdded value: +{ + "description": "Lessons derived from the returned experiences.", + "items": { + "$ref": "#/$defs/LessonSummary" + }, + "title": "Lessons", + "type": "array" +} - removed
Output schema / properties / resultRemoved value: -{ - "title": "Result", - "type": "string" -} - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "experiences", + "lessons" +] - changed
Output schema / titlePrevious value: -"search_experienceOutput"New value: +"SearchExperienceResult"
13 tool updates
v0.3.2- Changed
approve_project1 field changed- added
Input schema / properties / confirm / descriptionAdded value: +"true only when the user has explicitly asked to enable OpenReflex for this project; with false (the default) nothing changes."
- Added
check_progress - Changed
choose_path2 fields changed- added
Input schema / properties / steps / descriptionAdded value: +"For a custom strategy only: 2-4 short steps, e.g. ['Prototype the toml loader', 'Swap the callers']. Ignored for suggested strategies." - added
Input schema / properties / strategy / descriptionAdded value: +"The strategy being followed: one of the suggested names 'inspect-first', 'test-first' or 'incremental', or a short kebab-case name for your own, e.g. 'spike-then-rewrite'."
- Added
explain_decision - Changed
explain_node1 field changed- added
Input schema / properties / node_id / descriptionAdded value: +"Id of an Experience Graph node: an experience id from search_experience (starts with 'exp-'), the execution id from get_execution_context, or any node id from an earlier explain_node result."
- Added
forget_experience - Changed
get_execution_context4 fields changed- added
Input schema / properties / max_context_tokensAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional cap on tokens of tool output added to context; a positive integer. Default: no cap.", + "title": "Max Context Tokens" +} - added
Input schema / properties / max_minutesAdded value: +{ + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional cap on active working time in minutes; a positive number. Default: no cap.", + "title": "Max Minutes" +} - added
Input schema / properties / max_tool_callsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional cap on tool calls for this task; a positive integer. Default: no cap.", + "title": "Max Tool Calls" +} - added
Input schema / properties / task / descriptionAdded value: +"The task in one or two plain sentences, e.g. 'Fix the login redirect loop after logout'. Used to find similar past tasks; secrets are redacted before it is stored."
- Added
get_execution_trace - Added
get_project_insights - Added
get_reflex_score - Removed
project_insights - Changed
record_outcome2 fields changed- added
Input schema / properties / evidence / descriptionAdded value: +"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Up to 500 characters; secrets are redacted." - added
Input schema / properties / status / descriptionAdded value: +"'success' when tests, lint or build passed or the user confirmed the result; 'failure' when the task failed or was abandoned."
- Changed
search_experience2 fields changed- added
Input schema / properties / limit / descriptionAdded value: +"Maximum number of past tasks to return, 1-20; values outside the range are clamped. Default 5." - added
Input schema / properties / query / descriptionAdded value: +"Words describing the topic, e.g. 'expired token login bug'. Matched lexically against past task descriptions."
7 tool updates
v0.1.3- First observed
approve_project - First observed
choose_path - First observed
explain_node - First observed
get_execution_context - First observed
project_insights - First observed
record_outcome - First observed
search_experience
TDQS
Scored across 14 tools
Each tool targets a distinct action or view, and close pairs like explain_decision vs get_reflex_score are clearly separated by output format and explicit 'Not for' cross-references. The tools are easy for an agent to tell apart despite several operating on the same recent-task context.
All 14 tools follow a consistent snake_case verb_noun pattern: get_, explain_, search_, check_, choose_, record_, approve_, forget_. No mixed casing or stylistic drift is present.
At 14 tools, the set is well within the ideal range and each tool has a clear role in the OpenReflex workflow. The count feels appropriate for covering planning, execution monitoring, explanation, memory management, and metrics without redundancy.
The core lifecycle is well covered: approval, planning, path selection, progress checks, outcome recording, searching experiences, and forgetting tasks. Minor gaps exist, such as no MCP tool to revoke project approval or perform a full memory wipe, both of which are delegated to CLI commands.
Maintenance
Related MCP Connectors
Shared task board and knowledge base for AI coding agents Give your coding agents a shared task board and knowledge base, so the plan survives between sessions and across agents.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceProvides AI coding agents with persistent, long-term memory through local semantic search and SQLite storage. It enables agents to save and retrieve architectural decisions or project context across different conversation sessions without requiring cloud services.MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI coding assistants with persistent project memory to retain architectural decisions, code patterns, and domain knowledge across sessions. It stores data locally in a SQLite database, allowing agents to remember, recall, and manage project-specific context using full-text search.7 npmApache 2.0
- AlicenseNot gradedqualityCmaintenanceGives AI coding agents persistent memory by storing observations, decisions, and learnings in a local SQLite database with vector search, full-text search, and a rules engine.4MIT
- AlicenseNot gradedqualityAmaintenanceProvides long-term memory for LLMs via local SQLite storage with hybrid search (BM25, vectors, recency decay), enabling AI coding agents to persist and recall memories across sessions without cloud or API keys.53MIT