Skip to main content
Glama

Website · Docs · Walkthrough · PyPI · Issues · Buy me a coffee

PyPI Python CI License Glama score


What is OpenReflex

OpenReflex gives AI coding agents ambient project muscle memory across coding, investigation, and reasoning work. Lifecycle hooks quietly record how work actually goes, compile repeated project behaviour into reusable Reflexes, and inject only the relevant project procedure into new tasks. The model does not need to remember to call OpenReflex, and MCP diagnostics are opt-in rather than loaded into normal agent context.

You install it once and keep working normally. Core memory stays on your machine in a local SQLite database. There is no OpenReflex account or hosted service. Optional Claude token accounting uses Claude Code's local OpenTelemetry export to a loopback-only OpenReflex receiver and stores counts only.

Related MCP server: LumenCore

Why OpenReflex

  • It learns how the project behaves. OpenReflex compiles successful work into project-scoped Reflexes such as authentication changes, Helm configuration, database migrations or CI work. Resolution combines semantic intent, task family and module/file locality, so a novel task can reuse a project procedure without repeating an earlier task.

  • It catches loops while they happen. Repeated failing commands, identical retries, stalled progress, and runaway context growth raise one alert that says whether to continue, pivot to another approach, or stop and check in with you, never a stream of nags.

  • It learns after every task. Build work closes against tests/lint/build evidence; investigations can close against cross-checked sources; reasoning-only work can complete without external tools. A simple Path check only names a better option when comparable past tasks actually support one.

  • It is private by design. Only coarse, project-relative metadata is stored. File contents, commands, tool output, and transcripts never are, and capture is off until you approve a project.

Quick start

Requires Python 3.11 or newer.

Claude Code — recommended: install the plugin, reload it, inspect the current state, then explicitly enable project memory.

/plugin marketplace add vishnu-77/openreflex
/plugin install openreflex@openreflex
/reload-plugins
/openreflex
/openreflex approve

Claude may display the canonical plugin namespace as /openreflex:openreflex; both refer to the same user-invoked OpenReflex control surface when the bare alias is available. The no-argument command is read-only and compact. approve explicitly enables local memory for the current project. The plugin prepares an exact-version runtime under ~/.openreflex/runtime/ and normal agent work remains ambient after approval.

The Claude plugin does not execute a global openreflex from PATH. An older pipx/uv/pip installation can coexist without taking over the plugin runtime.

CLI / other agents: use a persistent package install.

pipx install openreflex            # or: uv tool install openreflex
cd your-project
openreflex install codex           # or cursor / opencode / claude-code

After a few tasks:

openreflex context "fix the login redirect bug"   # preview the context a task would receive
openreflex status                                 # what has been captured and learned
openreflex doctor                                 # installation checks and recent hook errors
openreflex update --check                         # check the latest stable release

Package updates preserve local memory and project configuration. Managed pipx and uv tool installs can use openreflex update, openreflex update --reinstall, or openreflex self reinstall. The Claude plugin manages its own pinned runtime automatically; these package-manager commands are only needed for standalone CLI installs.

Use it with your agent

OpenReflex installs per project with openreflex install <agent>, or as a plugin.

Agent

Connects through

Install

Verified

Claude Code

Plugin hooks, or project hooks

claude plugin marketplace add vishnu-77/openreflex then claude plugin install openreflex@openreflex, or openreflex install claude-code

Live sessions

Codex

Plugin hooks, or project hooks

codex plugin marketplace add vishnu-77/openreflex, or openreflex install codex, then trust the hooks once in /hooks

Live sessions

Cursor

Lifecycle hooks

openreflex install cursor

Protocol and fuzz tests

OpenCode

Local lifecycle plugin

openreflex install opencode

Protocol and fuzz tests

With the Claude plugin, /openreflex shows a compact read-only state view. Use /openreflex approve to enable local memory for the current project, then work normally. Other explicit controls include status, reflexes, memory, doctor, why, trace, and revoke. OpenReflex control turns are not learned as project experiences and do not trigger normal task verification. The control skill is user-invoked only and does not enable MCP diagnostics or participate in normal model routing. Claude may show the canonical namespaced form /openreflex:openreflex. Per-agent guides: Claude Code, Codex, Cursor, OpenCode.

How it works

Hooks send lifecycle events (prompt, tool start, tool end, compaction, stop) to the OpenReflex engine, which stores them in an Experience Graph:

Task -caused-> Execution -used-> Context -used-> Experience
CandidatePath -recommended_for-> Task            Lesson -recommended_for-> Task
Execution -failed_with-> ToolCall -resolved_by-> ToolCall
Execution -caused-> Outcome -caused-> Experience -caused-> Lesson
  • Before a task: similar experiences are retrieved and OpenReflex first identifies the work mode. BUILD uses inspect-first, test-first, and incremental; INVESTIGATE uses source-first, cross-check, and broad-then-deep; THINK uses reason-first, compare-options, and evidence-first. Paths are ranked using past evidence, estimated cost, risk, uncertainty and reversibility. A compact context is injected only when relevant experience exists.

  • During a task: when a detector finds a failure loop, repeated calls, stalled progress, context growth, or work past the budget, OpenReflex estimates the marginal value of more work. The current path's success estimate is updated with each call that makes no progress or fails, and compared with the cost of the work left and with the best untried alternative. The alert ends with a recommendation to continue, pivot to another strategy, or stop and ask the user. Each problem, pivot, or stop is raised once, with a cooldown between messages.

  • After a task: lifecycle hooks close the execution automatically from work-mode-appropriate evidence; the model does not need to call OpenReflex. Explicit MCP outcome/path operations remain available only for manual or diagnostic workflows. The Path check says either better option: <path> when comparable completed tasks support it or better option: none proven. Recommendations are never presented as paths that were actually executed.

Every recommendation is stored as a decision snapshot with a Reflex Score (0-100): how strong the recommendation is, which is separate from the estimated chance that the task succeeds. openreflex why explains the latest decision and openreflex trace shows the timeline. In Claude Code, a short recap appears when a task starts, when OpenReflex recommends a pivot or stop or detects trouble, and when the task completes.

Set optional limits for every task with OPENREFLEX_BUDGET, for example calls=40,minutes=20,tokens=60000. Routing, scoring, budget and recap settings come from a versioned policy: the packaged defaults, overridden by ~/.openreflex/config.toml and then by .openreflex.toml in the project.

Normal work does not load OpenReflex MCP tools. If you explicitly want model-facing introspection for a project, enable it with openreflex diagnostics enable <agent>. Disable it again with openreflex diagnostics disable <agent>. The standalone MCP Registry package remains available for users who deliberately install OpenReflex as an MCP server.

Tool

What it does

get_execution_context

Plan a task: similar past tasks, the suggested strategy with alternatives, a budget, likely files, lessons. Optional max_tool_calls, max_minutes, max_context_tokens.

check_progress

Whether more work on the current path is worth it: continue, pivot or stop.

choose_path

Declare the strategy being followed when it differs from the suggestion.

record_outcome

Record a confirmed outcome (tests/builds, cross-checked research, user confirmation, or failure) and learn from it.

explain_decision

Why the latest recommendation was made: Reflex Score, signals, confidence, next-best route.

get_execution_trace

The decision timeline of the most recent task.

get_reflex_score

The latest Reflex Score and its components as JSON.

search_experience

Past tasks in the project by description similarity, with their lessons.

explain_node

One Experience Graph node and its relations.

get_project_insights

What has been recorded, reused and learned in the project.

approve_project

Enable capture, only when the user explicitly asks.

forget_experience

Delete one past task's memory, only when the user explicitly asks.

CLI

Command

Purpose

install <agent> [--dry-run]

Install the hooks-only ambient runtime and enable the project

diagnostics enable/disable <agent>

Explicitly opt in/out of model-facing MCP diagnostics

uninstall <agent> [--dry-run]

Remove OpenReflex's integration entries; captured memory is kept

update [--check] [--reinstall]

Check or update a managed pipx/uv-tool installation; optionally force a reinstall

self reinstall

Repair/reinstall the managed OpenReflex package while preserving memory/config

self uninstall --yes

Remove only the managed OpenReflex package; memory/config remain on disk

approve / revoke

Enable or disable capture for the current project

context "<task>"

Preview the Execution Context a task would receive

status [--json]

What has been captured, reused, learned, and how model-token usage compares

why / trace

Explain the latest recommendation, or show the decision timeline

doctor

Installation, project resolution, and recent hook activity checks

forget --yes

Delete the project's data

tokens enable / tokens status / tokens disable

Opt in to local Claude Code token accounting, inspect it, or remove OpenReflex-owned telemetry settings

benchmark

Run the simulated benchmark

hook <agent> <event>

Used by ambient agent integrations

mcp

Run the optional diagnostics/admin MCP server explicitly

Privacy

Stored

Never stored

Tool name and a coarse category (read, edit, search, test, ...)

File contents

A fingerprint of the arguments

Command text

Project-relative file paths

Tool output

Pass or fail, duration, output size

Transcripts and model output

Optional model-token counts, model/source label, estimated cost

Prompt/response/tool contents from Claude telemetry

A masked one-line error signature

Paths outside the project

The prompt as a task description (up to 1,000 characters, secrets redacted)

Anything sent to an OpenReflex-hosted service: there is none

Data lives in ~/.openreflex/projects/<hash>/experience.sqlite3. Set OPENREFLEX_HOME to move it, OPENREFLEX_DISABLE=1 to turn capture off everywhere, openreflex forget --yes to delete a project's data, or ask your agent to forget_experience a single task.

Research

OpenReflex is also a research project in budget-aware execution: instead of treating success as a yes or no, it studies how agents choose execution paths, spend tool calls, model tokens and context, respond to uncertainty, and whether past work makes similar future work cheaper without reducing outcome quality. The loop is experience retrieval, evidence-aware path selection, budget-aware execution, work-mode-specific completion and plain-language path comparison. The Researcher view on the site explains the idea, lets you step through one reflex forming in the graph, and places it next to related work.

The reproducible evaluation protocol, pinned real-agent corpus, exact prompts, and claim boundaries are documented in docs/evaluation.md.

Community & Contributing

  • Issues and ideas: GitHub Issues

  • Support the project: Buy me a coffee

  • Develop locally:

    git clone https://github.com/vishnu-77/openreflex && cd openreflex
    pip install -e ".[dev]"
    pytest && ruff check src tests scripts

OpenReflex is listed in the MCP Registry as io.github.vishnu-77/openreflex and on Glama:

License

OpenReflex is released under the MIT License.

Available Tools

14 tools
approve_projectEnable projectA
Idempotent

Enable OpenReflex capture for this project. Returns: one line confirming the project is enabled, or 'Not approved' when confirm is false. Use when: the user explicitly asks to enable OpenReflex here. Never call it on your own initiative; the user can also run openreflex approve, and openreflex revoke disables capture again. Not for: anything else; every other tool answers with a 'not enabled' message until this has happened. Side effects: writes the project's approval to the local OpenReflex home; enabling an already enabled project changes nothing. Touches no project files. Errors: none.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue only when the user has explicitly asked to enable OpenReflex for this project; with false (the default) nothing changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses side effects ('writes the project's approval to the local OpenReflex home'), idempotence ('enabling an already enabled project changes nothing'), non-destructiveness ('Touches no project files'), and errors ('none'). These details go well beyond the annotations, which only state idempotentHint and destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with labeled sections (Returns, Use when, Never, Not for, Side effects, Errors) and every sentence carries useful information. The most important usage constraint is front-loaded, making it easy for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single optional parameter with an output schema, the description covers the return value, invocation conditions, exclusions, side effects, idempotence, and error behavior. Nothing an agent needs to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents `confirm` clearly with its default and meaning. The description adds the return behavior tied to `confirm=false`, but does not add substantial parameter semantics beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Enable OpenReflex capture for this project.' This clearly differentiates it from sibling tools like get_execution_context or record_outcome, so an agent can tell what this tool is for at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it ('the user explicitly asks to enable OpenReflex here'), when not to call it ('Never call it on your own initiative'), and names the CLI alternatives `openreflex approve` and `openreflex revoke`. This is exemplary routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_progressCheck progressA
Read-onlyIdempotent

Estimate whether more work on the current task's path is still worth it. Returns: plain text starting with 'Recommendation: continue', 'pivot' or 'stop', then the success estimates for the current path and the best alternative, the marginal value of each option, the execution budget and how much of it is used, and any detected problems (failure loop, repeated calls, stalled progress, context growth, budget overrun). Use when: unsure mid-task whether to keep going. Follow 'pivot' by switching to the named strategy; follow 'stop' by summarizing what was tried and asking the user. Not for: planning a task (use get_execution_context). Side effects: none; read-only on the most recent task. Errors: asks for get_execution_context first when no task exists; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description aligns with annotations by stating 'Side effects: none; read-only on the most recent task.' It adds valuable behavioral context beyond annotations: error behavior when no task exists, the 'not enabled' state until project approval, and the form of the returned recommendation. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action, then organized into labeled sections: return format, usage conditions, follow-up actions, exclusions, side effects, and errors. Every sentence adds necessary information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, zero-parameter tool with an output schema, the description covers everything an agent needs: when to call it, what the response looks like, how to react to each recommendation, and failure/error modes. This is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to document; schema coverage is vacuously 100%. The description instead explains what will happen when the tool is invoked, which is the appropriate contribution at this level.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Estimate whether more work on the current task's path is still worth it.' It also identifies the unique output ('Recommendation: continue', 'pivot' or 'stop'), which clearly distinguishes it from planning tools like get_execution_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it ('unsure mid-task whether to keep going'), what to do after 'pivot' and 'stop', and what it is not for, naming the alternative tool get_execution_context. This is complete usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

choose_pathChoose pathA
Idempotent

Declare the strategy you are following for the most recent task. Returns: one line, 'Recorded chosen path: '. Use when: deliberately departing from the suggested path, so OpenReflex remembers what was actually done rather than confusing the recommendation with the observed work. Not for: tasks that follow the suggestion; the path is then inferred from tool activity. Side effects: writes the choice to the local Experience Graph; calling it again replaces the earlier choice. Touches no project files. Errors: 'No current execution' before any task was planned (call get_execution_context first); a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNoFor a custom strategy only: 2-4 short steps, e.g. ['Prototype the toml loader', 'Swap the callers']. Ignored for suggested strategies.
strategyYesThe strategy being followed: one of the suggested names 'inspect-first', 'test-first' or 'incremental', or a short kebab-case name for your own, e.g. 'spike-then-rewrite'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that the tool writes to the local Experience Graph, replaces earlier choices on repeated calls, and touches no project files. It also lists two specific error conditions, giving the agent a clear behavioral model. The idempotentHint annotation is consistent with replacing an earlier choice with the same choice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into labeled segments — Returns, Use when, Not for, Side effects, Errors — that are easy to scan. Every sentence carries operational value, and the core purpose is front-loaded before details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers when to use the tool, what it returns, side effects, prerequisites, and failure modes. Given the rich schema and output schema, nothing essential is missing for an agent to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both strategy and steps already explained in the input schema. The description adds context about the strategy's role but no additional parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Declare the strategy you are following for the most recent task.' It clearly distinguishes the tool from siblings like record_outcome or check_progress by framing this as the act of recording a deliberately chosen path rather than observing activity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' and 'Not for' guidance, stating it is for departing from the suggested path and not for tasks that follow the suggestion. It also references the prerequisite get_execution_context in the error section, giving the agent actionable routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_decisionExplain decisionA
Read-onlyIdempotent

Explain OpenReflex's latest recommendation for the most recent task. Returns: plain text with the recommended strategy, the Reflex Score (0-100: how strong the recommendation is, not the chance of success), the estimated success probability, confidence, the signals behind the score, the next-best strategy with its route advantage, and the policy version. Use when: the user asks why OpenReflex suggested a path, pivot or stop. Not for: the same numbers as structured data (use get_reflex_score) or every decision in order (use get_execution_trace). Side effects: none; read-only. Errors: 'No decision snapshot recorded yet' before any task was planned; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavioral context beyond annotations: it explicitly states 'Side effects: none; read-only' and discloses specific error messages ('No decision snapshot recorded yet', 'not enabled'). It also clarifies that the Reflex Score is a recommendation strength, not success probability, which is a valuable interpretive detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with labeled sections ('Returns', 'Use when', 'Not for', 'Side effects', 'Errors') that front-load the most important information. Every sentence adds distinct value, and the length is justified by the rich behavioral and error context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter input and the presence of an output schema, the description fully covers what an agent needs: what is returned, when to use it, exclusions, side effects, and error conditions. No gaps are apparent for safe and correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description need not explain parameter semantics. The schema coverage is 100% and there is nothing to add; the baseline of 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Explain') and resource ('OpenReflex's latest recommendation for the most recent task'), clearly stating the tool's scope. It also distinguishes itself from siblings by explicitly naming get_reflex_score and get_execution_trace as alternatives for structured numbers or chronological decisions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' and 'Not for' sections, giving clear conditions for invocation and excluding the sibling tools. It leaves no ambiguity about when to choose this tool over the alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_nodeExplain graph nodeA
Read-onlyIdempotent

Show one Experience Graph node with its direct relations. Returns: structured 'root', 'nodes' (Task, Execution, ToolCall, Outcome, Experience, Lesson, Context or CandidatePath and their stored fields) and typed 'edges' (used, caused, failed_with, resolved_by, recommended_for). Embeddings are omitted. Use when: tracing why a lesson or recommendation exists, after finding an id with search_experience or get_execution_context. Not for: searching by topic (use search_experience). Side effects: none; read-only. Errors: an MCP error reports 'Unknown graph node' for an id that does not exist, and an MCP error is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
node_idYesId of an Experience Graph node: an experience id from search_experience (starts with 'exp-'), the execution id from get_execution_context, or any node id from an earlier explain_node result.

Output Schema

ParametersJSON Schema
NameRequiredDescription
rootYesId of the graph node that was requested.
edgesYesDirect typed relationships to or from the root node.
nodesYesThe root node and directly related nodes; embeddings are omitted.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, and the description reinforces with 'Side effects: none; read-only.' It also adds non-redundant behavioral details: embeddings are omitted, and error behavior (unknown id yields 'Unknown graph node', and errors until project approval) is disclosed. This exceeds the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured with labeled sections (Returns:, Use when:, Not for:, Side effects:, Errors:). The main purpose is front-loaded, and every section contributes either to usage or safety. Slightly verbose but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (per context signals), the description needn't enumerate return fields, but it summarizes types (root, nodes, edges) and omissions (embeddings). It also covers error conditions and usage context, making it complete for a single-parameter read-only tool. Minor gap: output schema description is not in the text, but that is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description for node_id already explains valid id sources (experience id, execution id, or previous explain_node result). The description adds little beyond reusing this guidance, so it stays at the baseline for well-covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Show one Experience Graph node') and explicitly differentiates from the sibling search_experience in 'Not for: searching by topic'. The return structure (root, nodes, edges) is also outlined, making the tool's exact scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' guidance ('tracing why a lesson or recommendation exists, after finding an id with search_experience or get_execution_context') and a clear exclusion ('Not for: searching by topic (use search_experience)'), giving agents a definitive decision rule among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

forget_experienceForget experienceA
DestructiveIdempotent

Delete one past task from this project's memory. Returns: one line listing what was removed (task, executions, tool calls, outcomes, experience, lessons), or 'Not forgotten' when confirm is false. Use when: the user explicitly asks to forget a specific past task, for example one that involved sensitive work or taught a wrong lesson. Never call it on your own initiative. Not for: deleting everything; the user runs openreflex forget --yes for that. Side effects: permanently removes that task's nodes from the local Experience Graph; other tasks are untouched. Touches no project files. Errors: 'Unknown experience id' for an id that does not exist, is not an experience, or was already forgotten; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue only when the user has explicitly asked to forget this task; with false (the default) nothing is deleted.
experience_idYesThe experience id to delete, as returned by search_experience; starts with 'exp-'.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, and the description meaningfully enriches both: it specifies the exact destruction scope ('that task's nodes from the local Experience Graph; other tasks are untouched'), asserts side-effect boundaries ('Touches no project files'), states permanence ('permanently removes'), enumerates error modes ('Unknown experience id' for non-existent/already-forgotten; 'not enabled' until project approval), and discloses the confirm=false no-op path via the 'Not forgotten' return. All of this goes well beyond the structured annotation flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured into clearly labeled sections (Returns, Use when, Not for, Side effects, Errors) with the core action front-loaded. Every section earns its place for a permanently destructive tool, and the labels make it highly scannable. Slight redundancy exists between 'Never call it on your own initiative' and confirm's schema description, keeping it just shy of maximal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a permanently destructive tool, the description covers the deletion target, return format in both confirm states, every error condition including the approval prerequisite ('not enabled' until project approved), the side-effect boundary, and the user-consent requirement. With an output schema also present to carry return-structure detail, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — experience_id is documented with its 'exp-' pattern and source (returned by search_experience), and confirm with its guardrail semantics ('true only when the user has explicitly asked'). The description reinforces confirm's default-false behavior through the 'Not forgotten' return case but does not materially extend the schema's parameter explanations. The 100%-coverage baseline of 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with 'Delete one past task from this project's memory' — a specific verb, a specific resource, and an explicit scope ('one'). The 'Not for: deleting everything' clause further carves out the boundary versus bulk deletion, and the resource ('past task', 'experience graph nodes') distinguishes it from record_outcome, search_experience, and other siblings. No tautology; the title 'Forget experience' is expanded into an operational statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit triggering condition ('Use when: the user explicitly asks to forget a specific past task') with a concrete example ('sensitive work or taught a wrong lesson'), an absolute exclusion ('Never call it on your own initiative'), and a named alternative for the excluded case (user runs `openreflex forget --yes`). This is complete routing guidance with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_candidate_pathsGet candidate pathsA
Read-onlyIdempotent

Show every candidate path OpenReflex considered for the most recent task side by side. Returns: plain text with each strategy's score, success probability, cost/risk numbers, and whether it was dominated; which path was recommended; and which one was actually followed - explicit (via choose_path), inferred from tool-call evidence, or not yet determined while the execution is still running. Use when: the user asks what other approaches were considered, or whether the agent followed the suggested path. Not for: a single recommendation's reasoning (use explain_decision) or the decision timeline (use get_execution_trace). Side effects: none; read-only. Errors: 'No execution recorded yet' before any task was planned; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

While annotations already declare readOnlyHint and destructiveHint false, the description adds substantial context: it details the exact content returned (score, success probability, cost/risk, domination status), explains how the followed path is determined (explicit, inferred, or undetermined), lists specific error messages ('No execution recorded yet', 'not enabled'), and states side effects. This goes far beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections: main purpose, returns, use-when, not-for, side effects, and errors. Each sentence adds unique value with no redundancy. It is appropriately detailed without being verbose, and the primary purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the tool takes no parameters, the description covers all necessary aspects: purpose, usage conditions, alternatives, errors, side effects, and return details. Nothing an agent needs to decide whether to invoke this tool and interpret its result is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema fully covers parameter semantics (trivially). The baseline for 0 params is 4, and the description adds no parameter-specific information because none is needed. It correctly focuses on output and behavior instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Show'), a clear resource ('every candidate path OpenReflex considered'), and a scope ('for the most recent task'). It explicitly differentiates from siblings by naming what it is not for and pointing to alternatives, leaving no ambiguity about its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit 'Use when' and 'Not for' conditions, naming the exact user intents that should trigger this tool and the sibling tools (explain_decision, get_execution_trace) that should be used instead. This is the gold standard for usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_execution_contextGet execution contextA

Plan a task from this project's past experience. Returns: plain text with the learned project Reflex procedure when one applies; otherwise a neutral working approach, relevant past experience, budget, likely files/sources and lessons. Internal strategy names and candidate scores are deliberately omitted from this normal surface; use get_candidate_paths for diagnostics. Use when: starting any substantial coding, investigation, research, analysis, review, planning or reasoning task and no [OpenReflex] block was injected. Not for: looking up history (use search_experience) or checking progress mid-task (use check_progress). Side effects: starts or re-plans the current task in the local Experience Graph; touches no project files, runs no commands, sends nothing over the network. Calling it again for the same task returns the same plan unless new limits are given. Errors: invalid non-positive limits are rejected by the input schema; a 'not enabled' message is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesThe task in one or two plain sentences, e.g. 'Fix the login redirect loop after logout'. Used to find similar past tasks; secrets are redacted before it is stored.
max_minutesNoOptional cap on active working time in minutes; a positive number. Default: no cap.
max_tool_callsNoOptional cap on tool calls for this task; a positive integer. Default: no cap.
max_context_tokensNoOptional cap on tokens of tool output added to context; a positive integer. Default: no cap.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only mark non-read-only/non-destructive, which is thin; the description carries the behavioral burden. It discloses that the call starts/re-plans a task in the local Experience Graph, touches no files/commands/network, is repeatable (same plan unless new limits are given), and can return a 'not enabled' message. This significantly exceeds annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Though longer than typical tool descriptions, it is broken into labeled sections (Returns, Use when, Not for, Side effects, Errors) that each add distinct value. The core purpose is front-loaded in the first sentence, and no sentence is redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers every decision an agent needs: what it returns, when to use it, which siblings to prefer, side effects, repeat-call behavior, and error conditions. Combined with the output schema, nothing important is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter (task, max_minutes, max_tool_calls, max_context_tokens) already has a clear schema description. The tool description adds only minor context ('unless new limits are given'), so it meets the baseline but does not meaningfully compensate beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb-object ('Plan a task from this project's past experience') and clarifies the return format. It explicitly distinguishes itself from sibling tools like get_candidate_paths, search_experience, and check_progress, so an agent can select it confidently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Contains an explicit 'Use when' condition (starting substantial work with no [OpenReflex] block injected) and a 'Not for' section naming alternatives (search_experience for history, check_progress for mid-task checks). This is exactly the when/when-not/alternatives guidance requested.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_execution_traceGet execution traceA
Read-onlyIdempotent

Return the decision timeline of the most recent task. Returns: plain text, one line per decision in order, with elapsed time, phase (start, runtime, complete), action, strategy, Reflex Score and the event that triggered it. Never includes prompts, commands or tool output. Use when: reviewing how a task unfolded. Not for: only the latest decision (use explain_decision). Side effects: none; read-only. Errors: 'No decision snapshots recorded yet' before any task was planned; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and idempotent behavior, and the description adds valuable behavioral context: plain text output format, one line per decision, content exclusions, and specific error messages. It also confirms no side effects, complementing the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet well-structured with labeled sections: Returns, Use when, Not for, Side effects, and Errors. Every sentence adds useful information, and the main purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a read-only, parameterless operation, and the description covers purpose, output format, usage boundaries, side effects, and error conditions. An output schema exists, so return-value details need not be repeated. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is trivially complete and the description does not need to explain parameter meaning. The baseline of 4 applies; the description adds useful output details even though parameters are absent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return the decision timeline of the most recent task.' It also distinguishes itself from explain_decision by explicitly noting it is not for only the latest decision. This makes it easy for an agent to select among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance: 'Use when: reviewing how a task unfolded' and 'Not for: only the latest decision (use explain_decision).' This clearly tells the agent when to call this tool and when to choose an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_insightsGet project insightsA
Read-onlyIdempotent

Summarize what OpenReflex has recorded and learned in this project. Returns: structured activation, engagement, experience reuse, outcomes, observational efficiency with and without prior experience, path-comparison trends, routing agreement, live alerts, execution-control metrics and lesson count. Efficiency comparisons are observational and are not presented as causal evidence. Use when: the user asks how OpenReflex is doing in this project or a program needs project-level metrics. Not for: individual past tasks (use search_experience). Side effects: none; read-only. Errors: an MCP error is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
lessonsYesNumber of extracted lessons retained in this project.
projectYesLocal project path represented by this Experience Graph.
routingYesRecommendation agreement with retrospective realised performance.
outcomesYesKnown and verified outcome coverage.
activationYesActivation and first-use metrics.
engagementYesCapture and usage volume.
path_checkYesSimple evidence-backed path comparison counts.
live_alertsYesCounts of loop, repetition, stagnation, context and budget alerts.
experience_reuseYesHow often prior experience was reused.
project_reflexesYesProject-specific reusable procedures learned from execution evidence.
execution_controlYesRuntime verdict and budget-adherence metrics.
efficiency_observationalYesObserved efficiency with versus without reused experience; this is not a controlled comparison.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive behavior, and the description reinforces this with 'Side effects: none; read-only.' It adds genuinely useful context beyond annotations: the error condition until project approval and the caveat that efficiency comparisons are observational, not causal evidence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with labeled sections, front-loads the core purpose, and every section earns its place: purpose, return content, usage conditions, exclusions, side effects, and errors. The length is justified given the breadth of insights returned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This tool has no parameters and a rich output schema, and the description covers everything needed for correct invocation: when to use it, when not to, the read-only guarantee, the project-approval error condition, and the causal-evidence caveat. No critical context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is no schema gap for the description to compensate for. The baseline for a zero-parameter tool is 4, and the description appropriately focuses on what the tool returns and when to use it rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Summarize what OpenReflex has recorded and learned in this project.' It also distinguishes itself from search_experience by clarifying it covers project-level insights, not individual past tasks. The detailed Returns list further disambiguates the tool's scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Not for' sections give clear invocation conditions and explicitly name search_experience as the alternative for individual past tasks. This gives an agent direct routing guidance with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_reflexesGet project ReflexesA
Read-onlyIdempotent

Show the project Reflex and specialised procedures OpenReflex is learning from repeated work. Returns: the root Project Reflex after the first captured experience, plus specialised Reflexes with learning/learned/proven/stale maturity, support/verification counts and procedure steps. Use when: the user asks what OpenReflex has learned specifically about this project, which reusable procedures exist, or how a named Reflex works. Not for: backend routing internals or generic strategies (use get_candidate_paths for those diagnostics). Side effects: may refresh derived local Reflex data from already-captured evidence; touches no project files. Errors: a 'not enabled' message is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, but the description adds valuable context: a potential refresh side effect (non-destructive) and an error condition ('not enabled' until approval). This goes beyond the structured annotations and clarifies behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Returns, Use when, Not for, Side effects, Errors) and every sentence adds value. It is front-loaded with the primary purpose and remains concise despite its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter tool and the presence of an output schema, the description provides sufficient context: it describes what is returned, when to use, exclusions, side effects, and error conditions. Nothing critical is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline for parameter semantics is 4. The description correctly avoids any parameter explanations since none exist, and the schema coverage is 100% (trivially).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool shows project Reflexes and specialised procedures, with a specific verb and resource. It explicitly differentiates from sibling get_candidate_paths by stating what it is not for, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit 'Use when' scenarios and a 'Not for' section that names the alternative tool (get_candidate_paths). This fully guides an agent on when to select this tool versus others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reflex_scoreGet Reflex ScoreA
Read-onlyIdempotent

Return the latest decision's Reflex Score and machine-readable components. Returns: structured fields for availability, Reflex Score (0-100 recommendation strength, not success probability), success probability, confidence, strategy, next-best strategy, route advantage, evidence, context cost, budget use, component signals and policy version. If no decision exists, available=false and reason explains why. Use when: a program or agent needs decision numbers it can inspect without parsing prose. Not for: a readable explanation (use explain_decision) or the full decision timeline (use get_execution_trace). Side effects: none; read-only. Errors: an MCP error is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
reasonNoWhy no score is available, when available is false.
signalsNoNormalised component signals used to compute the Reflex Score.
strategyNoRecommended strategy for the latest decision.
availableYesWhether a decision snapshot is available for the most recent decided task.
budget_usedNoLargest fraction of the task budget consumed across time, calls and context.
reflex_scoreNoStrength of the recommendation on a 0-100 scale; not success probability.
context_tokensNoEstimated tokens injected from prior experience for this task.
evidence_countNoNumber of relevant prior experiences supporting the decision.
policy_versionNoOpenReflex decision-policy version used for the score.
route_advantageNoNormalised advantage of the recommended route over the next best route.
next_best_strategyNoHighest-ranked alternative strategy, if any.
decision_confidenceNoConfidence in the recommendation from the available evidence.
success_probabilityNoEstimated probability of success for the recommended strategy.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only, idempotent, and non-destructive, and the description adds value by stating side effects are none, errors occur until project approval, and availability behavior when no decision exists. It also clarifies that the Reflex Score is recommendation strength, not success probability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into purpose, return contents, usage guidance, exclusions, side effects, and errors. Each section earns its place, and the core purpose is front-loaded before the detailed field list.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description still covers availability, error behavior, side effects, and sibling-tool distinctions. Nothing needed to invoke or interpret the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter burden for the description to carry. The description instead clarifies the output contract, which is the relevant semantic content for this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb 'Return', the resource 'latest decision's Reflex Score', and the machine-readable component fields. It explicitly distinguishes itself from explain_decision and get_execution_trace, so an agent can identify its unique purpose without inspecting siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' guidance for programmatic inspection of decision numbers, and 'Not for' guidance with named alternatives for prose explanations and full timelines. This leaves no ambiguity about when to select this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_outcomeRecord outcomeA
Idempotent

Record the verified outcome of the most recent task and learn from it. Returns: the recorded status plus a plain-language Path check. A better path is only named when comparable completed tasks provide evidence; otherwise the result says that no better option is proven. Use when: the result is confirmed: tests/lint/build passed, research was cross-checked against relevant evidence, the user confirmed the answer, or the task failed/was abandoned. Not for: declaring the strategy (use choose_path). Side effects: finalizes the task in the local Experience Graph and updates its experience, lessons and path comparison; calling it again for the same task replaces the recorded outcome. Touches no project files. Errors: 'No execution to record an outcome for' before any task was planned; a 'not enabled' message until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes'success' when the result is confirmed (tests/builds passed, research was cross-checked, or the user confirmed it); 'failure' when the task failed or was abandoned.
evidenceYesShort proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Secrets are redacted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already note idempotentHint=true and destructiveHint=false, and the description adds meaningful context without contradicting them: it finalizes the task in the local Experience Graph, updates experience/lessons/path comparison, replaces prior recorded outcomes on repeat calls, and touches no project files. It also discloses error conditions like 'No execution to record an outcome for' and the 'not enabled' project approval state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then organized into clear labeled sections: Returns, Use when, Not for, Side effects, and Errors. Despite being detailed, every sentence earns its place and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return details are partially covered, and the description adds the necessary decision context, side effects, error conditions, and prerequisites. It tells an agent when to call it, what will happen, what will not happen, and what errors to expect, making it fully actionable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already fully explains status and evidence, including the status enum semantics. The description reinforces the status meaning in the 'Use when' section but adds no new parameter-level syntax or format details beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: 'Record the verified outcome of the most recent task and learn from it.' The 'Not for' line explicitly contrasts it with choose_path, making the tool's scope unambiguous and differentiating it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit 'Use when' conditions with concrete examples such as tests/lint/build passing, research cross-checked, user confirmation, or task failure/abandonment. It also gives an explicit exclusion, 'Not for: declaring the strategy (use choose_path),' which is clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_experienceSearch experienceA
Read-onlyIdempotent

Search this project's past tasks by description similarity. Returns: structured 'experiences' (id, score, description, class, agent, strategy, status, tool_calls, minutes, files), best match first, plus structured 'lessons' (text, confidence, support). Both lists are empty when nothing is similar enough. Use when: the user asks what was learned or what worked before, or to find an experience id for explain_node or forget_experience. Not for: planning a new task (use get_execution_context). Side effects: none; read-only. Matching is lexical, on task descriptions only. Errors: invalid limit values are rejected by the input schema; an MCP error is returned until the project is approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of past tasks to return. Default 5; allowed 1-20.
queryYesWords describing the topic, e.g. 'expired token login bug'. Matched lexically against past task descriptions.

Output Schema

ParametersJSON Schema
NameRequiredDescription
lessonsYesLessons derived from the returned experiences.
experiencesYesMatching past tasks ordered from best to worst match.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds meaningful behavior beyond the annotations: describes return ordering, empty-list behavior, lexical matching scope, side-effect-free read-only nature, and error conditions including the approval requirement. This is rich contextual information that helps the agent predict tool behavior accurately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with labeled sections: Returns, Use when, Not for, Side effects, Errors. Every sentence provides distinct value, and the most important usage guidance is front-loaded near the top.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Combined with a fully documented input schema, rich annotations, and an output schema, the description covers purpose, usage scenarios, exclusions, side effects, return behavior, and error cases. Nothing critical for an agent to select and invoke the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema descriptions cover 100% of parameters, so the description is not required to re-document them. It adds minor context about lexical matching and validation errors, but the schema already explains query and limit adequately, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Search this project's past tasks by description similarity.' It clearly identifies what the tool does and distinguishes itself from siblings by naming get_execution_context as the alternative for planning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' conditions: asking what was learned or worked before, or finding an experience id for explain_node or forget_experience. It also states what it is not for and names the alternative tool, leaving no ambiguity about when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.9.6
    • Changedget_project_insights3 fields changed
      • addedOutput schema / $defs / ProjectReflexMetrics
        Added value: +{
        +  "properties": {
        +    "credit_signals": {
        +      "description": "Action-level execution-credit signals stored across surfaced Reflexes.",
        +      "minimum": 0,
        +      "title": "Credit Signals",
        +      "type": "integer"
        +    },
        +    "credit_spine_actions": {
        +      "description": "Action signals with enough repeated evidence to participate in compressed Reflex procedures.",
        +      "minimum": 0,
        +      "title": "Credit Spine Actions",
        +      "type": "integer"
        +    },
        +    "learned": {
        +      "description": "Specialised Reflexes promoted after repeated successful executions.",
        +      "minimum": 0,
        +      "title": "Learned",
        +      "type": "integer"
        +    },
        +    "learning": {
        +      "description": "Specialised Reflexes still accumulating evidence.",
        +      "minimum": 0,
        +      "title": "Learning",
        +      "type": "integer"
        +    },
        +    "project_state": {
        +      "description": "Current maturity of the root Project Reflex.",
        +      "enum": [
        +        "cold",
        +        "learning",
        +        "learned",
        +        "proven",
        +        "stale"
        +      ],
        +      "title": "Project State",
        +      "type": "string"
        +    },
        +    "project_support": {
        +      "description": "Captured experiences supporting the root Project Reflex.",
        +      "minimum": 0,
        +      "title": "Project Support",
        +      "type": "integer"
        +    },
        +    "proven": {
        +      "description": "Specialised Reflexes backed by repeated explicitly verified evidence.",
        +      "minimum": 0,
        +      "title": "Proven",
        +      "type": "integer"
        +    },
        +    "specialised_total": {
        +      "description": "Number of specialised area/procedure Reflexes.",
        +      "minimum": 0,
        +      "title": "Specialised Total",
        +      "type": "integer"
        +    },
        +    "stale": {
        +      "description": "Previously learned Reflexes contradicted by newer evidence.",
        +      "minimum": 0,
        +      "title": "Stale",
        +      "type": "integer"
        +    },
        +    "visible": {
        +      "description": "Project and specialised Reflexes currently surfaced.",
        +      "minimum": 0,
        +      "title": "Visible",
        +      "type": "integer"
        +    }
        +  },
        +  "required": [
        +    "visible",
        +    "project_state",
        +    "project_support",
        +    "specialised_total",
        +    "learning",
        +    "learned",
        +    "proven",
        +    "stale",
        +    "credit_signals",
        +    "credit_spine_actions"
        +  ],
        +  "title": "ProjectReflexMetrics",
        +  "type": "object"
        +}
      • addedOutput schema / properties / project_reflexes
        Added value: +{
        +  "$ref": "#/$defs/ProjectReflexMetrics",
        +  "description": "Project-specific reusable procedures learned from execution evidence."
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "project",
        -  "activation",
        -  "engagement",
        -  "experience_reuse",
        -  "outcomes",
        -  "efficiency_observational",
        -  "path_check",
        -  "routing",
        -  "live_alerts",
        -  "execution_control",
        -  "lessons"
        -]New value: +[
        +  "project",
        +  "activation",
        +  "engagement",
        +  "experience_reuse",
        +  "outcomes",
        +  "efficiency_observational",
        +  "path_check",
        +  "routing",
        +  "live_alerts",
        +  "execution_control",
        +  "project_reflexes",
        +  "lessons"
        +]
    • Addedget_project_reflexes
  2. 1 tool updatev0.5.3
    • Addedget_candidate_paths
  3. 8 tool updatesv0.5.1
    • Changedchoose_path2 fields changed
      • addedInput schema / properties / strategy / maxLength
        Added value: +60
      • addedInput schema / properties / strategy / minLength
        Added value: +1
    • Changedexplain_node9 fields changed
      • addedInput schema / properties / node_id / maxLength
        Added value: +160
      • addedInput schema / properties / node_id / minLength
        Added value: +1
      • addedOutput schema / $defs
        Added value: +{
        +  "GraphEdge": {
        +    "properties": {
        +      "relation": {
        +        "description": "Typed relationship between the source and target nodes.",
        +        "enum": [
        +          "used",
        +          "caused",
        +          "failed_with",
        +          "resolved_by",
        +          "recommended_for"
        +        ],
        +        "title": "Relation",
        +        "type": "string"
        +      },
        +      "source": {
        +        "description": "Source node id.",
        +        "title": "Source",
        +        "type": "string"
        +      },
        +      "target": {
        +        "description": "Target node id.",
        +        "title": "Target",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "source",
        +      "relation",
        +      "target"
        +    ],
        +    "title": "GraphEdge",
        +    "type": "object"
        +  },
        +  "GraphNode": {
        +    "additionalProperties": true,
        +    "properties": {
        +      "id": {
        +        "description": "Stable node id within this project's local Experience Graph.",
        +        "title": "Id",
        +        "type": "string"
        +      },
        +      "kind": {
        +        "description": "Experience Graph node type.",
        +        "enum": [
        +          "Task",
        +          "Context",
        +          "CandidatePath",
        +          "Execution",
        +          "ToolCall",
        +          "Outcome",
        +          "Experience",
        +          "Lesson"
        +        ],
        +        "title": "Kind",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "kind",
        +      "id"
        +    ],
        +    "title": "GraphNode",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / properties / edges
        Added value: +{
        +  "description": "Direct typed relationships to or from the root node.",
        +  "items": {
        +    "$ref": "#/$defs/GraphEdge"
        +  },
        +  "title": "Edges",
        +  "type": "array"
        +}
      • addedOutput schema / properties / nodes
        Added value: +{
        +  "description": "The root node and directly related nodes; embeddings are omitted.",
        +  "items": {
        +    "$ref": "#/$defs/GraphNode"
        +  },
        +  "title": "Nodes",
        +  "type": "array"
        +}
      • removedOutput schema / properties / result
        Removed value: -{
        -  "title": "Result",
        -  "type": "string"
        -}
      • addedOutput schema / properties / root
        Added value: +{
        +  "description": "Id of the graph node that was requested.",
        +  "title": "Root",
        +  "type": "string"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "root",
        +  "nodes",
        +  "edges"
        +]
      • changedOutput schema / title
        Previous value: -"explain_nodeOutput"New value: +"ExplainNodeResult"
    • Changedforget_experience3 fields changed
      • addedInput schema / properties / experience_id / maxLength
        Added value: +160
      • addedInput schema / properties / experience_id / minLength
        Added value: +5
      • addedInput schema / properties / experience_id / pattern
        Added value: +"^exp-"
    • Changedget_execution_context5 fields changed
      • changedInput schema / properties / max_context_tokens / anyOf
        Previous value: -[
        -  {
        -    "type": "integer"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "exclusiveMinimum": 0,
        +    "type": "integer"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / max_minutes / anyOf
        Previous value: -[
        -  {
        -    "type": "number"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "exclusiveMinimum": 0,
        +    "type": "number"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / max_tool_calls / anyOf
        Previous value: -[
        -  {
        -    "type": "integer"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "exclusiveMinimum": 0,
        +    "type": "integer"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / task / maxLength
        Added value: +1000
      • addedInput schema / properties / task / minLength
        Added value: +1
    • Changedget_project_insights15 fields changed
      • addedOutput schema / $defs
        Added value: +{
        +  "ActivationMetrics": {
        +    "properties": {
        +      "approved_at": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Unix timestamp when capture was approved for this project.",
        +        "title": "Approved At"
        +      },
        +      "first_session_captured": {
        +        "anyOf": [
        +          {
        +            "type": "boolean"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Whether the first observed session produced an experience.",
        +        "title": "First Session Captured"
        +      },
        +      "seconds_to_first_task": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Seconds from approval to the first captured task.",
        +        "title": "Seconds To First Task"
        +      }
        +    },
        +    "required": [
        +      "approved_at",
        +      "seconds_to_first_task",
        +      "first_session_captured"
        +    ],
        +    "title": "ActivationMetrics",
        +    "type": "object"
        +  },
        +  "EfficiencyObservational": {
        +    "properties": {
        +      "model_token_change": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Relative mean model-token change when real telemetry exists; observational, not causal.",
        +        "title": "Model Token Change"
        +      },
        +      "time_change": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Relative mean active-time change; observational, not causal.",
        +        "title": "Time Change"
        +      },
        +      "tool_call_change": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Relative mean tool-call change; observational, not causal.",
        +        "title": "Tool Call Change"
        +      },
        +      "with_prior_experience": {
        +        "$ref": "#/$defs/EfficiencyProfile",
        +        "description": "Observed metrics for tasks that reused prior experience."
        +      },
        +      "without_prior_experience": {
        +        "$ref": "#/$defs/EfficiencyProfile",
        +        "description": "Observed metrics for tasks without prior experience."
        +      }
        +    },
        +    "required": [
        +      "with_prior_experience",
        +      "without_prior_experience",
        +      "tool_call_change",
        +      "model_token_change",
        +      "time_change"
        +    ],
        +    "title": "EfficiencyObservational",
        +    "type": "object"
        +  },
        +  "EfficiencyProfile": {
        +    "properties": {
        +      "known_outcomes": {
        +        "description": "Number of known outcomes in the group.",
        +        "minimum": 0,
        +        "title": "Known Outcomes",
        +        "type": "integer"
        +      },
        +      "minutes": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Mean active execution time in minutes.",
        +        "title": "Minutes"
        +      },
        +      "model_tokens": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Mean real model-token usage when Claude telemetry is available.",
        +        "title": "Model Tokens"
        +      },
        +      "n": {
        +        "description": "Number of experiences in this observational group.",
        +        "minimum": 0,
        +        "title": "N",
        +        "type": "integer"
        +      },
        +      "success_rate": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Success rate among known outcomes in the group.",
        +        "title": "Success Rate"
        +      },
        +      "tool_calls": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Mean tool calls in the group.",
        +        "title": "Tool Calls"
        +      },
        +      "tool_output_tokens_estimate": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Mean legacy estimate of tool-output tokens.",
        +        "title": "Tool Output Tokens Estimate"
        +      }
        +    },
        +    "required": [
        +      "n",
        +      "tool_calls",
        +      "model_tokens",
        +      "tool_output_tokens_estimate",
        +      "minutes",
        +      "success_rate",
        +      "known_outcomes"
        +    ],
        +    "title": "EfficiencyProfile",
        +    "type": "object"
        +  },
        +  "EngagementMetrics": {
        +    "properties": {
        +      "active_week_ratio": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Share of elapsed weeks containing activity.",
        +        "title": "Active Week Ratio"
        +      },
        +      "active_weeks": {
        +        "description": "Distinct weeks containing captured executions.",
        +        "minimum": 0,
        +        "title": "Active Weeks",
        +        "type": "integer"
        +      },
        +      "agents": {
        +        "description": "Agents observed in this project.",
        +        "items": {
        +          "type": "string"
        +        },
        +        "title": "Agents",
        +        "type": "array"
        +      },
        +      "executions": {
        +        "description": "Total execution records.",
        +        "minimum": 0,
        +        "title": "Executions",
        +        "type": "integer"
        +      },
        +      "experiences": {
        +        "description": "Completed executions retained as reusable experiences.",
        +        "minimum": 0,
        +        "title": "Experiences",
        +        "type": "integer"
        +      },
        +      "substantial_tasks": {
        +        "description": "Tasks eligible for planning and experience retrieval.",
        +        "minimum": 0,
        +        "title": "Substantial Tasks",
        +        "type": "integer"
        +      },
        +      "tasks": {
        +        "description": "Total captured tasks, including non-substantial tasks.",
        +        "minimum": 0,
        +        "title": "Tasks",
        +        "type": "integer"
        +      },
        +      "weeks_since_start": {
        +        "description": "Weeks elapsed since activation or first task.",
        +        "minimum": 1,
        +        "title": "Weeks Since Start",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "tasks",
        +      "substantial_tasks",
        +      "executions",
        +      "experiences",
        +      "active_weeks",
        +      "weeks_since_start",
        +      "active_week_ratio",
        +      "agents"
        +    ],
        +    "title": "EngagementMetrics",
        +    "type": "object"
        +  },
        +  "ExecutionControlMetrics": {
        +    "properties": {
        +      "success_after_pivot_or_stop": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Observed success rate after pivot/stop advice; not causal.",
        +        "title": "Success After Pivot Or Stop"
        +      },
        +      "tasks_within_tool_call_budget": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Share of measured tasks finishing within the tool-call budget.",
        +        "title": "Tasks Within Tool Call Budget"
        +      },
        +      "verdicts": {
        +        "additionalProperties": {
        +          "type": "integer"
        +        },
        +        "description": "Counts of runtime continue, pivot and stop verdicts.",
        +        "title": "Verdicts",
        +        "type": "object"
        +      }
        +    },
        +    "required": [
        +      "verdicts",
        +      "success_after_pivot_or_stop",
        +      "tasks_within_tool_call_budget"
        +    ],
        +    "title": "ExecutionControlMetrics",
        +    "type": "object"
        +  },
        +  "ExperienceReuseMetrics": {
        +    "properties": {
        +      "benefit_rate": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Deprecated compatibility alias of reuse_rate; it does not measure causal benefit.",
        +        "title": "Benefit Rate"
        +      },
        +      "reuse_rate": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Share of substantial tasks that received relevant prior experience; this measures coverage, not benefit.",
        +        "title": "Reuse Rate"
        +      },
        +      "tasks_with_prior_experience": {
        +        "description": "Number of substantial tasks that received prior experience.",
        +        "minimum": 0,
        +        "title": "Tasks With Prior Experience",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "reuse_rate",
        +      "benefit_rate",
        +      "tasks_with_prior_experience"
        +    ],
        +    "title": "ExperienceReuseMetrics",
        +    "type": "object"
        +  },
        +  "OutcomeMetrics": {
        +    "properties": {
        +      "known": {
        +        "description": "Experiences whose outcome is success or failure rather than unknown.",
        +        "minimum": 0,
        +        "title": "Known",
        +        "type": "integer"
        +      },
        +      "success_rate": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Success rate among experiences with known outcomes.",
        +        "title": "Success Rate"
        +      },
        +      "verified": {
        +        "description": "Outcomes explicitly verified by the agent or user.",
        +        "minimum": 0,
        +        "title": "Verified",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "known",
        +      "verified",
        +      "success_rate"
        +    ],
        +    "title": "OutcomeMetrics",
        +    "type": "object"
        +  },
        +  "PathCheckMetrics": {
        +    "properties": {
        +      "better_option_found": {
        +        "description": "Tasks where comparable past evidence indicated a better option.",
        +        "minimum": 0,
        +        "title": "Better Option Found",
        +        "type": "integer"
        +      },
        +      "by_class": {
        +        "additionalProperties": {
        +          "type": "integer"
        +        },
        +        "description": "Number of evidence-backed path checks by task class.",
        +        "title": "By Class",
        +        "type": "object"
        +      },
        +      "comparisons": {
        +        "description": "Completed tasks with enough comparable past evidence to check another path.",
        +        "minimum": 0,
        +        "title": "Comparisons",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "comparisons",
        +      "better_option_found",
        +      "by_class"
        +    ],
        +    "title": "PathCheckMetrics",
        +    "type": "object"
        +  },
        +  "RoutingMetrics": {
        +    "properties": {
        +      "agreement": {
        +        "anyOf": [
        +          {
        +            "type": "number"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Share of comparable executions where the recommendation matched retrospective best.",
        +        "title": "Agreement"
        +      },
        +      "compared_executions": {
        +        "description": "Number of executions eligible for routing-agreement comparison.",
        +        "minimum": 0,
        +        "title": "Compared Executions",
        +        "type": "integer"
        +      },
        +      "retrospective_best": {
        +        "additionalProperties": {
        +          "type": "string"
        +        },
        +        "description": "Best realised strategy per task class when enough evidence exists.",
        +        "title": "Retrospective Best",
        +        "type": "object"
        +      }
        +    },
        +    "required": [
        +      "retrospective_best",
        +      "agreement",
        +      "compared_executions"
        +    ],
        +    "title": "RoutingMetrics",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / properties / activation
        Added value: +{
        +  "$ref": "#/$defs/ActivationMetrics",
        +  "description": "Activation and first-use metrics."
        +}
      • addedOutput schema / properties / efficiency_observational
        Added value: +{
        +  "$ref": "#/$defs/EfficiencyObservational",
        +  "description": "Observed efficiency with versus without reused experience; this is not a controlled comparison."
        +}
      • addedOutput schema / properties / engagement
        Added value: +{
        +  "$ref": "#/$defs/EngagementMetrics",
        +  "description": "Capture and usage volume."
        +}
      • addedOutput schema / properties / execution_control
        Added value: +{
        +  "$ref": "#/$defs/ExecutionControlMetrics",
        +  "description": "Runtime verdict and budget-adherence metrics."
        +}
      • addedOutput schema / properties / experience_reuse
        Added value: +{
        +  "$ref": "#/$defs/ExperienceReuseMetrics",
        +  "description": "How often prior experience was reused."
        +}
      • addedOutput schema / properties / lessons
        Added value: +{
        +  "description": "Number of extracted lessons retained in this project.",
        +  "minimum": 0,
        +  "title": "Lessons",
        +  "type": "integer"
        +}
      • addedOutput schema / properties / live_alerts
        Added value: +{
        +  "additionalProperties": {
        +    "type": "integer"
        +  },
        +  "description": "Counts of loop, repetition, stagnation, context and budget alerts.",
        +  "title": "Live Alerts",
        +  "type": "object"
        +}
      • addedOutput schema / properties / outcomes
        Added value: +{
        +  "$ref": "#/$defs/OutcomeMetrics",
        +  "description": "Known and verified outcome coverage."
        +}
      • addedOutput schema / properties / path_check
        Added value: +{
        +  "$ref": "#/$defs/PathCheckMetrics",
        +  "description": "Simple evidence-backed path comparison counts."
        +}
      • addedOutput schema / properties / project
        Added value: +{
        +  "description": "Local project path represented by this Experience Graph.",
        +  "title": "Project",
        +  "type": "string"
        +}
      • removedOutput schema / properties / result
        Removed value: -{
        -  "title": "Result",
        -  "type": "string"
        -}
      • addedOutput schema / properties / routing
        Added value: +{
        +  "$ref": "#/$defs/RoutingMetrics",
        +  "description": "Recommendation agreement with retrospective realised performance."
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "project",
        +  "activation",
        +  "engagement",
        +  "experience_reuse",
        +  "outcomes",
        +  "efficiency_observational",
        +  "path_check",
        +  "routing",
        +  "live_alerts",
        +  "execution_control",
        +  "lessons"
        +]
      • changedOutput schema / title
        Previous value: -"get_project_insightsOutput"New value: +"ProjectInsightsResult"
    • Changedget_reflex_score16 fields changed
      • addedOutput schema / properties / available
        Added value: +{
        +  "description": "Whether a decision snapshot is available for the most recent decided task.",
        +  "title": "Available",
        +  "type": "boolean"
        +}
      • addedOutput schema / properties / budget_used
        Added value: +{
        +  "anyOf": [
        +    {
        +      "minimum": 0,
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Largest fraction of the task budget consumed across time, calls and context.",
        +  "title": "Budget Used"
        +}
      • addedOutput schema / properties / context_tokens
        Added value: +{
        +  "anyOf": [
        +    {
        +      "minimum": 0,
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Estimated tokens injected from prior experience for this task.",
        +  "title": "Context Tokens"
        +}
      • addedOutput schema / properties / decision_confidence
        Added value: +{
        +  "anyOf": [
        +    {
        +      "maximum": 1,
        +      "minimum": 0,
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Confidence in the recommendation from the available evidence.",
        +  "title": "Decision Confidence"
        +}
      • addedOutput schema / properties / evidence_count
        Added value: +{
        +  "anyOf": [
        +    {
        +      "minimum": 0,
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Number of relevant prior experiences supporting the decision.",
        +  "title": "Evidence Count"
        +}
      • addedOutput schema / properties / next_best_strategy
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Highest-ranked alternative strategy, if any.",
        +  "title": "Next Best Strategy"
        +}
      • addedOutput schema / properties / policy_version
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "OpenReflex decision-policy version used for the score.",
        +  "title": "Policy Version"
        +}
      • addedOutput schema / properties / reason
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Why no score is available, when available is false.",
        +  "title": "Reason"
        +}
      • addedOutput schema / properties / reflex_score
        Added value: +{
        +  "anyOf": [
        +    {
        +      "maximum": 100,
        +      "minimum": 0,
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Strength of the recommendation on a 0-100 scale; not success probability.",
        +  "title": "Reflex Score"
        +}
      • removedOutput schema / properties / result
        Removed value: -{
        -  "title": "Result",
        -  "type": "string"
        -}
      • addedOutput schema / properties / route_advantage
        Added value: +{
        +  "anyOf": [
        +    {
        +      "maximum": 1,
        +      "minimum": 0,
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Normalised advantage of the recommended route over the next best route.",
        +  "title": "Route Advantage"
        +}
      • addedOutput schema / properties / signals
        Added value: +{
        +  "additionalProperties": {
        +    "type": "number"
        +  },
        +  "description": "Normalised component signals used to compute the Reflex Score.",
        +  "title": "Signals",
        +  "type": "object"
        +}
      • addedOutput schema / properties / strategy
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Recommended strategy for the latest decision.",
        +  "title": "Strategy"
        +}
      • addedOutput schema / properties / success_probability
        Added value: +{
        +  "anyOf": [
        +    {
        +      "maximum": 1,
        +      "minimum": 0,
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Estimated probability of success for the recommended strategy.",
        +  "title": "Success Probability"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "available"
        +]
      • changedOutput schema / title
        Previous value: -"get_reflex_scoreOutput"New value: +"ReflexScoreResult"
    • Changedrecord_outcome4 fields changed
      • changedInput schema / properties / evidence / description
        Previous value: -"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Up to 500 characters; secrets are redacted."New value: +"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Secrets are redacted."
      • addedInput schema / properties / evidence / maxLength
        Added value: +500
      • addedInput schema / properties / evidence / minLength
        Added value: +1
      • changedInput schema / properties / status / description
        Previous value: -"'success' when tests, lint or build passed or the user confirmed the result; 'failure' when the task failed or was abandoned."New value: +"'success' when the result is confirmed (tests/builds passed, research was cross-checked, or the user confirmed it); 'failure' when the task failed or was abandoned."
    • Changedsearch_experience11 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of past tasks to return, 1-20; values outside the range are clamped. Default 5."New value: +"Maximum number of past tasks to return. Default 5; allowed 1-20."
      • addedInput schema / properties / limit / maximum
        Added value: +20
      • addedInput schema / properties / limit / minimum
        Added value: +1
      • addedInput schema / properties / query / maxLength
        Added value: +1000
      • addedInput schema / properties / query / minLength
        Added value: +1
      • addedOutput schema / $defs
        Added value: +{
        +  "ExperienceSummary": {
        +    "properties": {
        +      "agent": {
        +        "description": "Agent that produced the experience.",
        +        "title": "Agent",
        +        "type": "string"
        +      },
        +      "class": {
        +        "description": "OpenReflex task class.",
        +        "title": "Class",
        +        "type": "string"
        +      },
        +      "description": {
        +        "description": "Redacted task description stored for the past execution.",
        +        "title": "Description",
        +        "type": "string"
        +      },
        +      "files": {
        +        "description": "Project-relative files associated with the execution, capped for compactness.",
        +        "items": {
        +          "type": "string"
        +        },
        +        "title": "Files",
        +        "type": "array"
        +      },
        +      "id": {
        +        "description": "Experience Graph id for the past task.",
        +        "title": "Id",
        +        "type": "string"
        +      },
        +      "minutes": {
        +        "description": "Active execution time in minutes.",
        +        "minimum": 0,
        +        "title": "Minutes",
        +        "type": "number"
        +      },
        +      "score": {
        +        "description": "Similarity/retrieval score for this query; higher ranks first.",
        +        "title": "Score",
        +        "type": "number"
        +      },
        +      "status": {
        +        "description": "Observed outcome status: success, failure or unknown.",
        +        "title": "Status",
        +        "type": "string"
        +      },
        +      "strategy": {
        +        "anyOf": [
        +          {
        +            "type": "string"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "description": "Strategy used for the past task when known.",
        +        "title": "Strategy"
        +      },
        +      "tool_calls": {
        +        "description": "Number of captured tool calls in the execution.",
        +        "minimum": 0,
        +        "title": "Tool Calls",
        +        "type": "integer"
        +      }
        +    },
        +    "required": [
        +      "id",
        +      "score",
        +      "description",
        +      "class",
        +      "agent",
        +      "strategy",
        +      "status",
        +      "tool_calls",
        +      "minutes",
        +      "files"
        +    ],
        +    "title": "ExperienceSummary",
        +    "type": "object"
        +  },
        +  "LessonSummary": {
        +    "properties": {
        +      "confidence": {
        +        "description": "Combined confidence in this lesson.",
        +        "maximum": 1,
        +        "minimum": 0,
        +        "title": "Confidence",
        +        "type": "number"
        +      },
        +      "support": {
        +        "description": "Number of experiences supporting the lesson.",
        +        "minimum": 1,
        +        "title": "Support",
        +        "type": "integer"
        +      },
        +      "text": {
        +        "description": "A compact lesson extracted from one or more past executions.",
        +        "title": "Text",
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "text",
        +      "confidence",
        +      "support"
        +    ],
        +    "title": "LessonSummary",
        +    "type": "object"
        +  }
        +}
      • addedOutput schema / properties / experiences
        Added value: +{
        +  "description": "Matching past tasks ordered from best to worst match.",
        +  "items": {
        +    "$ref": "#/$defs/ExperienceSummary"
        +  },
        +  "title": "Experiences",
        +  "type": "array"
        +}
      • addedOutput schema / properties / lessons
        Added value: +{
        +  "description": "Lessons derived from the returned experiences.",
        +  "items": {
        +    "$ref": "#/$defs/LessonSummary"
        +  },
        +  "title": "Lessons",
        +  "type": "array"
        +}
      • removedOutput schema / properties / result
        Removed value: -{
        -  "title": "Result",
        -  "type": "string"
        -}
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "experiences",
        +  "lessons"
        +]
      • changedOutput schema / title
        Previous value: -"search_experienceOutput"New value: +"SearchExperienceResult"
  4. 13 tool updatesv0.3.2
    • Changedapprove_project1 field changed
      • addedInput schema / properties / confirm / description
        Added value: +"true only when the user has explicitly asked to enable OpenReflex for this project; with false (the default) nothing changes."
    • Addedcheck_progress
    • Changedchoose_path2 fields changed
      • addedInput schema / properties / steps / description
        Added value: +"For a custom strategy only: 2-4 short steps, e.g. ['Prototype the toml loader', 'Swap the callers']. Ignored for suggested strategies."
      • addedInput schema / properties / strategy / description
        Added value: +"The strategy being followed: one of the suggested names 'inspect-first', 'test-first' or 'incremental', or a short kebab-case name for your own, e.g. 'spike-then-rewrite'."
    • Addedexplain_decision
    • Changedexplain_node1 field changed
      • addedInput schema / properties / node_id / description
        Added value: +"Id of an Experience Graph node: an experience id from search_experience (starts with 'exp-'), the execution id from get_execution_context, or any node id from an earlier explain_node result."
    • Addedforget_experience
    • Changedget_execution_context4 fields changed
      • addedInput schema / properties / max_context_tokens
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional cap on tokens of tool output added to context; a positive integer. Default: no cap.",
        +  "title": "Max Context Tokens"
        +}
      • addedInput schema / properties / max_minutes
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional cap on active working time in minutes; a positive number. Default: no cap.",
        +  "title": "Max Minutes"
        +}
      • addedInput schema / properties / max_tool_calls
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Optional cap on tool calls for this task; a positive integer. Default: no cap.",
        +  "title": "Max Tool Calls"
        +}
      • addedInput schema / properties / task / description
        Added value: +"The task in one or two plain sentences, e.g. 'Fix the login redirect loop after logout'. Used to find similar past tasks; secrets are redacted before it is stored."
    • Addedget_execution_trace
    • Addedget_project_insights
    • Addedget_reflex_score
    • Removedproject_insights
    • Changedrecord_outcome2 fields changed
      • addedInput schema / properties / evidence / description
        Added value: +"Short proof of the result, e.g. 'pytest tests/test_auth.py passed' or 'user confirmed the fix'. Up to 500 characters; secrets are redacted."
      • addedInput schema / properties / status / description
        Added value: +"'success' when tests, lint or build passed or the user confirmed the result; 'failure' when the task failed or was abandoned."
    • Changedsearch_experience2 fields changed
      • addedInput schema / properties / limit / description
        Added value: +"Maximum number of past tasks to return, 1-20; values outside the range are clamped. Default 5."
      • addedInput schema / properties / query / description
        Added value: +"Words describing the topic, e.g. 'expired token login bug'. Matched lexically against past task descriptions."
  5. 7 tool updatesv0.1.3
    • First observedapprove_project
    • First observedchoose_path
    • First observedexplain_node
    • First observedget_execution_context
    • First observedproject_insights
    • First observedrecord_outcome
    • First observedsearch_experience

TDQS

A4.7/5.0

Scored across 14 tools

Disambiguation5/5

Each tool targets a distinct action or view, and close pairs like explain_decision vs get_reflex_score are clearly separated by output format and explicit 'Not for' cross-references. The tools are easy for an agent to tell apart despite several operating on the same recent-task context.

Naming Consistency5/5

All 14 tools follow a consistent snake_case verb_noun pattern: get_, explain_, search_, check_, choose_, record_, approve_, forget_. No mixed casing or stylistic drift is present.

Tool Count5/5

At 14 tools, the set is well within the ideal range and each tool has a clear role in the OpenReflex workflow. The count feels appropriate for covering planning, execution monitoring, explanation, memory management, and metrics without redundancy.

Completeness4/5

The core lifecycle is well covered: approval, planning, path selection, progress checks, outcome recording, searching experiences, and forgetting tasks. Minor gaps exist, such as no MCP tool to revoke project approval or perform a full memory wipe, both of which are delegated to CLI commands.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Provides AI coding agents with persistent, long-term memory through local semantic search and SQLite storage. It enables agents to save and retrieve architectural decisions or project context across different conversation sessions without requiring cloud services.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides AI coding assistants with persistent project memory to retain architectural decisions, code patterns, and domain knowledge across sessions. It stores data locally in a SQLite database, allowing agents to remember, recall, and manage project-specific context using full-text search.
    7 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Gives AI coding agents persistent memory by storing observations, decisions, and learnings in a local SQLite database with vector search, full-text search, and a rules engine.
    4
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides long-term memory for LLMs via local SQLite storage with hybrid search (BM25, vectors, recency decay), enabling AI coding agents to persist and recall memories across sessions without cloud or API keys.
    53
    MIT