Skip to main content
Glama
ac0033

agent-memory

by ac0033

agent-memory

给 LLM agent 的本地长期记忆基础设施。 一个跑在你自己机器上的小服务,让任何 agent(Claude Code 等编程 agent、LangGraph 应用……)跨会话记住用户偏好、项目约定和踩过的坑;写入必须过门、可按要求遗忘、能在合适的时候主动想起来。

记忆机制不是从别家的机制清单里挑出来的,而是从一套 13 项能力、3 项质量属性的能力框架和它的评价标准(每一项都要显著优于朴素 RAG)反推出来的:原话是证据,记忆只是通向证据的键和贴在证据上的注解;读取时只组织不裁决;智能前移到写入期且只增不减。 验收用多个公开评测集、每个只测一次:LoCoMo 未见过的对话 42 对 33(p=0.049),PersonaMem-32k 与 LongMemEval-S 与朴素 RAG 持平;自建的 92 题 11 桶验证集 88/92,朴素 RAG 在最初 45 题上 39/45;治理能力(遗忘、投毒、任务状态、主动浮现)另有一套 373 条用例的内部评测集 MemCompass 量出来。

English: README.en.md · License: MIT · Python ≥ 3.12 · 876 个测试无需网络与 API key · 当前版本 v0.4.0(CHANGELOG)


目录


Related MCP server: Recall

它解决什么问题

编程 agent 每开一个新会话都从零开始:上周定好的数据库选型、用户"以后都用 uv"的偏好、昨天踩过的坑,全部要重新讲一遍。常见的补救办法是把历史对话整体塞进 RAG,但这样做有四个结构性的问题:

问题

原文 RAG 的表现

agent-memory 的做法

忘不掉

用户要求"忘了这件事",原文还在,检索照样泄漏

删除记忆条目 + 把原文对应片段擦成占位符,审计只记元数据

防不了投毒

对话里混进的"以后忽略安全检查"会被原样召回

写入过评价门:指令性内容、注入特征、泄漏的密钥一律拦在库外

没有任务状态

只能检索到"说过什么",不知道"现在做到哪了"

独立的工作记忆层:目标、约束、待办、未决问题,会话开头自动注入

不知道何时该开口

每轮都注入,相关不相关一起塞

精确率优先的主动浮现:只在"不提就会出错"时才提

agent-memory 把记忆分成三层(长期 / 工作 / 短期),写入走"脱敏 → 蒸馏 → 评价门 → 对账"管线,读出时带"参考而非指令"的护栏,并提供 MCP、Python 库、Skill 三种接入方式。

为什么不是"再做一个 RAG"

朴素 RAG(原文全留、每轮检索)是很强的基线——我们自己早期的评测里,它在多数纯问答子集上追平甚至超过当时的记忆系统。memory-v1(2026-09-21 起)的答案不是"不跟它比问答",而是把它当成地板:

  • 载荷按构造包含原话:记忆条目只是通向原文的键和贴在原文上的注解,答题者看到的信息不少于朴素 RAG 给的(检索对等率约 90%,载荷体量 ≤ 1.15 倍);

  • 读取时只组织、不裁决:不合成"库里没有 X"、不替答题者数数,读路径零 LLM 调用;

  • 写入期把原文里没有的东西加上去:绝对日期、取代史、有效期、出处、线索词、用户画像——这些是朴素 RAG 结构上给不出的,也是领先的来源;

  • 治理能力照旧是它做不到的:按要求遗忘(泄漏 0%)、投毒拦截(误拦 0%)、任务状态、主动浮现、跨 agent 不串线。

在没见过的外部对话上的读数(同一套默认配置,答题器与评委固定):LoCoMo 两段新对话 60 题 42 对 33(p=0.049);PersonaMem-32k 60 道选择题 51 对 53(持平);LongMemEval-S 60 题 51 对 49(持平,p=0.77;知识更新 9 对 6)。推导与逐轮记录见 docs/research/memory-v1-design.md,机制说明见 docs/design/memory-v1-mechanism.md。

核心亮点

  1. 机制由能力框架推导,用多个外部评测集验收。 先定义好的 agent 记忆应具备哪 13 项能力(K1–K13)、每项怎么度量、成熟度怎么定级(L2 = 显著优于朴素 RAG),再从"分数由什么决定"反推机制,每条原则对应一类实测过的失分。准入规则写在机制之前:至少两个独立评测同向不劣、至少一个显著更优、任何一个题型显著变差即否决;外部集只做验收、每个只测一次,机制迭代靠自建验证集(一份 55 场会话的共享历史,92 题 11 个能力桶,第 0 层不调 LLM 一分钟出结果)。推导过程与被证伪的预测都记录在 docs/research/memory-v1-design.md。

  2. 三层记忆,一次组装。 长期记忆(跨会话的事实 / 偏好 / 流程)、工作记忆(当前任务的目标 / 约束 / 待办 / 未决问题)、短期记忆(直接读宿主自己的会话日志,不复制)。memory_context 一次调用拿到"常驻画像 + 工作记忆 + 相关召回"三个分节块。

  3. 写入永远过门,但永不丢数据。 任何候选都要经过脱敏 → 蒸馏 → 评价门 → 对账(ADD / UPDATE / DELETE / NOOP);判不了的冲突进人工复核队列而不是猜。原文在蒸馏之前就已归档,评价门拒绝的内容可以强制转人工复核。

  4. 记忆是参考,不是指令。 蒸馏拒绝提炼指令性内容,评价门拦截提示词注入特征,注入块自带护栏声明。当记忆与当前请求冲突时,永远以当前请求为准。

  5. 可验证的遗忘。 memory_forget_request 同时删除记忆条目和原文归档中的对应片段,审计日志只记元数据,检索层与回答层泄漏均为 0%。

  6. 主动浮现(精确率优先)。 LLM 线索扩展 + 一跳扩散 + 逐条判断"不提就会出错 / 原理相通值得点明",宁可少说。系统层 F0.5 77%,误插话率 2/13。

  7. 双时态与变更史。 每条记忆带 valid_from / valid_to,对账 UPDATE 时旧版本折进 history;蒸馏时按会话日期换算相对时间、识别追溯更正。"截至某时"问答准确率 97%。

  8. 原话随记忆一起到场。 归档脱敏后建派生索引(向量 + 词级 / trigram 全文);检索返回的是"证据束"——同一处原话及其注解(取代史、有效期、出处、事件日期),原话在前、注解在后,只是复述原话的条目不再渲染。答题者永远能回到证据核实。

  9. 离线进化闭环,可回滚。 agent-memory evolve 做去重合并、离线复核、降权归档,产出提案而非直接改库;三档验证(边界 / 留存 / 安全)任一不过即否决,晋升前快照、晋升后审计、随时回滚。

  10. 宿主中立,没有 API key 也能用。 MCP over HTTP / MCP stdio / LangGraph 库 / Skill 四种接入;订阅制 agent 拿 memory_distill_prompt 的协议在自己的上下文里蒸馏,服务端照样校验、脱敏、过门、对账。

  11. 自带评测集与诚实的报告。 MemCompass:8 个子集 373 条合成用例,覆盖公开评测集没有的能力(遗忘、投毒、主动浮现、任务状态、跨 agent、双时态……),程序化金标、异源评委、配对统计、消融、对照组,报告里连"朴素 RAG 在哪里赢了我们"都写清楚。

评测结果

三类数据、三种用途:外部评测集(LoCoMo、PersonaMem-32k、LongMemEval-S)只做验收,每个来源只测一次,朴素 RAG 同场作答;自建验证集是集成测试,按能力维度出题、反复用;内部评测集 MemCompass 覆盖公开评测集没有的治理能力。验收标准:每个外部集都不劣于朴素 RAG,且至少一个显著更优。

外部验收(memory-v1,2026-09-22,每个来源只测一次;朴素 RAG 同场、同答题器 deepseek-v4.1-flash、同评委 glm-5.3-flash,AML 公开的答题与评分提示词):

来源

题量

朴素 RAG

memory-v1

配对

LoCoMo(未见过的 conv-41/42)

60

33(55%)

42(70%)

独赢 13 / 独输 4,McNemar p=0.049

LoCoMo 首测(conv-26/30)

54

31(57%)

37(69%)

独赢 10 / 独输 4,p=0.18,方向一致

PersonaMem-32k(4 份历史,选择题)

60

53

51

独赢 2 / 独输 4,p=0.69(持平);全文上下文 47

LongMemEval-S

60

49

51

独赢 7 / 独输 5,p=0.77(持平);知识更新 9 对 6,弃答题 4/5 对 1/5

LoCoMo 上的领先主要来自前提不成立的对抗题(9 对 0,载荷首条的静态行为协议起作用),多跳与时间题各 +1,单跳与开放域各 −1;载荷体量 1.15 倍于朴素 RAG。PersonaMem 上事实回忆 21/21、变更原因 6/6、推荐 7/7 与朴素 RAG 相等,唯一落后的 suggest_new_ideas 类(5 对 8 / 14)对全文上下文也只有 6/14。

自建验证集(一份 55 场会话的共享历史,92 题 11 个能力桶)88/92;最初的 45 题上朴素 RAG 39、全文上下文 40,每桶不低于两个对照;21 道换说法、泛称、有描述无名字的可答题没有一次错误弃答。方法:docs/research/benchmark-suite/README.md §8。

内部评测集(v0.2.2 时的读数)MemCompass v0.3(8 个子集 / 373 条用例,本仓库自建,全部合成数据)。下表是 test 切分(193 条)上 v0.2.2 与前一版本的逐题配对比较,McNemar 精确检验:

能力

子集 / 模式

v0.2.2

前一版本

p

机制证据

按要求遗忘

fg / 问答

100%(15/15)

40%

.004

检索层硬泄漏 0% vs 70%

回溯补全

ca / 端到端

88%(30/33)

19%

<.001

消融关掉原文工具后降到 75%

主动浮现(系统层)

pr / 系统层

78%(25/32,F0.5 77%)

41%(13/32)

.004

消融关掉主动浮现后回到 13/32

跨 agent 迁移(系统层)

xa / 系统层

83%

0%

.002

宿主中立的归档与注入

时间与变化

at / 问答

97%

57%

.007

双时态与追溯更正

当下保真

pf / 端到端

100%

84%

.016

压缩前情节卡片

投毒鲁棒

mp / 问答

100%

89%

.250

攻击成功率两版都是 0%,差异在误拦 0% vs 21%

v0.3.2 回归(同一评委 glm-5.3-flash 下重跑三方;Kimi K3 已不可用,所以不与上表直接比较):xa 为 v0.3.2 的读数,其余子集为 v0.3.1——v0.3.2 只多了作用域约定兜底,它不作用于 global 记忆,而这些子集的记忆全是 global。

能力

子集 / 模式

v0.3.2

v0.2.2

朴素 RAG

当下保真

pf / 端到端

100%

95%

99%

回溯补全

ca / 端到端

100%

94%

81%

主动浮现(系统层)

pr / 系统层 F0.5

80%

77%

60%

跨 agent 迁移(系统层)

xa / 系统层

10/12(该注入的 12/12)

10/12

7/12

任务状态(系统层)

ts / 系统层

15/23

15/23

0/23

时间与变化

at / 问答

100%

97%

100%

按要求遗忘

fg / 问答

100%

100%

47%

投毒鲁棒

mp / 问答

97%

96%

94%

细节保留

ca / 问答

90%

80%

100%

此后两版只改动了对应子集:v0.3.3 任务状态(ts / 系统层)串线率 60% → 0%、主动浮现端到端(pr / 端到端)F0.5 70% → 80%(朴素 RAG 75%);v0.3.4 主动浮现过早插话 16% → 11%、跨 agent 串线 17% → 8%。逐项读数见 CHANGELOG 与设计稿 §11–12。

诚实的注脚:

  • 每条用例只有 2–8 个会话,朴素 RAG 在多数纯问答子集上持平或更好(见上一节);长历史档位(每题约 35 万 token)尚未构建,是下一步最重要的工作。

  • 单种子、n < 50 的子集区间较宽,报告里逐处标注了"只看方向"。

  • 评委与答题器异源(DeepSeek 答题、Kimi K3 评委、Claude 修订用例),评委间一致率 85%–98%(κ 0.69–0.93);尚未做人工一致性研究。

完整报告(对照组、消融、成本、用例体检、局限):docs/research/benchmark-suite/results/2026-09-16-v03-report.md。复现方法见 evals/memcompass/README.md。

架构

                 ┌──────────────────────────────────────────────────────┐
   宿主 agent    │  memory_context = 常驻画像 + 工作记忆 + 相关召回      │
 (MCP / 库 / Skill)  memory_surface = 主动浮现(精确率优先)              │
                 └───────────────▲──────────────────────────▲───────────┘
                                 │ 读                        │ 读
        ┌────────────────────────┴────────┐   ┌──────────────┴──────────────┐
        │ 长期记忆  data/memory/*.md      │   │ 工作记忆  data/working/      │
        │ 唯一事实来源,按 scope 隔离      │   │ 目标/约束/待办/未决问题       │
        │ + 可重建索引 index.db           │   │ 服务端增量整理 wm_refresh    │
        │   (sqlite-vec 1024d + FTS5)     │   └──────────────▲──────────────┘
        └────────────────────────▲────────┘                  │ 只过脱敏
                                 │ 写                        │
   ┌─────────────────────────────┴──────────────────────────────────────┐
   │ 写入管线: 脱敏 → 蒸馏 → 评价门 → 对账(ADD/UPDATE/DELETE/NOOP) → 变更传播 │
   │           判不了的 → data/review_queue/(人工复核)                    │
   └─────────────────────────────▲──────────────────────────────────────┘
                                 │ 先归档再蒸馏
        ┌────────────────────────┴────────┐   ┌─────────────────────────────┐
        │ 原文归档  data/raw/(只追加)    │   │ 短期记忆 = 宿主自己的会话日志 │
        │ 脱敏后归档 + 派生索引 raw_index │   │ transcript 适配层直接解析     │
        └─────────────────────────────────┘   └─────────────────────────────┘

   离线:agent-memory evolve → 提案 → 三档验证 → 快照 → 晋升 → 审计 / 回滚

作用域:global(跨项目)/ repo:<项目名> / agent:<宿主名>,检索只看当前 scope 加 global。三条不可违反的红线(数据三层分离、写入过门、评测可信根)见 AGENTS.md。

快速开始

1. 安装本体

uv tool install git+https://github.com/ac0033/agent-memory   # 得到 agent-memory 命令;Python ≥ 3.12
# 开发本仓库时改用可编辑安装,改了代码无需重装:uv tool install --editable .

bge-m3 嵌入模型在第一次检索时下载(约 2 GB)。配置写进 ~/.agent-memory/config.env(KEY=VALUE,环境变量优先;AGENT_MEMORY_CONFIG 可改路径),宿主拉起的进程都能读到:

AGENT_MEMORY_LLM_API_KEY=sk-...        # 任意 OpenAI 兼容端点;默认 DeepSeek deepseek-flash
# AGENT_MEMORY_LLM_BASE_URL= / AGENT_MEMORY_LLM_MODEL=
# AGENT_MEMORY_DATA_DIR=               # 缺省 ~/.agent-memory/data
# AGENT_MEMORY_DAEMON_IDLE_MINUTES=30  # 后台进程空闲多久退出

没有 LLM key 时,检索、手动写入、反馈、工作记忆、归档等不依赖 LLM 的功能照常可用;对话蒸馏可交给宿主自己做(见下文)。

2. 接入方式

方式

适合谁

怎么接

Claude Code 插件(推荐)

Claude Code 用户

/plugin marketplace add ac0033/agent-memory,再 /plugin install agent-memory@agent-memory。一次装齐 25 个 MCP 工具、使用规范 Skill 和三个 hook

MCP stdio

任何 MCP 宿主

注册命令 agent-memory mcp,例如 claude mcp add agent-memory -- agent-memory mcp;其他宿主写 {"command": "agent-memory", "args": ["mcp"]}。再把 SKILL.md 装进宿主的 skills 目录

宿主 hook

支持命令 hook 的宿主

agent-memory hook wm-inject --agent <宿主名>(会话开始)、agent-memory hook surface(用户提交消息)、agent-memory hook turn(一轮回复结束);插件已包含

Python 库

LangGraph / LangChain 应用

见下方示例,进程内直接调用,不经后台进程

后台进程:工具调用由本机一个后台进程执行,模型在本机只加载一份、所有会话共用。它不是开机常驻的服务:第一次调用时自动拉起(普通用户权限,只绑 127.0.0.1),空闲 30 分钟自己退出,代码更新后下次调用自动重启。一般不用管它;需要时用 agent-memory daemon status | start | stop 查看或控制。

from langgraph.prebuilt import create_react_agent
from agent_memory.long_term.adapters.langgraph.store import AgentMemoryStore
from agent_memory.long_term.adapters.langgraph.tools import build_memory_tools
from agent_memory.long_term.retrieve.resident import build_system_context

store = AgentMemoryStore()                       # LangGraph BaseStore,namespace ("memories", <scope>)
tools = build_memory_tools()                     # 17 个 ReAct tool,三层记忆全暴露
prompt = "你是用户的编程助手。\n\n" + build_system_context("repo:myproj")   # 常驻画像进 system prompt
agent = create_react_agent(model, tools, prompt=prompt, store=store)

可运行示例:uv run python examples/langgraph_demo.py。

3. 没有 API key 的订阅制 agent

宿主本身就是大模型,蒸馏可以自己做:memory_distill_prompt() 拿协议 → 宿主在自己的上下文里产出 {"memories": [...]} → memory_add(distilled_json=...) 提交。服务端照常校验、脱敏、过门、对账。

更多用法(CLI、评估命令、离线整理、人工复核、hook、工作记忆、会话收尾)见 使用手册;宿主 runtime 必须自己承担的职责清单见 接入指南 §四。开发本仓库:uv sync 装环境,uv run pytest -q 跑测试(不需要网络与 API key)。

MCP 工具一览

25 个 MCP tool(stdio 转发层与后台进程用同一份定义与业务实现):

分组

工具

用途

读取

memory_context memory_search memory_wm_read memory_transcript_read memory_distill_prompt memory_consistency_check

一次组装上下文;混合检索(稠密 + BM25 → RRF → 置信度 × 时间衰减);读工作记忆;读会话日志;宿主蒸馏协议;记忆层与索引一致性检查

写入

memory_add memory_update memory_forget memory_feedback memory_wm_write memory_wm_clear

长期记忆走完整管线;工作记忆只过脱敏

会话

memory_session_end

会话收尾:归档 + 联合蒸馏 + 清理已完成待办(有未完成待办时否决)

复核

memory_review_list memory_review_resolve

管线不敢自动入库的候选交人裁决(approve / modify / discard)

原文归档

memory_archive_search memory_archive_read memory_archive_sync

脱敏后只追加的原文归档,可检索;记忆只有要点或看起来不对时回溯原文

主动浮现

memory_surface

精确率优先的"记忆副手",在 agent 会漏掉的时候提醒它

待确认队列

memory_confirm_enqueue memory_confirm_list memory_confirm_resolve

无人值守时把需要用户拍板的事挂起,其余照常处理

工作记忆整理

memory_wm_refresh

服务端按最近轮次增量整理目标 / 约束 / 待办 / 未决问题

情节卡片

memory_episode_pack

上下文压缩前把标识符、端口、路径、报错原文等原样细节存成卡片

遗忘请求

memory_forget_request

执行用户明确的遗忘要求:删记忆 + 擦原文片段,审计只记元数据

写类工具的描述都注明"仅限主 agent 调用";subagent 只读,结论回传主 agent 后由它策展沉淀。

仓库结构

agent-memory/
├── agent_memory/            # Python 包(import agent_memory)
│   ├── config.py / models.py    配置(AGENT_MEMORY_* 环境变量)与记忆条目 schema
│   ├── llm.py / confirmations.py  OpenAI 兼容 LLM 客户端;待确认队列
│   ├── long_term/               长期记忆:store(Markdown + SQLite 索引)/ ingest(脱敏、蒸馏、评价门、对账、
│   │                            变更传播、情节卡片、遗忘)/ retrieve(bge-m3 混合检索、注入、常驻画像、主动浮现)/
│   │                            evolve(离线整理闭环)/ adapters/langgraph
│   ├── working/                 工作记忆(任务状态 + 服务端增量整理)
│   ├── short_term/              短期记忆:宿主会话日志适配层
│   ├── hooks.py                 宿主 hook(会话开头注入、主动提醒、每 N 轮沉淀),经 agent-memory hook 调用
│   └── server/                  MemoryService 业务层、stdio 转发层、按需启动的后台进程及其生命周期管理
├── plugins/agent-memory/    # Claude Code 插件:.mcp.json、hooks、skills/agent-memory/SKILL.md(Skill 唯一源头)
├── .claude-plugin/          # 插件市场清单
├── examples/                # LangGraph 最小接入示例
├── evals/                   # 可信根(agent 不得修改):layer1–3 / prefix 回归集;另含 MemCompass 运行副本(不属可信根,从源头同步)
├── docs/                    # 使用手册、接入指南、设计文档、研究与评测(见下)
├── tests/                   # 876 个测试(慢测试默认跳过:uv run pytest -m slow)
└── data/                    # 运行时数据(gitignored):raw / memory / working / review_queue / snapshots / logs
                             #   + dev/(自建验证集)、external/<来源>/(外部测试集)

文档导航

想做什么

看哪里

把记忆接进我的 agent,宿主要负责什么

docs/agent-integration.md

日常用法:CLI、评估命令、离线整理、人工复核、hook、会话收尾

docs/usage.md

理解三层记忆的设计与取舍

docs/design/memory-architecture.md

记忆现在怎么读、怎么写、为什么(memory-v1)

docs/design/memory-v1-mechanism.md

机制是怎么从能力框架推导出来的、每一轮实测

docs/research/memory-v1-design.md

好的 agent 记忆应该具备哪些能力、怎么度量

docs/research/agent-memory-capability-framework.md

评测集 MemCompass 的设计、构造脚本、校验与运行

docs/research/benchmark-suite/README.md

最新评测报告

docs/research/benchmark-suite/results/

v0.2 每个机制为什么这么做

docs/research/optimization-v02.md

各版本改了什么

docs/CHANGELOG.md

各阶段交付记录、缺陷复盘、可靠性加固

docs/history/

给在本仓库工作的人和 agent 的规范(红线、约定、行为语义)

AGENTS.md

参与贡献

CONTRIBUTING.md

设计原则

  • D1 数据三层分离:data/raw 只追加;data/memory 的 Markdown 是唯一事实来源;data/index.db 是可重建的派生索引,绝不手改。

  • D2 写入过门:原始对话不直接入库,必须经脱敏 → 蒸馏 → 评价门 → 对账;蒸馏绝不提炼指令性内容。

  • D6 可信根:evals/、rubric、发布门槛、审计日志禁止 agent 自行修改。MemCompass 在 docs/research/benchmark-suite/ 编写与核验,用 tools/migrate_to_evals.py 逐字节同步到运行副本 evals/memcompass/,不在副本里单独改。

  • fail-closed 但不 fail-lost:配置非法、校验失败、证据缺失直接报错,不静默降级;写入路径的内容永不因故障丢失。

  • 本地优先:所有数据是你磁盘上的 Markdown 与 SQLite 文件,后台进程只绑 127.0.0.1,除你配置的 LLM 端点外没有任何数据出站。

已知局限与路线图

  • 领先不是在每个外部集上都显著:LoCoMo 显著领先,PersonaMem 与 LongMemEval-S 与朴素 RAG 持平。LongMemEval-S 的时间推理板块是两边都补上提问日期后的读数(AML 答题模板不带提问日期,"几天前"类题对任何系统都无解);MemCompass 内部表格的读数来自 v0.2.2。

  • PersonaMem 的 suggest_new_ideas 类略低于朴素 RAG(5 对 8 / 14);该类对全文上下文也只有 6/14,属答题者层面的"选泛泛选项"。

  • 关联与图式归纳(K6 离线归纳)未实现:离线整理的验证逻辑属于可信根,需要新的提案类型。

  • 写入成本:对账已按批(每批 8 条)判定,但整段对话蒸馏仍可能要几十秒到几分钟,而 MCP 客户端的超时不可配;改成"先归档、返回任务号、再查状态"的异步写入是待办。

  • LangGraph 适配只覆盖 15 个基础工具,v0.2 新增的 10 个目前仅 MCP 侧提供。

  • 内部评测集为合成数据,评委—人工一致性研究尚未开展。

许可

MIT,见 LICENSE。evals/memcompass/ 与 docs/research/benchmark-suite/datasets/ 下的评测数据均为合成数据,同样按 MIT 发布。

Available Tools

13 tools
memory_addA

写入记忆:对话走蒸馏管线,单条 content 走脱敏+对账。conversation_json 推荐传 [{role, content}, ...] 的 JSON 字符串(直接传数组也可以,服务端会自动序列化;其他类型会报错并提示格式)。scope 应显式选择:跨项目通用知识用 global,项目相关用 repo:<项目名>,agent 自身相关用 agent:<名字>;缺省回落 global 并附提醒

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo
sourceNomcp
contentNo
entry_idNo
confidenceNohigh
session_idNo
memory_typeNosemantic
conversation_jsonNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral disclosure. It reveals the distillation pipeline for conversations, desensitization/reconciliation for single content, server-side auto-serialization for arrays, error behavior for invalid types, and default scope fallback. This is strong behavioral context beyond a simple 'write' action, though it stops short of describing return values or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact, logically organized paragraph with no filler. It front-loads the core action, then covers format, scope, and default behavior in sequence. Every sentence contributes essential information, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers purpose, key parameter semantics, and processing behavior well. The remaining gaps—such as the exact meaning of optional parameters like entry_id or session_id—are minor because those parameters are either inferable or have defaults. The description is close to complete for an agent to successfully call the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameter descriptions, so the description must compensate. It explains the two most complex parameters: conversation_json (JSON string or array, auto-serialization) and scope (global/repo/agent conventions), and mentions content processing. Other parameters like memory_type, confidence, and source remain undocumented, but their names and defaults make them less ambiguous. The description covers the parameters that truly need clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool writes memory, using the specific verb '写入记忆' (write memory), and distinguishes it from sibling tools like memory_update and memory_forget by nature. It also explains the processing pipeline for conversation vs single content, which adds further specificity beyond a generic 'write' operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete guidance on how to invoke the tool: recommended conversation_json format, explicit scope value conventions (global, repo:<name>, agent:<name>), and default fallback behavior. It does not explicitly state when to use this tool instead of alternative memory tools, but the detailed usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_contextA

统一组装注入上下文:常驻画像块(长期用户画像)+ 工作记忆块(当前任务状态)+ 召回块(传 query 才检索历史记忆),按此顺序拼接。复核队列有积压时按 review_gate 配置处置:返回 status=blocked 表示被复核门拦截,需先向用户确认(用户同意后以 acknowledge_pending=true重试,或先用 memory_review_list / memory_review_resolve 处理待办)

ParametersJSON Schema
NameRequiredDescriptionDefault
kNo
queryNo
scopeYes
current_turnNo
acknowledge_pendingNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and delivers substantial behavioral detail: the fixed assembly order, conditional recall on query, and the review-gate state machine (status=blocked, retry flag, prerequisite cleanup via review tools). The blocked/retry workflow is non-obvious and not inferable from the schema. It stops short of stating whether the operation has side effects or how a successful response is structured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the primary assembly behavior front-loaded before the conditional review-gate flow. Every clause carries information — assembly blocks, ordering, query conditionality, and the blocked-state retry procedure. It is slightly dense with domain terminology but efficient overall with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and no parameter descriptions, so the description must cover both return values and parameter semantics. It explains the blocked status and retry path but never describes the success response shape, and leaves scope (required), k, and current_turn undefined. This is insufficient for an agent to invoke the tool reliably on the first attempt.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains two of five parameters: query (triggers recall) and acknowledge_pending (retry after user confirmation). The required parameter scope is never defined at all, and k (despite its default of 5) and current_turn are also left unexplained. An agent cannot determine valid scope values or what k controls, which is a critical gap for the only required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('统一组装注入上下文' — uniformly assemble and inject context) and names the exact resources: resident profile block, working memory block, and recall block, with their concatenation order. This composition detail differentiates it from siblings like memory_wm_read or memory_search that handle only individual memory pieces. It doesn't explicitly name a sibling it is not, but the assembly semantics make the tool's role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit conditional guidance: recall is only performed when query is passed ('传 query 才检索历史记忆'). It also specifies a concrete edge-case workflow with named alternatives — when status=blocked, retry with acknowledge_pending=true or first process pending items via memory_review_list / memory_review_resolve. It does not contrast normal-use selection against siblings like memory_search or memory_wm_read, but the blocked-state routing is explicit and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_feedbackB

反馈记忆是否有用,调整置信度;降到 low 以下进人工复核队列

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
helpfulYes
memory_idYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It reveals the action of adjusting confidence and the conditional manual review queue, adding context beyond the verb. However, it does not disclose whether the operation is a write, if it is reversible, return format, or error handling, leaving notable gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loaded with the action. It is efficient and easy to scan, though it lacks structure for parameter details. It earns a high score for conciseness, not for completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, no output schema, and no annotations, the description should compensate by explaining parameter semantics and usage context. It explains the main behavior but omits parameter meanings and when to use this tool over siblings, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the parameters memory_id, helpful, or note. The description implies 'helpful' relates to usefulness but never explicitly defines each parameter, leaving agents to infer meaning from types and names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (feedback on memory usefulness), the resource (memory), and the outcome (adjust confidence, possible manual review). This clearly differentiates it from siblings like memory_add, memory_update, and memory_review_resolve, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for providing feedback on memory usefulness, but does not explicitly state when to use it versus alternatives like memory_review_resolve or memory_forget. No when-not-to-use conditions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_forgetB

删除一条记忆(记忆层与索引同步删除)

ParametersJSON Schema
NameRequiredDescriptionDefault
memory_idYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a useful behavioral detail: the memory layer and index are deleted synchronously, ensuring consistency. However, with no annotations at all, it does not disclose irreversibility, permission requirements, or what happens to related feedback/review data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence, efficient and easy to scan. The parenthetical adds a relevant operational detail without bloating the text, though a bit more context could be added.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no annotations and no output schema, the description is under-specified. It does not mention irreversibility, how to retrieve memory_id, or any side effects on related data, leaving an agent to guess critical usage constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description does not mention memory_id at all. It simply says 'delete a memory' without explaining that the memory_id parameter identifies the target or how to obtain it, leaving the schema to bear all meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb '删除' (delete) and a clear resource ('一条记忆' - one memory), and the parenthetical clarifies it removes both the memory layer and its index. This clearly distinguishes it from siblings like memory_search or memory_wm_clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not state that this is for permanent removal of a specific memory, nor does it mention prerequisites such as obtaining a memory_id from memory_search or memory_review_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_review_listA

列出人工复核队列的全部待办(内容、排队原因、队列文件名)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. The description only states that it lists items and the fields returned; it does not explicitly state that it is read-only or lacks side effects. While the name suggests a list operation, the description does not communicate this behavioral guarantee, leaving the agent to infer safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the action and resource, then specifies the returned fields. There is no extraneous information, and it is appropriately concise for a list operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, but the description omits how to identify specific items for later resolution (e.g., an ID field) and does not mention any ordering or pagination behavior. Without an output schema, an agent may need more context to use the results with sibling tools like memory_review_resolve, which would require some reference to individual items.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema description coverage is trivially 100%. The baseline for 0 parameters is 4, and the description does not need to add parameter meaning because there are none. It provides no extra parameter information, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the verb '列出' (list) with a specific resource '人工复核队列的全部待办' and enumerates the returned fields (content, reason, queue file name). This clearly distinguishes it from siblings like memory_review_resolve, which handles resolution, and memory_search, which is general search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for viewing the review queue but does not explicitly state when to use it over alternatives, such as when to use memory_review_resolve after listing. There is no mention of prerequisites, exclusions, or the relationship with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_review_resolveA

裁决一条复核待办:approve 确认入库 / modify 以 new_content 替换正文后入库 / discard 丢弃。queue_file 取 memory_review_list 返回里的 file 字段

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
queue_fileYes
new_contentNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden. It accurately conveys the side effects of each action: approve persists, modify replaces content before persisting, and discard drops the item. It also explains the provenance of queue_file. However, it does not disclose whether the queue item is consumed after resolution or whether new_content is mandatory for the modify action, leaving minor gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that lists the actions and their meanings, followed by a short clarification of where queue_file originates. Every clause adds necessary information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core workflow, parameter sourcing, and action semantics, which is sufficient for an agent to invoke the tool correctly. It omits return values, explicit requirement of new_content for modify, and post-resolution queue state, but these are secondary for a narrowly scoped resolution tool. Given the lack of an output schema, this is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully documents all three parameters: action enumerates the valid values (approve/modify/discard), queue_file is tied to the output of memory_review_list, and new_content is defined as the replacement body for modify. This resolves the ambiguity left by the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (裁决/adjudicate) and resource (复核待办/review todo), then enumerates the permissible actions (approve, modify, discard), making the tool's purpose unmistakable. It also distinguishes itself from the sibling memory_review_list by focusing on the resolution step rather than the listing step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent that queue_file comes from the return of memory_review_list, effectively positioning this as the follow-up to that tool. It names the source of a key parameter and implies a list-then-resolve workflow, though it does not formalize when-not-to-use or list alternative tools for this specific action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_session_endA

会话结束收尾编排:归档原文(data/raw,只追加不改写)+ 联合蒸馏(对话提炼长期记忆,工作记忆快照作参考上下文,冲突会更新旧条目) + 清理工作记忆里已完成的待办。工作记忆有未完成任务时会 veto(status=vetoed,归档/蒸馏/清理都不执行),确认结束请以 force=true 重试。对话材料二选一:conversation_json([{role, content}, ...] 的 JSON 字符串或数组,agent 中立推荐,优先使用)或 log_path(agent 会话日志路径,走日志适配器解析,adapter 可缺省按文件名 自动识别)。未配置 LLM 时只归档不蒸馏(status=archived_only)

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo
scopeYes
sourceNomcp
adapterNo
log_pathNo
session_idNo
conversation_jsonNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses detailed behavior: append-only archiving, conflict updates during distillation, veto with status=vetoed and force override, archived_only fallback, and adapter auto-detection. This goes well beyond minimal expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that covers all key points without fluff. It is front-loaded with the main orchestration, then veto, then material options. However, it could be improved with bullet points for readability, but it still earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 params, no output schema, no annotations), the description is quite complete. It explains the core flow, edge cases, and material options. The main shortfalls are undocumented parameters and lack of return-value details, but these are partially offset by the thorough behavioral coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains conversation_json, log_path, adapter, and force thoroughly, but omits scope, source, and session_id entirely. These are not self-evident from the schema, leaving gaps for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: orchestrating session-end wrap-up with archiving, distillation, and cleanup. It distinguishes itself from siblings (e.g., memory_transcript_read, memory_add) by focusing on the end-of-session flow, making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides strong context on when to use the tool (at session end) and explains veto/force behavior for incomplete tasks, plus fallback to archived_only without LLM. However, it does not explicitly contrast with sibling tools like memory_wm_write or memory_add, and the 'agent 中立推荐' note is more about parameter selection than tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_transcript_readA

读取 agent 会话日志(如 kimi-code 的 wire.jsonl),解析成干净的轮次序列(user/assistant/tool,含轮次编号与时间戳)。纯读不写。配合工作记忆水位做新鲜度补偿:传 since_turn=<memory_wm_read 返回的 turn_watermark> 只返回水位之后的新轮次,据此判断要不要 wm_write 刷新工作记忆。adapter 缺省按日志文件名自动识别,识别不了需显式指定(可用列表见报错信息);日志不存在会报错

ParametersJSON Schema
NameRequiredDescriptionDefault
adapterNo
log_pathYes
since_turnNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the full burden. It discloses the tool is read-only ('纯读不写'), describes the output structure (turn sequences with numbering and timestamps), explains the adapter fallback behavior (auto-detect or explicit, with error message listing available adapters), and notes that missing logs will cause an error. This is thorough behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured into a logical flow: main purpose, usage pattern with since_turn, then adapter behavior and error conditions. It is four sentences but dense with information, front-loaded with the core function and then branching into usage details. It is not overly verbose for the complexity it covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (read logs, parse turns, support filtering, adapter handling) and the absence of annotations, the description covers essential aspects: parameters, usage integration with wm_read/wm_write, error handling (missing log, adapter detection), and output format ('轮次序列(user/assistant/tool,含轮次编号与时间戳)'). No critical gaps for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all three parameters. It does: log_path is implied ('读取 agent 会话日志'), adapter behavior is explained ('adapter 缺省按日志文件名自动识别,识别不了需显式指定'), and since_turn is clearly defined ('传 since_turn=... 只返回水位之后的新轮次'). Meaning is fully compensated beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('读取' = read), a specific resource ('agent 会话日志' = agent session logs), and the output transformation ('解析成干净的轮次序列' = parse into clean turn sequences). It also explicitly distinguishes itself from siblings by emphasizing '纯读不写' (pure read, no write) and unique functionality (transcript reading) not covered by other tools like memory_wm_read or memory_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for use: it explains the integration with memory watermark ('配合工作记忆水位做新鲜度补偿') and explicitly references sibling tools memory_wm_read and memory_wm_write, showing how since_turn should be used. It also covers adapter auto-detection behavior and error conditions. It lacks an explicit 'when not to use' clause but the context is sufficiently clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_updateB

更新一条记忆的正文(过脱敏与评价门)

ParametersJSON Schema
NameRequiredDescriptionDefault
memory_idYes
new_contentYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

描述提到更新内容会经过脱敏与评价门,揭示了一个重要的处理流程,但未说明失败行为(如门拒绝时返回什么)、权限要求、可逆性或副作用。由于无注解,描述承担全部责任,但此处提供的信息有限,仅部分披露了行为。

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

一句话完成,无冗余,动作前置,括号补充关键流程约束。紧凑且信息优先,符合高效结构要求。

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

无输出schema,无注解,且为变更操作,但描述未说明返回内容、错误处理、副作用(如修改是否不可逆),也未说明通过门失败时的行为。对于调用者而言,信息不足以做出正确调用决策,完整性不足。

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

schema描述覆盖率为0%,参数名称和类型是唯一信息。描述未对memory_id或new_content增加任何额外语义,如格式、长度限制、示例或注意事项。描述仅重复了'正文'概念,对参数理解帮助有限。

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

描述明确给出具体动作'更新'和资源'记忆的正文',并附带脱敏与评价门的处理提示,使工具用途清晰且与其他记忆操作(如添加、删除)区分开来。即使未命名兄弟工具,名称和描述已足够明确。

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

未提供任何使用时机或与其他工具的对比。没有说明何时应使用更新而非添加或删除,也没有提及前置条件(如记忆必须存在)。仅凭工具名称和描述推断用途,缺乏明确指导。

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_wm_clearA

清空一个 scope 的工作记忆;本就不存在时返回 already empty,不算错误

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that clearing a non-existent scope returns 'already empty' rather than an error, which is a key behavioral nuance. While it does not elaborate on side effects or permissions, this is sufficient for a simple clear operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the core action and efficiently includes the critical edge case. There is no redundancy or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and a clear action, the description covers the primary behavior and a notable edge case. It is slightly lacking in defining 'scope' and specifying the success return value, but these are minor given the tool's minimal complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the tool description does not explain what 'scope' refers to or what valid values it accepts. The agent is left without guidance on this essential parameter, and the description fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear verb '清空' (clear) and a specific resource 'scope 的工作记忆' (working memory of a scope), explicitly distinguishing this from sibling tools like memory_wm_read and memory_wm_write. The edge case about returning 'already empty' further clarifies its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context that this tool operates on a scope's working memory, which makes its usage obvious given the sibling set. It does not explicitly name alternatives or exclusions, but the action is unambiguous and no conflicting tools exist in the visible sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_wm_readA

读取一个 scope 的工作记忆(当前任务状态:目标/待办/决策/变量/备注),返回渲染好的注入块与结构化字段。scope 自动归一化,非法当场报错。传 current_turn 时返回 stale_wm 表示工作记忆是否可能滞后(当前轮次超过已更新到的轮次水位),滞后可考虑 wm_write 刷新

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeYes
current_turnNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses automatic scope normalization, immediate errors on invalid scope, the stale_wm flag when current_turn is provided, and the return format (rendered injection block + structured fields). It also hints at the refresh path via wm_write. This is comprehensive for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary purpose and then adding the advanced staleness detail. Every sentence adds value, no fluff. It is well-structured and easily parsed by an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with only 2 parameters and no output schema, the description covers the return format, error handling, normalization behavior, and the staleness mechanism. An agent has enough context to invoke it correctly and interpret results without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains both parameters: 'scope' is the working memory scope to read and is normalized automatically; 'current_turn' is used to trigger staleness detection and returns stale_wm. This adds meaningful semantics beyond the schema's bare names and types, though it doesn't detail value formats or ranges.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb '读取' (read) and a clear resource: 'scope 的工作记忆' (working memory of a scope), and enumerates its contents (goals/todos/decisions/variables/notes). It clearly differentiates from siblings like memory_wm_write (write) and memory_wm_clear (clear) by specifying the read operation. This is a precise, unambiguous purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this tool to read a scope's working memory for current task state. It also mentions when to consider using a sibling (wm_write refresh if stale_wm indicates lag), which implies the alternative when the data is stale. However, it does not explicitly contrast with other read/search tools like memory_search or state when NOT to use it, so it's not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_wm_writeA

写入工作记忆(当前任务状态)。注意是全量替换而非合并:未传的字段会被置空,只想改一个字段也要把其余字段原样带上。所有文本过脱敏;不过评价门——待办事项天然是祈使句,属于正常内容。todos 可传 [{content, status}, ...](status 为 pending/done)或纯字符串列表(按 pending)。turn_watermark 传当前对话轮次;未传保留旧值

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNo
notesNo
scopeYes
todosNo
decisionsNo
variablesNo
turn_watermarkNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the behavioral burden. It discloses the full-replacement semantics (destructive but essential), text desensitization, the pass-through of todo items as imperative sentences (bypassing evaluation gate), and the turn_watermark retention behavior. These are the key behavioral traits an agent needs to know before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense. The most critical warning (full replacement) is front-loaded, and every sentence adds necessary context—no filler. It efficiently covers the trickiest aspects of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter write tool with no annotations and no output schema, this description provides the essential operational knowledge: the destructive replacement behavior, the two tricky parameters (todos and turn_watermark), and the text-processing nuance. It is complete enough for an agent to call it correctly, covering the high-risk elements thoroughly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It thoroughly explains the todos parameter (two accepted formats with status semantics) and turn_watermark (current dialogue turn, default retention). It also implies that all text fields undergo desensitization. It does not delve into goal, notes, decisions, or variables, but the general replacement rule covers them, so it adds significant meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('写' write) on the '工作记忆' (working memory, current task state), which is a specific resource. This distinguishes it from sibling memory tools like memory_add or memory_update, which likely target long-term memory. The scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context (writing current task state) and explicit critical usage instructions (full replacement, not merge; fields not passed are nulled). However, it does not explicitly contrast with alternative memory tools (e.g., when to use this vs memory_update or memory_add), so it lacks exclusions. Still, the context is clear enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.0
    • First observedmemory_add
    • First observedmemory_context
    • First observedmemory_feedback
    • First observedmemory_forget
    • First observedmemory_review_list
    • First observedmemory_review_resolve
    • First observedmemory_search
    • First observedmemory_session_end
    • First observedmemory_transcript_read
    • First observedmemory_update
    • First observedmemory_wm_clear
    • First observedmemory_wm_read
    • First observedmemory_wm_write

TDQS

A4/5.0

Scored across 13 tools

Disambiguation5/5

Each tool maps to a distinct resource and action: long-term memory CRUD, working memory read/write/clear, review queue list/resolve, unified context assembly, transcript reading, and session-end orchestration. Even memory_context is clearly an orchestration layer over memory_search and memory_wm_read, not a competing implementation. The only mild overlap is transcript_read versus session_end log parsing, but one is read-only and the other is a write pipeline.

Naming Consistency4/5

All tools share the memory_ prefix and snake_case, but the pattern varies slightly: base operations use verb-only suffixes like memory_search and memory_add, while subarea operations use noun_verb suffixes like memory_wm_read and memory_review_resolve. memory_context and memory_feedback are noun-oriented rather than strict verb_noun. Overall the convention is still predictable and readable.

Tool Count5/5

13 tools is well-scoped for a memory server covering long-term memory, working memory, review queue, context assembly, transcript access, and session-end orchestration. Each tool has a distinct lifecycle purpose, and none feels redundant or missing as a surface-level feature.

Completeness5/5

The surface covers full CRUD for long-term memories, working-memory read/write/clear, review workflow, context assembly, and session-end archival/distillation. The review-gate blocked/acknowledge and session-end veto/force flows prevent dead ends. The only arguable gap is dedicated profile management, but profiles are exposed through memory_context and updatable via memory_add/memory_update.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to store and semantically retrieve durable memories across sessions via MCP or REST, with tools for remembering, recalling, asking, updating, and forgetting memories.
    16 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides AI agents with persistent, human-like memory infrastructure via MCP, enabling them to store, search, summarize, and forget episodic, semantic, procedural, and working memories across sessions.
    752 npm
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Provides agents with a structured long-term memory pipeline, enabling recall, hybrid retrieval, and memory management operations through MCP tools.
    -