doc-fine-tuning-mcp
doc-fine-tuning-mcp —— opencode オフィス文書精密修正アノテーター
opencode 用 MCP サーバーです。LLM がオフィス文書に対して精密な修正を行う必要があるとき、独立したブラウザアプリウィンドウで文書を視覚的に開き、 あなたが段落全体をクリック(または連続する文字列をドラッグ選択してより精密にアノテーション)し、各位置の修正プロンプトを入力します。一度に複数箇所をアノテーションできます。完了後、エージェントに返却し、 LLM があなたのプロンプトに基づいて各箇所を順に修正します。アノテーションウィンドウはエージェントのキャンセル/終了に伴って自動的に閉じ、同一文書の複数回のアノテーションでは同じウィンドウを再利用して自動リロードします。
対応形式: .docx / .xlsx / .pptx(OOXML)。v1 では旧形式(.doc/.xls/.ppt)は対象外です。
ワークフロー
用户对 opencode 说:把 D:\报告.docx 的第 3 段和标题改得更正式
│
▼
opencode 插件 doc-edit-listener 检测到"文档精细修改意图",注入引导
│
▼
LLM 调用 doc-edit_annotate_document("D:\报告.docx")
│ → 本地 HTTP 服务启动,独立浏览器应用窗口打开标注页 http://127.0.0.1:<port>/?session=xxx
│ (同一文档再次标注:复用该窗口自动重载,不重开新窗口)
▼
你在页面上:看到文档 → 点击整段/单元格/形状(或拖选连续字符串)→ 输入提示词 → 继续标注下一处 → 点【完成】
│ (标注以 loc + prompt 形式提交给服务端;完成后 agent 窗口重新聚焦)
▼
LLM 调用 doc-edit_wait_for_annotations 拿到全部标注
│ (若窗口被关闭:返回 status=closed + reason,LLM 据此决定重开或询问你)
▼
LLM 对每个标注:doc-edit_read_location 读取原文 → 依据提示词生成新内容
│ → doc-edit_apply_edit 应用修改(首次自动备份 .bak-<时间戳>)
▼
标注窗口自动重载(服务端推送),展示修改后的最新文档供你检查
│ → LLM 再次 doc-edit_wait_for_annotations 等待下一轮标注(多轮循环)
▼
你点【完成】且无标注 / 点【取消】/ 关闭窗口 → 修改流程结束Related MCP server: MCP Word Commander
アーキテクチャ
opencode
├─ 插件 doc-edit-listener(监听消息 → 检测修改意图 → 注入引导)
└─ MCP 客户端 ──stdio──► doc-edit MCP server(Node/TS, @modelcontextprotocol/sdk)
├─ 标注工具:annotate_document / wait_for_annotations / cancel_session
├─ 编辑原语:read_structure / read_location / apply_edit
├─ 本地 HTTP 服务(127.0.0.1:动态端口)→ H5 标注页 + API
└─ Python 编辑引擎(python-docx / openpyxl / python-pptx)位置特定の一貫性: ページ上でクリックした位置(
loc)は JS 側で OOXML を走査して生成され、Python エンジンは同じ走査アルゴリズムで解析・特定します。tests/parityテストにより、両側の loc→テキストマッピングが完全に一致することを保証します(特に Word の表内段落、PPT の複数シェイプ)。実行主体は LLM: 新しいコンテンツを生成する各ステップは LLM があなたのプロンプトに基づいて行います。MCP サーバーは「ページを開く、アノテーションを収集する、位置に基づいてファイルを読み書きする」だけを担当します。
インストール
1. プロジェクトをクローンして依存関係をインストール
cd D:\develop\doc-fine-tuning-mcp
npm install
# Python 编辑引擎(创建 .venv 并安装 python-docx/openpyxl/python-pptx)
cmd //c scripts\\setup_venv.bat2. ビルド
npm run build # 编译 src → dist/(MCP server)
cd web && npm install && npm run build # 构建 H5 标注页 → web/dist/3. MCP サーバーを opencode に登録
~/.config/opencode/opencode.jsonc の mcp ブロックに追加します(開発モード、ローカルのビルド成果物を指す):
"doc-edit": {
"type": "local",
"command": ["node", "D:/develop/doc-fine-tuning-mcp/dist/index.js"],
"enabled": true
}リリースモード(任意):
npm packで tgz を生成後、npx -y --package <パス>/doc-fine-tuning-mcp-<バージョン>.tgz doc-fine-tuning-mcpを使用します。プロジェクト内の他の MCP と同様です。
4. リスナープラグインのインストール
プラグインは2つのファイルに分かれています: doc-edit-listener.ts と、その依存する検出モジュール lib\detect.ts。
opencode のプラグインディレクトリは、その下の .ts ファイルを自動的にプラグインとしてスキャンしますが、サブディレクトリは再帰的にスキャンしません。
そのため、detect.ts を lib\ サブディレクトリに置いても、独立したプラグインとして誤って読み込まれることはありません。
copy plugin\doc-edit-listener.ts %USERPROFILE%\.config\opencode\plugin\doc-edit-listener.ts
mkdir %USERPROFILE%\.config\opencode\plugin\lib
copy plugin\lib\detect.ts %USERPROFILE%\.config\opencode\plugin\lib\detect.ts⚠️ プラグインモジュールはプラグイン自体のみをエクスポートできます(default)。
doc-edit-listener.tsに名前付き関数エクスポートを追加しないでください。 そうしないと、opencode がそれらを追加のフック/プラグインとみなし、読み込みに失敗して "Unexpected server error" が発生します (hooks が空にされ、Provider.defaultModel のクラッシュに連鎖します)。検出ロジックは必ずlib/detect.tsに置いてください。
プラグインは @opencode-ai/plugin に依存します(~/.config/opencode/node_modules に 1.18.11 が組み込み済み)。opencode を再起動すると有効になります。
5. エンドツーエンドの自己チェック
npm test # 引擎 / parity / mcp 客户端 / 插件 全部测试
node scripts/e2e-verify.ts # 输出 PASS 即闭环可用MCP ツールの説明
ツール | 入参 | 説明 |
|
| 独立したアノテーションウィンドウを開き、アノテーションセッションを作成/再利用し、 |
|
|
|
|
| アノテーションセッションをキャンセルします |
|
| 文書のアウトライン(docx は段落ごと / xlsx は各シートのサンプル / pptx は各ページのシェイプ) |
|
| 指定位置の原文 + 隣接コンテキスト(docx の |
|
| 修正を適用します(docx の |
|
|
|
|
| 全文検索・置換(docx のみ、通常の文字列で正規表現ではありません; |
|
| 修正を一括プレビューします(メモリ内で適用し、ディスクには書き込みません。各箇所の |
|
| バージョン履歴を一覧表示します(各 |
|
| 指定バージョンにロールバックします(ロールバック前に現在の状態をスナップショットし、可逆です) |
mode: replace / append / prepend / insert_after / delete。
style(任意): { bold, italic, sizePt, color }。
注:
template_replace/find_replaceの置換は run レベルの書式保持を行います——置換範囲が複数の異なる書式の run にまたがる場合、新しいテキストはソース run の文字数重みで分割され、各セグメントが対応する run の書式を継承します(最初の run の書式のみに退化しません)。
位置記述子 loc
ページのクリックで生成される loc は「どこを修正するか」の唯一の証明書であり、6つのタイプがあります:
| { kind: "docx-paragraph", paraIndex } // Word 段落(body 文档序,0-based)
| { kind: "docx-cell", tableIndex, rowIndex, colIndex, paraIndex } // Word 表格内段落
| { kind: "xlsx-cell", sheet, row, col } // Excel 单元格(1-based,同 A1)
| { kind: "xlsx-range", sheet, row1, col1, row2, col2 }
| { kind: "pptx-shape", slideIndex, shapeIndex } // PPT 形状(1-based)
| { kind: "pptx-shape-paragraph", slideIndex, shapeIndex, paraIndex }使用上の注意とヒント
LLM は: まず
annotate_documentでユーザーにアノテーションさせ →wait_for_annotationsでアノテーションを取得 → 各アノテーションについてread_locationで原文を確認 → プロンプトに基づいて新しいコンテンツを生成 →apply_editを実行します。修正位置を推測しないでください。バージョン履歴による修正: 各
apply_editの前に文書は<path>.versions/にスナップショットされます。修正後に内容が異常な場合(置換結果がプロンプトと一致しない、ユーザーが特定のラウンドの結果に不満があるなど)、list_versionsでスナップショットを確認し、restore_versionで修正前に戻してから再生成します。ロールバック後、古いアノテーションの loc インデックスが無効になる可能性があるため、read_locationで再確認するか、ユーザーに再アノテーションを依頼してください(最初のラウンドでのロールバックが最も価値があります——アノテーションは元の文書に基づいています)。アノテーションウィンドウのサイドバーの**「履歴」タブ**からもバージョンチェーンを直接表示(各項目に操作説明/時刻/サイズを含む)し、ワンクリックでロールバックできます。ユーザーのロールバック操作はセッションのuser_actionsに記録され、wait_for_annotationsとともに返されます——LLM はuser_actionsに restore が含まれている場合、文書がロールバックされ、以降の修正が無効になる可能性があることを認識し、確認してから続行してください。待機ポーリング(タイムアウトなし):
wait_for_annotationsはデフォルトでタイムアウトしません(ユーザーが送信 / キャンセル / ウィンドウを閉じる、またはエージェントが終了するまで待機します)。一部のクライアントでは単一のツール呼び出しにタイムアウト上限(約60秒)があり、その場合単一の呼び出しは中断されます——そのツールを再度呼び出せば待機を続行できます。セッションは中断によってキャンセルされません。ウィンドウを閉じる: ユーザーがアノテーションウィンドウを閉じた場合、
wait_for_annotationsはclosed+reason(window_closed/page_unload/window_lost/agent_cancelled/agent_exited)を返します。LLM は理由に基づいて、アノテーションを再度開く(再度annotate_document)か、ユーザーに問い合わせるかを決定します。ウィンドウの自動リロード(プッシュモード): 各
apply_edit/template_replace/find_replace/restore_versionの修正が成功すると、サーバーはその文書のアノテーションウィンドウに自動的にリロードをプッシュします(約2秒のデバウンス、連続修正は1回にまとめられます)。ウィンドウは修正後の最新内容を即座に表示します——LLM が手動でannotate_documentを再度呼び出す必要はありません。スクロール位置の保持: リロード後、ページは以前の閲覧位置に戻ります(ビューポートのアンカー——ビューポート上部付近のコンテンツを記録し、オフセットに基づいて復元します。単純に先頭に戻るのではありません)。同時に xlsx の現在のシートと pptx の現在のページも保持します。
複数会話で共有されるサーバーのウィンドウ所有権: WorkBuddy などのクライアントはグローバルに同じ MCP サーバープロセスを共有し、すべての会話のアノテーションセッションが同じプロセスに混在します。修正後の自動リロードが他の会話のウィンドウに影響しないように、LLM は
wait_for_annotationsでアノテーションを取得した後、返されたsession_idをそのままapply_editなどの修正ツールに渡す必要があります(ツールはこのオプション引数をサポートしています)。サーバーはこれに基づいて正確にこのセッションのウィンドウをリロードします。渡されない場合は「最近アクティブなセッション」でフォールバックし、パスに曖昧さがある場合はリロードしません(誤操作を防ぐため)。同一文書のアノテーションウィンドウは共有サーバー下で再利用されます——同じ時刻に複数の会話で同じ文書を操作しないことをお勧めします。同一文書の再利用: 同じ文書に対して再度
annotate_documentを呼び出すと、既存のウィンドウを再利用して自動リロードします(ページは最新ファイルを再取得し、前回のアノテーションをクリアします)。新しいウィンドウは開きません。複数ラウンドのアノテーションループ: エージェントが1ラウンドのアノテーションを処理した後、ウィンドウは自動リロードされ、修正を確認してアノテーションを続行できます。エージェントは再度
wait_for_annotationsを呼び出して次のラウンドを待機する必要があります。ウィンドウで【完了】をクリックし、アノテーションを追加しない場合、そのラウンドで修正不要という意味です。【キャンセル】をクリックするとウィンドウを直接閉じてそのラウンドを終了します(wait_for_annotationsはこれに基づいて停止します)。文字列レベルのアノテーション: Word ページで連続する文字列をドラッグ選択すると、アノテーションはその文字列に正確になります(
loc.range)。apply_editはその部分のみ置換します。キャンセル: ユーザーがページで【キャンセル】をクリックするか、LLM が
cancel_sessionを呼び出すと、セッションはcancelledになり、アノテーションウィンドウは直接閉じます。エージェントプロセスが終了すると、すべてのアノテーションウィンドウが自動的に閉じます。修正は非破壊的です: 初回編集前に自動的に
文書名.bak-<タイムスタンプ>を生成し、各修正前に<path>.versions/バージョンチェーンに自動スナップショットします(list_versions/restore_versionで任意のステップにロールバック可能)。
既知の制限
.docx/.xlsx/.pptx(OOXML)のみ対応。Word: docx-preview のレンダリング順序と body の走査順序が 1:1 であると仮定しています(極端なレイアウト要素ではずれが生じる可能性があります。
web/src/viewers/docxViewer.tsの先頭コメントを参照)。PPT: シェイプの位置特定はテキストマッチングでフォールバックするため、シェイプのテキストが重複している場合は不正確になる可能性があります。
Excel: SheetJS + 軽量グリッド。セルをダブルクリックしてページ内で編集可能(修正項目はアノテーションとして LLM に返され、確定されます)。結合セルの非左上隅は読み取り専用。数式セルを直接編集すると、数式が通常の値に置き換わります。
セッション状態はメモリ内にあり、MCP サーバーを再起動すると失われます。
FAQ
Q: ブラウザが自動的に開かない?
A: アノテーションページは Chrome/Edge の独立したアプリウィンドウで開きます(アドレスバー/ツールバー/タブなし)。Chrome/Edge が見つからないか起動に失敗した場合、ツールは url を返すので、LLM がリンクを送信し、手動で開くことができます(この場合、ウィンドウのクローズ検出はページのハートビートにフォールバックします)。
Q: アノテーションウィンドウが自動的に閉じるのはなぜ? A: Esc でエージェントをキャンセルした場合、またはエージェント/opencode プロセスが終了した場合、サーバーはアノテーションウィンドウを能動的に閉じます。ページで【完了】または【キャンセル】をクリックした後、ウィンドウは読み取り専用のまま再利用を待機します(同一文書の次のラウンドでは直接リロードされます)。
Q: アノテーションページが開けない / 白画面?
A: cd web && npm run build を実行したことを確認してください(サーバーは web/dist がない場合、プレースホルダーの案内ページのみを返します)。文書パスが絶対パスで、ファイルが存在することを確認してください。
Q: なぜプラグインが必要なのか?
A: プラグインはメッセージ層で「文書を精密に修正する」意図を検出し、ガイダンスを注入して、LLM が手を動かす前にまずあなたにアノテーションさせ、位置を推測して直接修正するのを防ぎます。注意: プラグインは opencode でのみ有効です。WorkBuddy などの他の MCP クライアントでは、LLM のフローガイダンスはツールの説明と wait_for_annotations が返す next フィールドから得られます(複数ラウンドのループ規約が組み込まれています)。
Q: 修正完了後にウィンドウが自動リロードされ、再度アノテーションして【完了】をクリックすると、エージェントは処理を続行しますか? A: エージェントがまだ待機しているかどうかによります:
エージェントがガイダンスに従って「アノテーション処理 → 再度
wait_for_annotations」のループにある場合、あなたが送信した新しいアノテーションはすぐに取得され、処理が続行されます(複数ラウンドがシームレスに接続されます)。エージェントがすでにターンを終了している場合(ツールを呼び出していない場合)、あなたのアノテーションはサーバーに一時保存されますが、エージェントは自動的に起動されません——その場合はエージェントに「アノテーションを送信しました。処理を続行してください」と伝えてください。 これは「実行主体は LLM」というアーキテクチャの固有の境界です。サーバーはエージェントに代わって次のラウンドを能動的に引き継ぐことはできません。
Q: 2つの会話が同時に同じ文書を編集すると、ウィンドウが混ざりますか?
A: WorkBuddy はグローバルに同じサーバープロセスを共有するため、同じ文書のアノテーションウィンドウは共有・再利用されます。修正後の自動リロードは現在セッションに正確にバインドされています。LLM が apply_edit などのツールに wait_for_annotations が返した session_id を含めると、その会話のウィンドウのみがリロードされます。含めない場合は最近アクティブなセッションでフォールバックし、パスに複数の生存セッションがある場合はリロードしないことを選択します(誤操作を防ぐため)。同じ時刻に複数の会話で同じ文書を操作しないことをお勧めします。
Available Tools
11 toolsannotate_documentA
打开文档标注页面(H5):创建标注会话并启动本地 HTTP 服务,在独立浏览器应用窗口打开标注页(无地址栏/工具栏/标签页)。若同一文档已有存活窗口则直接复用(原位重载、清空上轮标注),不重新开窗。返回 {session_id, url}。多轮循环:窗口会在文档被修改后自动重载最新内容(apply_edit 成功后服务端主动推送),无需每次手动调用本工具;也可在本轮开始或需要立即刷新时主动调用。同一文档重复调用会复用现有窗口。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 目标文档绝对路径(.docx / .xlsx / .pptx) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it discloses session creation, local HTTP service startup, window reuse/in-place reload, clearing of previous annotations, automatic reload after apply_edit via server push, and the returned {session_id, url}. This goes well beyond what the schema alone communicates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loads the primary action and return value before the multi-turn loop explanation. Minor redundancy exists: the final sentence about reusing existing windows repeats the earlier reuse statement, slightly padding an otherwise efficient description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description states the return value explicitly. It also covers the window lifecycle, the automatic reload mechanism, and the exact conditions under which the agent should call the tool again. For a tool that starts a local service and manages a browser window, nothing needed for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already describes path as an absolute .docx/.xlsx/.pptx path. The description adds meaning by making path the identity used for window reuse ('同一文档'), which clarifies that the same path maps to the same session/window. That is value beyond the schema's type-level description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('打开文档标注页面'), explains the underlying work (create annotation session, start local HTTP service, open an H5 page in a standalone browser window), and distinguishes the tool's key reuse behavior from a simple one-shot opener. It leaves no ambiguity about the resource it operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear when-to-use and when-not-to-use guidance: call at the start of a round or when an immediate refresh is needed, but not after every edit because the window auto-reloads after apply_edit. It does not explicitly name alternatives among the sibling tools, but the behavioral loop guidance is strong enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apply_editA
对指定位置应用修改(首次编辑前自动生成 .bak-<时间戳> 备份;同一标注轮次内只有第一次编辑会快照到版本历史,其余编辑复用该版本——一轮多个标注只产生一个版本,版本描述会汇总本轮全部修改)。mode 缺省 replace;引擎错误返回结构化错误。返回 {ok, loc, new_content, version}(version 为 null 表示本轮已快照过、未新增版本;version.desc 表达该快照为'修改前/该轮标注前的状态'——会话场景为'第 N 轮标注前的状态(本轮修改:…)',无会话为'修改前:<编辑摘要>',回退到该版本即恢复为此内容)。修改成功后,该文档的标注窗口会自动重载展示最新内容(无需再手动调用 annotate_document);用户检查后可能继续标注,此时应再次调用 wait_for_annotations 获取下一轮标注。若修改后发现内容异常,可用 list_versions / restore_version 回退到修改前再重新生成。
| Name | Required | Description | Default |
|---|---|---|---|
| loc | Yes | 统一位置描述符(docs/contracts.md §2):docx-paragraph / docx-cell / xlsx-cell / xlsx-range / pptx-shape / pptx-shape-paragraph | |
| mode | No | 编辑模式,缺省 replace | |
| path | Yes | 文档绝对路径 | |
| style | No | 格式样式提示(可选字段;缺省表示保持原样) | |
| session_id | No | (建议)标注会话 ID(wait_for_annotations 返回)。提供后窗口重载精确绑定该会话,且同一轮内的多次编辑会合并为一个版本快照 | |
| new_content | No | 新内容(mode=delete 时忽略,可为空字符串) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully discloses the mutation behavior: automatic .bak-<timestamp> backup before first edit, version snapshot only on the first edit of a round, later edits reusing that snapshot, and version.desc representing the pre-edit state. It also discloses structured engine errors, the exact return shape, the auto-reload side effect, and a rollback path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core purpose in the first clause and every subsequent sentence adds a distinct operational fact: backup generation, version collapsing, return values, auto-reload behavior, and rollback. It is dense but contains no filler or repeated schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with six parameters, nested objects, and no output schema, the description defines the return tuple, version semantics, backup and rollback path, and follow-up workflow with wait_for_annotations. No critical input needed to call the tool or react to its result is missing; only the exact structured-error shape is summarized rather than enumerated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description adds meaningful semantics beyond the schema: mode defaults to replace, delete ignores new_content, and session_id causes same-round edits to merge into one version snapshot. It also explains the meaning of version and version.desc in the response, which the input schema does not cover.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation: '对指定位置应用修改' (apply modifications at a specified location), immediately clarifying that this is a location-targeted edit tool. It also distinguishes its workflow from annotate_document by stating that the annotation window reloads automatically, and the loc-kind list reinforces that this is structural-position editing rather than a template or search operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong workflow guidance: call this after receiving annotations, do not manually call annotate_document afterward, call wait_for_annotations again for the next round, and use list_versions/restore_version if the edit result is wrong. It does not explicitly compare against template_replace or find_replace, so the selection boundary among sibling edit tools is left somewhat to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_sessionB
取消标注会话(用户点了“取消”)。返回 {session_id, status}。
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | 标注会话 ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions cancellation and the return shape, but does not explain side effects (e.g., whether in-flight requests are aborted, whether the session is permanently deleted or recoverable), authentication requirements, or idempotency. For a mutating action without annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the action and immediately states the return value. There is no wasted text, and it is appropriately sized for a simple cancellation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (one parameter, no output schema, no annotations), the description provides the essential action and return format. It is minimally sufficient for an agent to invoke the tool, but lacks contextual details like side effects or error handling. For a mutation tool, this is adequate but not rich; a 3 reflects the missing behavioral context without over-penalizing given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter 'session_id' is fully documented in the schema. The description adds no extra meaning about the parameter (e.g., format, constraints, or how to obtain it). Per the rubric, baseline 3 applies when schema covers the parameter, and the description does not need to repeat it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb '取消' (cancel) and a clear resource '标注会话' (annotation session). This distinguishes it from siblings like 'annotate_document' and 'wait_for_annotations', though it does not explicitly name an alternative like the highest-quality examples. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It only states the action and return value. It implies usage for canceling a session, but provides no exclusion criteria, prerequisites, or references to related tools like 'wait_for_annotations' or 'preview_edits'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_replaceA
在 docx 文档中全文查找并替换字符串(普通字符串,非正则;匹配范围含表格内单元格)。match_case=false 时大小写不敏感,替换文本原样插入。返回 {matched, replaced, locations}。
| Name | Required | Description | Default |
|---|---|---|---|
| find | Yes | 要查找的原文(普通字符串,非正则) | |
| path | Yes | 目标 .docx 文档绝对路径 | |
| replace | Yes | 替换为的内容 | |
| match_case | No | 是否大小写敏感,默认 false | |
| session_id | No | (建议)标注会话 ID(wait_for_annotations 返回)。窗口重载精确绑定该会话 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses plain-string matching, table-cell coverage, case-insensitive behavior when match_case=false, verbatim insertion of replacement text, and the return shape. It does not explicitly warn that the file is modified in place or discuss reversibility, but the core mutation semantics are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler. Every clause adds information: scope, regex exclusion, table-cell inclusion, case sensitivity, replacement semantics, and return contract. The most important operation and scope come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the action, scope, case behavior, replacement semantics, and return values, which is enough for an agent to invoke it correctly. It lacks only an explicit statement about whether the document is overwritten in place or how to recover, but the core usage context is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value beyond the schema: it clarifies that match_case=false means case-insensitive, that replacement text is inserted as-is, and that the result includes matched/replaced counts and locations. This deepens understanding of the parameters' runtime behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete operation: full-text find-and-replace in .docx files. It explicitly scopes the behavior as plain-string (non-regex) and includes table cells, which clearly distinguishes it from related tools like template_replace or scoped edits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is for whole-document plain-string replacement with an optional case-sensitivity flag. It does not explicitly name alternatives or state when not to use it, but the intended use case is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_versionsA
列出文档的版本历史(.versions/ 快照链:同一标注轮次内多个编辑合并为一个版本,即每轮首个编辑前快照;新版本在前)。返回 [{index, path, size, mtime, kind?, desc?, ts?}],无版本时返回 []。每条 desc 表达'修改前/该轮标注前的状态'(如'第 1 轮标注前的状态'或'修改前:替换…'),回退到该版本即恢复为该内容。配合 restore_version 实现任意步回退。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 文档绝对路径(.docx / .xlsx / .pptx) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it reveals the snapshot-chain model, version merging per annotation round, ordering (new versions first), empty-list behavior, and the meaning of each desc as 'state before that annotation round'. This goes far beyond what the schema alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: purpose, snapshot semantics, output shape, desc meaning, and companion tool. It is front-loaded with the core action and maintains a logical flow from behavior to return format to usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one required parameter, no output schema, and no annotations, this description is complete: it explains the return fields, empty result behavior, version semantics, and how to combine with restore_version. An agent can correctly invoke it and interpret results without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% coverage for the single 'path' parameter, so the baseline is 3. The description does not add parameter-specific semantics beyond the schema, but none are strictly needed here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: '列出文档的版本历史' (list document version history). It further clarifies semantics with the snapshot-chain explanation, output array shape, and explicitly names restore_version as the companion tool, so an agent can distinguish it from siblings like restore_version or read_location.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates this tool is for reading version history and works together with restore_version for arbitrary rollback, giving a concrete workflow context. However, it does not explicitly state when not to use it, e.g., for reading current document content instead of history.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_editsA
批量预览修改(仅内存计算,不写盘):按顺序在内存中对同一文档应用 edits 列表,返回每处的 {loc, before, after}。适合在真正 apply_edit 前检查一批修改是否符合预期。edits 元素为 {loc, new_content?, mode?, style?}。返回 [{loc, before, after}]。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 文档绝对路径 | |
| edits | Yes | 要预览的修改列表(按顺序应用,与逐个 apply_edit 一致) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral disclosure and does a good job: it explicitly states the operation is in-memory only and does not write to disk, and that edits are applied sequentially to the same document. It stops short of describing failure modes or precise before/after semantics for edge cases such as delete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main safety and purpose information is front-loaded in a compact opening phrase, followed by use case, parameter summary, and return shape. The return format is stated twice, which is mildly redundant, but overall the description is appropriately sized and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a preview tool with no output schema, the description covers what it returns, how it applies edits, and the critical no-write behavior, and it points to apply_edit as the actual mutation path. Minor gaps remain around edge-case behavior such as delete-mode results or invalid locations, but nothing essential blocks correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all parameters with 100% coverage, so the baseline is 3. The description adds meaningful semantics beyond the schema by clarifying that the edits list is applied sequentially in memory to the same document, and by summarizing the element shape as {loc, new_content?, mode?, style?}.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: batch preview edits, memory-only, with no write to disk. It explicitly returns per-edit {loc, before, after} and differentiates itself from the sibling apply_edit by positioning itself as a pre-check step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use this tool: before actually applying edits via apply_edit, to verify whether a batch of modifications is as expected. It names apply_edit as the real-write alternative, though it does not enumerate when-not-to-use conditions or other sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_locationA
读取指定位置的原文与相邻上下文,供生成新内容前核对。返回 {loc, text, context}。
| Name | Required | Description | Default |
|---|---|---|---|
| loc | Yes | 统一位置描述符(docs/contracts.md §2):docx-paragraph / docx-cell / xlsx-cell / xlsx-range / pptx-shape / pptx-shape-paragraph | |
| path | Yes | 文档绝对路径 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. The verb '读取' clearly implies a read-only operation, and the explicit return shape {loc, text, context} tells the agent what to expect. It could additionally mention failure behavior or permissions, but the non-destructive nature is evident.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that communicates purpose, usage context, and return shape without any wasted words. Every part contributes to helping an agent decide whether and how to call the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with two well-documented parameters, the description is mostly complete. It compensates for the lack of an output schema by specifying the return shape. It does not elaborate on the six loc kinds, but the schema enum already covers that, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema. The description adds no additional semantic information about the parameters themselves; it only describes the return value. The schema's descriptions of path and loc are sufficient, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (read), the resource (a specified location), and the scope (original text plus adjacent context), making the tool's purpose understandable. It does not explicitly differentiate from sibling tools like read_structure, but the focus on content and context is distinct enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage context: it should be used to check the original text and surrounding context before generating new content. It does not mention alternatives or when not to use this tool, but the stated scenario provides actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_structureA
读取文档大纲:docx 逐段 {loc,text};xlsx 每 sheet 维度与抽样;pptx 每页形状文本。返回 {format, items|sheets|slides}。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 文档绝对路径 | |
| max_items | No | 限制返回条目数,默认 500 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It clearly communicates that this is a read operation and specifies the exact return shape per format: paragraph-level {loc,text} for docx, sheet dimensions/sampling for xlsx, and shape text for pptx. It stops short of describing error behavior or limits beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, covering three file formats and the return envelope in a single sentence. Every clause earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by explicitly stating the return structure '{format, items|sheets|slides}' and per-format content. It is complete enough for an agent to know what to expect. Minor gaps like max_items interaction with sampling are not critical for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'path' and 'max_items' are described in the schema. The tool description adds no parameter-level detail beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('read') and resource ('document outline') and goes beyond that by detailing format-specific behavior for docx, xlsx, and pptx. This clearly distinguishes it from read_location and other sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for inspecting document structure before performing edits or annotations, but it does not explicitly name alternatives or state when not to use it. The format breakdown gives implicit context, but there is no direct when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_versionA
将文档直接回退到 list_versions 给出的某个版本(回到该版本锚点:文档被覆盖为该版本内容,回退前不会快照当前状态,因此版本记录不会增加)。⚠️ 警告:该版本之后的全部修改将丢失,且回退本身不可通过新快照撤销(但既有版本链仍保留,可再回退到其它版本)。回退后旧标注的 loc 索引可能失效,应重新 read_structure 遍历后再继续编辑。返回 {restored_index, path, version_path, versions}。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 文档绝对路径 | |
| index | Yes | 要恢复的版本序号(list_versions 返回的 index) | |
| session_id | No | (建议)标注会话 ID(wait_for_annotations 返回)。窗口重载精确绑定该会话 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility and does so thoroughly: it discloses that no snapshot is taken before rollback, version history will not increase, later modifications are lost, the rollback cannot be undone via a new snapshot, and old annotation loc indices may become invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: the core action is bolded and front-loaded, followed by the critical warning, the post-rollback guidance, and the return shape. Every sentence earns its place, with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive version-restore tool with no annotations and no output schema, the description is remarkably complete: it covers the action, the exact behavior regarding version history, data loss, irreversibility, follow-up steps, and the return payload.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the schema already explains path, index, and session_id. The description adds behavioral context around the index parameter (it comes from list_versions) but does not materially extend the schema's parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('直接回退' / directly restore), a clear resource (document version from list_versions), and the exact result ('document is overwritten with that version's content'). It also names sibling tools list_versions and read_structure, making the tool's role distinct from listing or editing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly identifies list_versions as the source for the version index and instructs the agent to re-run read_structure after rollback before continuing edits. It warns about data loss and irreversibility, but does not explicitly contrast this tool with alternatives like apply_edit or template_replace for non-restore modifications.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
template_replaceA
批量替换 docx 文档中的 {{变量}} 模板占位符。variables 为 {变量名: 值} 映射;未提供值的变量保持原样并在 missing_vars 中列出。同段落多处占位符按从右到左应用保证坐标正确;跨 run 断裂的占位符也能正确替换。返回 {matched, replaced, missing_vars, applied}。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | 目标 .docx 文档绝对路径 | |
| variables | Yes | 模板变量映射,如 {"name": "张三", "date": "2026-08-26"} | |
| session_id | No | (建议)标注会话 ID(wait_for_annotations 返回)。窗口重载精确绑定该会话 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
描述揭示了重要的非显而易见行为:未提供值的变量保持原样并在 missing_vars 中列出;同段落多处占位符按从右到左应用保证坐标正确;跨 run 断裂的占位符也能正确替换。这些细节超出 schema 和 annotations 提供的范围(annotations 未提供,schema 未提及这些行为)。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述紧凑且信息密集,每句话都有明确用途。核心行为在前两个分句中,边界情况(跨 run、从右到左)透明披露,返回结构也说明了。没有冗余。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
描述涵盖了返回结构(matched, replaced, missing_vars, applied)、边界情况行为和参数含义。但没有明确说明失败模式(如文件不存在、权限问题)或是否修改原文件或生成新文件。考虑到描述已经提供大量行为细节,缺少这些不影响基本使用。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema 描述覆盖率 100%,所有参数在 schema 中都有描述。描述本身没有为参数添加额外语义,但 variables 参数的行为(未提供值则保持原样)在描述中提到了,这补充了 schema 描述。整体上 schema 已承担主要负担,描述添加了边际价值。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确说明工具行为:批量替换 docx 文档中的 {{变量}} 模板占位符,并提到返回值。与兄弟工具如 find_replace 和 apply_edit 有区分(模板占位符替换 vs 通用查找替换/编辑)。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述隐含了使用场景(模板填充、缺失变量报告),但没有明确说明何时使用 vs 兄弟工具(如 find_replace)。没有提供排除条件,但通过描述变量映射和缺失变量行为暗示了用途。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_annotationsA
阻塞等待标注会话完成 / 取消 / 窗口关闭,返回用户提交的标注列表。返回 {session_id, status, reason?, annotations, next?},status 为 done | cancelled | closed | timeout。缺省永不超时(timeout_seconds 不传则一直等待直到用户提交/取消/关闭窗口,或 agent 退出);如客户端对单次调用有超时上限,可传 timeout_seconds 设一个安全值,被截断或超时后再次调用本工具继续等待即可(会话在服务端持续存在,不会因单次调用被截断而取消)。status=closed 表示标注窗口已关闭:reason=window_closed/page_unload 是用户主动关闭——本轮标注流程已结束,不要再调用本工具或 annotate_document(除非用户明确要求继续),直接总结结果即可;reason=window_lost(窗口崩溃/被强杀)或 agent_cancelled/agent_exited 时才考虑询问用户是否重开。status=done 且存在标注时:处理这些标注时请在 apply_edit / template_replace / find_replace / restore_version 中带上本返回值中的 session_id(确保标注窗口重载到正确的会话,多对话共享同一 server 时尤为重要);处理完后(窗口会自动重载修改后的内容)应再次调用本工具等待用户下一轮标注。status=done 且标注为空:用户确认本轮无需修改,标注流程已结束(窗口已自动关闭),不要再调用本工具或 annotate_document,直接总结结果即可。
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | 标注会话 ID(annotate_document 返回) | |
| timeout_seconds | No | 本次等待超时秒数,默认 1800(30 分钟);若客户端有工具超时上限请设一个安全值并轮询调用 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and delivers: blocking semantics, server-side session persistence across calls, no cancellation on truncation, status expansion, and automatic window reload behavior. However, it contradicts the input-schema default for timeout_seconds (description says 'never time out' by default; schema says 'default 1800'). This inconsistency creates agent-facing ambiguity about the actual default wait behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The paragraph is long and dense but nearly every clause carries a distinct behavioral rule spanning multiple statuses and call-pattern scenarios, so the length is justified. Core blocking behavior and default timeout are front-loaded, but the wall of text could be better structured into bullets or status-to-action pairs for faster parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param tool with no output schema and no annotations, the description covers return shape, per-status meaning, follow-up actions, and session persistence — essentially everything needed for correct invocation. The only substantive gap is the timeout-default contradiction with the schema, which leaves an agent uncertain about the real server-side default.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds meaningful semantics beyond the schema: when and why to pass timeout_seconds (client call limits), what to do if the call is truncated (call the tool again), and that session_id must be forwarded into apply_edit/template_replace/find_replace/restore_version. The added value is slightly offset by the timeout-default contradiction with the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: blocks waiting for an annotation session to complete and returns the user's submitted annotations. Explicitly lists the return shape and all possible status values (done | cancelled | closed | timeout), which clearly distinguishes it from sibling edit/apply tools like apply_edit or find_replace. No tautology and no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides exhaustive when-to-use and when-not-to-use rules: call again after processing done-annotations, do not call again when status is closed (user-initiated) or done with empty annotations, and only ask the user about reopening for window_lost/agent_cancelled/agent_exited. It also instructs passing session_id into sibling tools and when to set timeout_seconds for client-limited calls. This is textbook-level usage guidance with explicit exclusions and alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Most tools have clearly distinct roles in the read-annotate-edit-version workflow. The main ambiguity is among apply_edit, find_replace, and template_replace, which all modify documents, though their descriptions clarify different use cases.
Tool names are consistently lowercase snake_case and mostly follow a verb_noun pattern like read_location, apply_edit, and list_versions. template_replace and find_replace deviate slightly from the verb-first style, but the naming remains predictable and readable.
11 tools is well within the ideal range, and each tool covers a necessary part of the fine-tuning workflow: reading, annotating, waiting, editing, previewing, and versioning. There are no obvious filler or redundant tools.
The toolset covers the full annotation loop: open a session, wait for annotations, apply edits, preview changes, and roll back via versions. Features like document creation or export are outside the stated fine-tuning purpose, so no significant gaps are apparent.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
An agent-first office suite Claude & ChatGPT read and write over one MCP URL.
E2LLM gives your AI eyes and hands in a real browser: structured perception (SiFR) plus action.
Versioned artifact review for people and AI agents, with contextual comments and human control.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to read, write, and edit Office documents via LibreOffice with token-efficient design. Supports multiple formats including DOCX, XLSX, PPTX, and legacy formats through LibreOffice bridge.2735MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to directly read, edit, and manipulate Word documents, supporting image and table operations, paragraph editing, and search/replace.202MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to create, edit, format, and analyze Microsoft Office documents (Excel, Word, PPT, PDF) using natural language, with features like table replication and automated styling.2MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to read, create, edit Word/Excel/PPT files and manage the filesystem on the user's computer via natural language.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/chaosst/doc-fine-tuning-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server