Skip to main content
Glama

Abyss MCP

AI 앱과 터미널에서 Abyss 프로젝트의 맥락을 이어서 사용할 수 있는 공식 연결 도구입니다.

공식 제품 식별

  • 정식 브랜드명은 Abyss이며 heyabyss.com은 공식 도메인입니다. 제품명은 항상 Abyss로 표기합니다.

  • Abyss MCP는 Claude Desktop, ChatGPT(Codex) Desktop 앱, Codex CLI와 Claude Code에서 사용자가 선택한 Abyss 프로젝트의 이전 결정, 근거와 다음 할 일을 이어서 사용할 수 있게 돕는 공식 MCP 서버입니다.

  • 이 저장소는 my-abyss-project/abyss-mcp의 공식 소스입니다. 동명의 telagod/abyss 코드 그래프·토큰 압축 프로젝트와 관련이 없습니다.

  • Abyss는 iPhone·Mac용 음성 비서가 아닙니다. Abyss MCP는 사용자 기기의 임의 명령을 실행하는 범용 원격 제어 도구가 아니며, AWS Bedrock을 MCP 패키지의 기능으로 제공하지 않습니다.

  • 신뢰·보안 안내: https://heyabyss.com/trust

  • AI가 확인할 공식 정보: 서비스 확인 · 간단 안내 · 전체 안내

  • npm 패키지: abyss-mcp

  • 공개 소스: my-abyss-project/abyss-mcp

  • 보안 제보: Security Policy

  • 문의: contact@heyabyss.com

설치 전에는 패키지 이름과 버전, 공개 소스, 신뢰·보안 안내, heyabyss.com 로그인 주소를 확인하세요. /connect/mcp는 로그인을 시작한 뒤 사용하는 연결 승인 경로이며, 공개 문서나 소스 확인 경로가 아닙니다. npm view abyss-mcp@0.5.10 version repository homepage dist.integrity로 배포 정보를 확인할 수 있습니다.

Related MCP server: PostIdentity MCP Server

어디에서 시작하나요?

현재 다음 환경에서 사용할 수 있습니다.

  • Claude Desktop 앱

  • ChatGPT(Codex) Desktop 앱

  • 터미널의 Codex CLI 또는 Claude Code

브라우저에서 사용하는 일반 ChatGPT와의 직접 연결은 2026년 9월 예정입니다. 지금은 위 환경 중 하나에서 시작하고, 웹 브라우저는 Abyss 가입·로그인과 연결 승인에만 사용하세요.

시작하기 전에

  • Node.js 20 이상이 필요합니다.

  • 설치할 패키지 이름은 abyss-mcp입니다.

  • 로그인 페이지가 heyabyss.com인지 확인하세요.

1분 연결

사용 중인 환경 하나를 골라 아래 요청문을 지정된 입력창에 그대로 붙여 넣으세요.

Claude Desktop 앱

로컬 도구 사용 권한이 있는 Claude Desktop 대화 입력창에 붙여 넣으세요.

공식 npm 패키지 abyss-mcp@0.5.10을 설치하고,
이 Claude Desktop에서 Abyss를 사용할 수 있게 설정해줘.
실행할 명령과 바꿀 설정을 먼저 보여주고 내 확인을 받아.
설정 후 앱에서 Abyss 연결을 다시 불러오고 로그인을 시작해줘.
공식 사이트는 https://heyabyss.com 인지 확인하고,
비밀번호나 토큰을 요청하거나 출력하지 마.

ChatGPT(Codex) Desktop 앱

ChatGPT(Codex) Desktop 앱에서 새 작업을 열고 입력창에 붙여 넣으세요.

공식 npm 패키지 abyss-mcp@0.5.10을 설치하고,
이 Codex에서 Abyss를 사용할 수 있게 설정해줘.
실행할 명령과 바꿀 설정을 먼저 보여주고 내 확인을 받아.
설정 후 Abyss 연결을 다시 불러오고 로그인을 시작해줘.
공식 사이트는 https://heyabyss.com 인지 확인하고,
비밀번호나 토큰을 요청하거나 출력하지 마.

터미널의 Codex CLI 또는 Claude Code

Codex CLI나 Claude Code를 실행한 터미널의 프롬프트에 붙여 넣으세요.

공식 npm 패키지 abyss-mcp@0.5.10을 설치하고,
지금 사용하는 Codex CLI 또는 Claude Code에서 Abyss를 쓸 수 있게 설정해줘.
실행할 명령과 바꿀 설정을 먼저 보여주고 내 확인을 받아.
설정 후 Abyss 연결을 다시 불러오고 로그인을 시작해줘.
공식 사이트는 https://heyabyss.com 인지 확인하고,
비밀번호나 토큰을 요청하거나 출력하지 마.

세 환경 모두 AI가 보여주는 명령과 설정 변경 범위를 확인한 뒤 허용하세요. 설치 완료 답변만 확인하고 끝내지 말고, 이어서 아래 로그인과 프로젝트 확인까지 진행하세요.

로그인과 연결 완료

  1. AI가 Abyss 로그인을 시작하면 열린 https://heyabyss.com 페이지로 이동합니다.

  2. 가입 또는 로그인합니다.

  3. 표시된 연결 요청을 확인하고 승인합니다.

  4. 사용하던 AI로 돌아와 “승인했어. 로그인을 완료해줘”라고 요청합니다.

  5. 이어서 “Abyss 연결 상태와 사용할 수 있는 프로젝트를 확인해줘”라고 요청합니다.

다음 두 가지가 확인되면 연결이 끝난 것입니다.

  • AI가 Abyss에 로그인되었다고 응답합니다.

  • AI가 내가 사용할 수 있는 프로젝트 목록을 불러옵니다.

브라우저에 승인 완료 화면만 보이는 상태는 아직 끝이 아닐 수 있습니다. 반드시 사용하던 AI로 돌아와 로그인 완료와 프로젝트 확인까지 진행하세요.

문제가 생겼나요?

1. 패키지를 찾지 못하거나 설치가 멈춤

404, EAI_AGAIN, DNS 또는 시간 초과 메시지가 보이면 AI 실행 환경이 npm에 접속하지 못한 경우가 많습니다. 내 컴퓨터의 터미널에서 다음 명령을 실행하세요.

npx -y abyss-mcp@0.5.10 --version

여기서는 버전이 나오는데 AI에서만 실패하면, 해당 앱의 네트워크 또는 명령 실행 권한을 확인한 뒤 앱을 완전히 종료하고 다시 실행하세요.

2. AI가 Abyss 로그인 기능을 찾지 못함

설치 요청이 끝까지 실행되었는지 확인하고 앱을 완전히 종료한 뒤 다시 실행하세요. 계속 찾지 못하면 아래 수동 설정을 사용하거나 터미널에서 로그인을 시작하세요.

npx -y abyss-mcp@0.5.10 login

3. 브라우저가 열리지 않음

AI 또는 터미널에 표시된 https://heyabyss.com/connect/mcp... 주소를 복사해 직접 여세요. 다른 도메인이 표시되면 로그인하지 말고 중단하세요.

4. 승인했지만 연결되지 않음

사용하던 AI로 돌아와 “로그인을 완료하고 연결 상태를 확인해줘”라고 요청하세요. 터미널로 로그인했다면 다음 명령으로 완료 상태를 확인할 수 있습니다.

npx -y abyss-mcp@0.5.10 complete-login
npx -y abyss-mcp@0.5.10 status

승인 시간이 만료되었다면 login부터 다시 시작하세요. 일시적인 서비스 또는 네트워크 오류라면 로그인 파일을 지우지 말고 잠시 후 status를 다시 실행하세요.

수동 설정과 추가 진단

Codex CLI:

codex mcp add abyss -- npx -y abyss-mcp@0.5.10
codex mcp get abyss

Claude Code:

claude mcp add --transport stdio abyss -- npx -y abyss-mcp@0.5.10
claude mcp get abyss

Claude Desktop 설정:

{
  "mcpServers": {
    "abyss": {
      "command": "npx",
      "args": ["-y", "abyss-mcp@0.5.10"]
    }
  }
}

설정을 저장한 뒤 사용 중인 앱을 완전히 종료하고 다시 실행하세요.

비밀 값을 출력하지 않는 진단 명령입니다.

npx -y abyss-mcp@0.5.10 doctor

지원팀에 문의할 때는 사용한 AI 앱 또는 터미널, 운영체제, 화면에 보이는 오류 메시지를 함께 알려주세요. 비밀번호, 토큰, 로그인 파일 내용은 보내지 마세요.

보안

  • 브라우저에서 Abyss 비밀번호를 입력하며 npm 패키지에 비밀번호를 전달하지 않습니다.

  • 로그인 주소가 heyabyss.com인지 확인하세요.

  • 로그인 파일이나 토큰을 AI 대화 또는 지원 문의에 붙여 넣지 마세요.

  • 프로젝트 접근 권한은 연결된 Abyss 계정을 기준으로 확인합니다.

  • 로그아웃하면 이 기기의 Abyss 연결을 해제할 수 있습니다.

질문이나 연결 문제가 계속되면 contact@heyabyss.com으로 사용한 AI 앱 또는 터미널, 운영체제, 화면에 보이는 오류를 보내주세요. 비밀번호, 토큰, 로그인 파일은 보내지 마세요.

Available Tools

55 tools
ask_if_this_should_be_rememberedAsk If This Should Be RememberedA
Read-onlyIdempotent

Use when: 지금 대화에서 남길 만한 판단을 발견했을 때, 저장 여부를 사용자에게 먼저 묻는 도구. 자동 저장은 절대 하지 않으며, 사용자가 명시적으로 동의한 경우에도 checkpoint 초안으로만 넘긴다. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNo선택 항목. 이 판단이 속한 프로젝트나 주제.
summaryYes방금 대화에서 발견한, 남길 만한 판단과 그 이유의 짧은 요약.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses key behaviors beyond annotations: it is a read or staging action, never auto-saves, requires no human approval for this step, and only passes to checkpoint draft. Also explains failure modes and follow-up actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with front-loaded usage conditions. Some redundancy (e.g., repeating 'no auto-save' in multiple forms) but overall concise for the amount of guidance provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description covers all necessary context: when to use, prerequisites, safety profile, failure handling, and post-invocation steps. Missing nothing essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so parameter meanings are already clear. The tool description adds no additional semantic detail beyond the schema, meeting the baseline expectation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to ask the user if a judgment from the current conversation should be remembered, and it never auto-saves—only creates a checkpoint draft. It distinguishes itself from sibling tools like commit_checkpoint by emphasizing the ask-first behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' conditions, including requirements (authenticated API authority) and failure handling instructions. This gives clear guidance on when to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

auth_statusAuth StatusA
Read-onlyIdempotent

Use when: Check the current Abyss MCP login status. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds value by specifying 'Effect: read,' human approval not required, and detailed failure handling (e.g., 'On failure: login_required → login; project_not_selected → list_projects/select_project; ...'). This contextualizes behavior beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Use when, Do not use when, Requires, Effect, etc.). It is front-loaded with the key purpose. However, it is somewhat verbose with failure handling details, which could be condensed without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, full annotation coverage, and the presence of an output schema, the description is complete. It covers purpose, prerequisites, failure modes, and procedural next steps, leaving no significant gaps for correct tool usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and schema coverage is 100%. Per guidelines, baseline is 4. The description does not need to add parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb and resource: 'Check the current Abyss MCP login status.' It distinguishes itself from siblings like login, logout, and whoami by focusing on status checking.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides use and misuse conditions: 'Use when: Check the current Abyss MCP login status.' and 'Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action.' Also states the requirement for authenticated API authority.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkpointCheckpointA
Idempotent

Use when: Create a reviewable V1 checkpoint draft from the currently visible conversation. Structure observation bounds, safe source spans, candidate changes, per-change content approval, and unresolved items. Do not claim host metadata you cannot see; record unavailableFields. Do not include the full conversation, credentials, private local paths, or raw identity fields. The server computes the draft digest. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: draft. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required at the later commit/apply boundary. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesCheckpoint title.
changesYesOne or more item-level changes to stage in a draft checkpoint.
summaryYesHuman-readable checkpoint summary.
metadataNo
sessionIdNoOptional session id. Defaults to the selected project session.
evidenceIdsNo
sourceSpansYes
schemaVersionYes
idempotencyKeyYes
unresolvedItemsYes
conversationProvenanceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate idempotentHint=true and non-destructive write. The description adds key behavior: 'Effect: draft', server-computed digest, idempotencyKey replay semantics ('never bypasses authorization'), and human approval at a later stage. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear section headers (Use when, Do not use when, Requires, Effect, Then, On failure) and front-loads the critical purpose. It is somewhat long, but every sentence adds necessary context; the organization earns a high score for clarity despite length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high complexity (11 parameters, rich nested objects, output schema present), the description covers purpose, usage conditions, behavioral traits, error handling, and post-invocation steps. It omits detailed parameter semantics but includes key constraints (no credentials, no full conversation). The presence of an output schema partially compensates, but low schema coverage limits completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at only 36%, the description carries responsibility for explaining parameters. Instead, it provides only high-level terms ('observation bounds', 'safe source spans', 'candidate changes') without clarifying individual fields like metadata, sessionId, evidenceIds, or nested structures like conversationProvenance. The schema's own descriptions are sparse, and the tool label does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a 'V1 checkpoint draft' from the visible conversation, using specific verbs and resources. It provides scope ('observation bounds, safe source spans...') and contrasts with 'narrower tool' usage, though sibling tools like commit_checkpoint are not named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly includes 'Use when' and 'Do not use when' sections with concrete conditions, plus 'Requires', 'Human approval', 'Then' guidance, and a full 'On failure' error-handling table. This leaves no ambiguity about when and how to invoke the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

commit_checkpointCommit CheckpointA
Idempotent

Use when: Human review tool: commit the exact reviewed draft after explicit user approval of the checkpoint commit. Tool-execution approval or content approval is not commit approval. Bind approvalEvidence to the server draft digest returned by checkpoint or review_pending. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesShort reason describing what the user approved.
confirmedNoMust be true after the user explicitly approves committing this checkpoint.
checkpointIdYesCheckpoint draft id returned by checkpoint or review_pending.
idempotencyKeyYes
contractVersionYes
approvalEvidenceYes
expectedCheckpointDigestYesServer-issued draftDigest from checkpoint or review_pending.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=false, idempotentHint=true, destructiveHint=false. The description adds crucial context: 'canonical mutation', idempotency behavior ('reuse the same idempotencyKey', 'replay returns the first canonical result ID and version'), and human approval requirement. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear 'Use when', 'Do not use when', 'Requires', 'Effect', 'Human approval', 'Then', 'On failure' sections. It is front-loaded and every sentence adds value, though slightly lengthy for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool complexity (7 params, nested objects, output schema exists), the description covers approval flow, idempotency, failure modes, prerequisites, and follow-up actions. It is comprehensive without needing to detail return values since output schema is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 57% schema coverage, the description adds value by explaining idempotencyKey reuse policy and binding approvalEvidence to server draft digest. While some parameters still lack detailed explanation, the description compensates well for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the purpose: 'commit the exact reviewed draft after explicit user approval of the checkpoint commit.' It distinguishes itself from siblings like 'checkpoint' and 'discard_checkpoint' by focusing on the commit action after human review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action.' It also provides guidance on idempotencyKey reuse and human approval requirements, giving clear when-to-use and when-not-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_wiki_versionsCompare Wiki VersionsA
Read-onlyIdempotent

Use when: Compare two immutable PageVersions as a stable typed block diff. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
toVersionYes
fromVersionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already show readOnlyHint=true and destructiveHint=false, and the description adds 'Effect: read' which aligns. It goes further by detailing failure modes and required authority, providing valuable context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with bullet points and front-loads the purpose, but it is somewhat verbose, especially in the failure handling section. It could be slightly more concise while retaining essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to explain return values. It covers purpose, usage, prerequisites, failure modes, and post-action steps, making it complete for an AI agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (one param described). The description does not explicitly detail each parameter beyond their names, but it clarifies the role of projectId (optional, defaults to selected project) and reinforces the use of fromVersion/toVersion as version numbers via the tool's purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Compare two immutable PageVersions as a stable typed block diff,' specifying the action, resources, and output. It distinguishes from sibling tools by highlighting the typed diff nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' conditions are provided, including when to avoid using the tool (narrower tool matches, unresolved project scope, user declined). It also lists prerequisites like authenticated API authority and selected project.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complete_loginComplete LoginA

Use when: Complete Abyss browser device login after the user approves the device in the browser. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description goes beyond annotations by noting this is a canonical mutation with no idempotency key (do not auto-retry), human approval not required, and detailed failure paths. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with clear sections and no redundancy, though slightly longer than strictly necessary; still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description covers purpose, usage, effects, and failure handling completely. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters, so schema coverage is 100%. Baseline for 0 params is 4; description does not need to add parameter info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool completes Abyss browser device login after user approval. It distinguishes itself from siblings like 'login' and 'auth_status' by specifying the exact step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' conditions are provided, along with alternatives and failure-handling instructions, giving comprehensive guidance for when to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_context_deliveryConfirm Context DeliveryA
Idempotent

Use when: Record authenticated transport acknowledgement and advance the cursor only for a valid delivered package. Handles gap, stale delivery, and replay explicitly. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes
scopeIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeYes
toVersionYes
failureCodeNo
fromVersionYes
connectionIdYes
clientEventIdYes
idempotencyKeyYes
contextPackageIdYes
expectedCursorVersionYes
transportAcknowledgementHashNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate idempotentHint true and destructiveHint false. The description adds significant context: it is a 'canonical mutation', explains idempotency key reuse ('reuse the same idempotencyKey'), replay behavior ('returns the first canonical result ID and version'), and safety checks. It also states human approval is not required and provides error handling logic for various failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with labeled sections (Use when, Do not use when, Requires, Effect, Human approval, Then, On failure). Every sentence adds value, though the length is substantial. It could be slightly more concise, but it is efficient for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (13 parameters, 2 enums, output schema exists), the description covers purpose, usage, behavioral effects, idempotency, error handling, and post-conditions. Parameter semantics are the main gap. The presence of an output schema reduces the need to explain return values. Overall, it is quite complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 8% (only projectId documented). The description mentions key concepts like idempotencyKey, expectedCursorVersion, fromVersion, toVersion, and status, but does not provide detailed semantics for each of the 13 parameters. For instance, transportAcknowledgementHash is not explained. The description partially compensates by contextualizing versioning and idempotency, but lacks per-parameter clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Record authenticated transport acknowledgement and advance the cursor only for a valid delivered package.' It uses a specific verb ('record') and resource ('transport acknowledgement'), and distinguishes itself from siblings like record_context_consumption by focusing on delivery confirmation and cursor advancement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides 'Use when' and 'Do not use when' conditions, including alternatives like 'a narrower tool better matches the intent.' It also lists prerequisites: 'authenticated API authority and an API-authorized selected Project.' Error handling instructions further clarify when to use which recovery action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

continue_previous_threadContinue Previous ThreadA
Read-onlyIdempotent

Use when: Read-only compatibility alias for recall intent current_thread. Prefer recall for new clients. Deprecated compatibility alias. Prefer resume or recall; do not select this alias for new workflows. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
topKNoOptional bounded recall count, 1-50.
topicNoOptional project, decision, or topic.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds context beyond annotations: states read effect, authentication and project requirements, human approval not needed, and detailed failure handling. Annotations already cover readOnly and destructive hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Structured with clear headings but somewhat lengthy; each section is justified for the deprecated nature. No redundant content despite length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Comprehensive coverage: usage context, requirements, effect, human approval, post-action steps, and failure modes. Output schema exists but is not needed due to detailed 'Then' section.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and descriptions already cover behavior (defaults, rejection). The tool description does not add new parameter information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly identifies the tool as a deprecated compatibility alias for 'recall intent current_thread', distinguishing it from siblings like 'recall' and 'resume'. It explicitly states not to use for new workflows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' conditions, including alternatives and specific scenarios to avoid (e.g., narrower tool match, unresolved project scope).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate ProjectA

Use when: Human-account tool: create a personal project or a project inside an existing team workspace after the user asks to create it. For team scope, first use team_onboarding_status to obtain teamWorkspaceId. The API enforces user identity, team membership, and createProject permission. After creation, call select_project with the returned projectId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that it is a canonical mutation without client-side idempotency, warns against automatic retry, requires human approval, and details failure handling paths. Annotations already indicate non-idempotent and non-read-only; description adds critical context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Structured with clear sections (Use when, Do not use when, etc.), front-loading critical info. Slightly long but justified by complexity; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all aspects: two scopes, idempotency, human approval, failure modes, follow-up actions, and error handling. Output schema exists for return values. No gaps for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear required fields. Description adds context: requiring teamWorkspaceId for team scope and how to obtain it, plus enforcement of permissions. This adds value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool creates a personal or team project, using specific verbs and resources. It distinguishes from siblings like 'select_project' (post-creation) and 'list_projects' (no creation).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide comprehensive guidance, including prerequisites (e.g., get teamWorkspaceId via team_onboarding_status) and follow-up (call select_project). It also advises against use when scope is unresolved or user declined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_wiki_pageCreate Wiki PageA
Idempotent

Use when: Human-account tool: create a child Page under the selected Project Brief or another Page after the user asks to create it. The API enforces selected project scope, editWiki permission, hierarchy limits, and idempotency. Project agent credentials must use propose_wiki_changes and a human review instead. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
documentNoOptional PageBlockSchema v1 document. Omit for an empty Page.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
parentPageIdNoOptional parent Page id. Omit to create directly under the Project Brief.
idempotencyKeyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses it's a canonical mutation, idempotency behavior, human approval requirement, and error handling for various failures. Adds value beyond annotations which already indicate idempotent and non-readOnly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections; front-loaded with 'Use when'. Slightly verbose in places but every sentence provides value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage, behavioral traits, parameter semantics, error handling, follow-up steps. Output schema handles return values, making description complete for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers 60% of params. Description adds important context for idempotencyKey (reuse vs new key) and default for projectId. Does not detail title or other params beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it creates a child Page under the selected Project Brief or another Page, distinguishing it from propose_wiki_changes and other siblings. Verb and resource are clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Includes explicit 'Use when:' and 'Do not use when:' sections with conditions like narrower tool match, unresolved project scope, user declined. Also states alternative for agent credentials.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

discard_checkpointDiscard CheckpointA
DestructiveIdempotent

Use when: Human review tool: discard a specific draft checkpoint after the user explicitly confirms it should not become memory. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmedYesMust be true after the user explicitly approves discarding this checkpoint.
checkpointIdYesCheckpoint draft id returned by checkpoint or review_pending.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds rich behavioral details beyond annotations: canonical mutation, idempotency key behavior, human approval requirement, and failure mode actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well structured with front-loaded purpose, but slightly lengthy due to comprehensive failure handling. Still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage, effect, idempotency, human approval, next steps, and failure recovery. Output schema exists, so no gap on return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and description adds no parameter details beyond what schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool discards a draft checkpoint after user confirmation, distinguishing it from siblings like commit_checkpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use and when not, referring to narrower tools and user intent. Also covers failure handling and preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_block_commandExecute Block CommandA
Idempotent

Use when: Use the API compound command to atomically save PageVersion, object revision, references, ChangeSet and outbox. Agent-created objects remain proposals. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
viewNo
blockNo
objectNo
pageIdYes
positionNo
targetIdNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
selectionNo
targetTypeNo
anchorBlockIdNo
clientDraftIdNo
idempotencyKeyYes
expectedPageVersionYes
expectedObjectRevisionNo
compensatesIdempotencyKeyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses key behaviors beyond annotations: it's a 'canonical mutation', explains idempotency key usage ('reuse the same idempotencyKey... replay returns the first canonical result'), requires human approval, and details error handling (login_required, stale_version, etc.). This adds significant context to the idempotentHint annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with clear sections (Use when, Do not use when, Requires, Effect, Human approval, Then, On failure). It front-loads the most critical usage guidance. Minor redundancy could be trimmed, but overall it is efficient for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (16 parameters, nested objects, output schema exists), the description covers usage constraints, error handling, idempotency, and follow-up steps. However, it lacks detailed parameter guidance for the 'kind' enum and other properties, which would be needed for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 6%, so description should compensate. However, it does not explain individual parameters beyond noting required ones (pageId, kind, expectedPageVersion, idempotencyKey) in context. The 'kind' enum is not elaborated, and many parameters lack description. Some value is added via the compound command context, but it's insufficient for 16 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Use the API compound command to atomically save PageVersion, object revision, references, ChangeSet and outbox.' It uses a specific verb ('execute block command' implied) and resource, and distinguishes from siblings by framing it as a compound command for atomic updates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections are provided, including conditions like 'a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action.' Prerequisites ('Requires') and alternative guidance are also included.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_why_linkedExplain Why LinkedA
Read-onlyIdempotent

Use when: Read-only compatibility tool that explains why a saved thought or connection exists. Deprecated compatibility alias. Prefer get_continuity_object or get_page_backlinks; do not select this alias for new workflows. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetIdNoOptional item id to explain.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds context: 'Effect: read', 'Human approval: not required', and detailed error handling. It also notes the tool is deprecated, providing additional behavioral insight beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Use when, Do not use when, Requires, etc.) and front-loads essential information. It is slightly verbose, but every sentence adds value for a deprecated tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to explain return values. It covers purpose, usage, error handling, and required auth completely. The deprecated status and alternatives are clearly communicated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters adequately. The description does not add meaningful meaning beyond what the schema provides, achieving the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a read-only compatibility tool that explains why a saved thought or connection exists. It distinguishes itself from siblings by explicitly naming preferred alternatives (get_continuity_object, get_page_backlinks) and noting it is deprecated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' and 'Do not use when' conditions, including when a narrower tool matches intent, project scope is unresolved, or user declined action. It also suggests alternative tools, giving clear guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_wikiExport WikiA
Read-onlyIdempotent

Use when: Export the authorized Project Wiki, a Page, or a subtree as Markdown, or the full hierarchy/version/reference manifest as structured JSON. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
pageIdNo
subtreeNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description adds 'Effect: read', 'Human approval not required', and detailed failure recovery steps (login_required, project_not_selected, etc.), providing richer behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with labeled sections (Use when, Do not use when, Requires, Effect, etc.) and front-loaded with the core purpose. Every sentence is necessary and contributes to clarity without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (4 parameters, existence of output schema, many sibling tools), the description covers purpose, usage conditions, prerequisites, effects, failure modes, and follow-up actions. The presence of output schema means return values are handled externally, so the description is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 25% schema description coverage (only projectId described in schema), the description compensates by explaining format (Markdown vs JSON) and the meaning of pageId and subtree through the main use line. This adds value beyond the schema, though not all parameters are explicitly detailed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Export' and the resources (Project Wiki, Page, subtree) and formats (Markdown, JSON). It distinguishes from sibling tools like get_wiki_page or list_wiki_pages by specifying the export functionality and the option to export hierarchy or manifest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'Use when' and 'Do not use when' conditions, including when a narrower tool is better. It also lists prerequisites (authenticated authority, selected project) and failure handling, guiding the agent effectively.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_block_candidatesGet Block CandidatesA
Read-onlyIdempotent

Use when: List permission-safe Project block candidates without raw source or conversation bodies. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNo
typeNo
limitNo
cursorNo
statusNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
currentPageIdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint, idempotentHint, destructiveHint), description states 'Effect: read' and 'Human approval: not required.' Provides detailed error handling mapping. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with bullet points and labeled sections, making it easy to scan. However, it is somewhat verbose; could be more concise without losing critical info.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, no required ones, and an output schema exists, the description covers usage, prerequisites, effects, and error handling. Parameter documentation is lacking, but overall completeness is good.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (only projectId documented). Description does not add meaning for the other 6 parameters (q, type, limit, cursor, status, currentPageId). Fails to compensate for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action: 'List permission-safe Project block candidates without raw source or conversation bodies.' It distinguishes the tool by specifying what is excluded and the context (permission-safe). This is specific and differentiates from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' conditions, including narrowing to intent, project scope, and user consent. Also lists prerequisites (authenticated API authority, selected project) and next steps. Comprehensive guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_context_packageGet Context PackageA
Read-onlyIdempotent

Use when: Read one project-scoped context package by id. The selected project must match any explicit projectId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
contextPackageIdYesContext package id returned by resume, handoff, or list_context_packages.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint, idempotentHint), the description adds effect 'read,' human approval note, and detailed failure handling (e.g., login_required -> login), enhancing transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with labeled sections, front-loaded key action, and every sentence adds value. Efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given output schema exists and annotations cover safety, the description is complete: covers purpose, guidelines, behavior, and failure cases. No missing crucial context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, baseline is 3. Description adds minor context (e.g., source of contextPackageId) but doesn't significantly extend schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Read one project-scoped context package by id.' It specifies the action (read), resource (context package), and scope, distinguishing it from sibling tools like list_context_packages or reuse_context_package.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Has explicit 'Use when' and 'Do not use when' sections with conditions and alternatives. Provides requirements and notes on human approval, offering comprehensive guidance for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_continuity_objectGet Continuity ObjectA
Read-onlyIdempotent

Use when: Read a built-in Decision, Principle, Assumption, Open Question, or Next Action identity plus immutable revision history in the selected project. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, idempotentHint, and non-destructive. Description adds 'Effect: read' and 'Human approval: not required for this read or staging action', plus detailed on-failure behaviors that clarify expected side effects and next steps. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-structured with sections for use cases, requirements, effects, approval, and error handling. Each sentence adds value, though some redundancy exists (e.g., repeating 'read' in multiple forms). Overall efficient for the information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description focuses on preconditions, effects, and error handling, which are fully covered. It addresses authentication, project selection, human approval, and failure scenarios, making it complete for a read tool without needing to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% (only projectId has a description). itemId lacks description, and the tool description does not explain what itemId refers to or how to obtain it. While the context is clear from the overall description, parameter semantics are not fully detailed beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'read' and the specific resources: 'built-in Decision, Principle, Assumption, Open Question, or Next Action identity plus immutable revision history in the selected project.' This specificity distinguishes it from sibling tools like get_wiki_page or get_continuity_operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'Use when:' and 'Do not use when:' conditions, including guidance to look for a narrower tool if applicable. Also details when human approval is not required and provides structured error handling steps with recommended follow-ups.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_continuity_operationsGet Continuity OperationsA
Read-onlyIdempotent

Use when: Read the selected Project durable outcomes, corrections, reviewer attribution, and audit activity. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds 'Effect: read' and 'Human approval: not required for this read or staging action,' plus detailed failure resolution steps. Annotations already indicate read-only, idempotent, non-destructive; description enriches context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections. While somewhat verbose, each section adds value and aids comprehension. Not overly concise but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given low complexity (1 optional parameter, output schema exists), the description fully covers purpose, usage, requirements, and failure modes. No gaps identified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single optional parameter with 100% schema description coverage. Description does not add extra meaning beyond schema, which is already thorough. Baseline score of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool reads selected Project durable outcomes, corrections, reviewer attribution, and audit activity. Distinguishes from siblings by advising not to use when a narrower tool matches intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when:' and 'Do not use when:' sections specify conditions. Also outlines requirements (authenticated API authority, API-authorized selected Project) and failure handling, providing complete guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_sync_activityGet Project Sync ActivityA
Read-onlyIdempotent

Use when: Read connection-scoped delivery, receipt, consumption, preview, and deferral history without changing the sync cursor. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeYes
connectionIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. Description adds beyond annotations by stating the effect is read-only (no sync cursor change) and provides extensive error-specific guidance, enhancing transparency without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with clear sections ('Use when', 'Do not use when', 'Requires', etc.) and front-loads the core purpose. However, it is verbose with extensive detail on failure modes and post-actions, which could be condensed to improve conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking explicit return value documentation (though output schema exists), the description adequately sets expectations by listing the types of history returned (delivery, receipt, etc.) and covering error handling. It is largely complete for a read operation with good annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only projectId has a description). The description does not detail what each parameter means, only using terms like 'connection-scoped' without clarifying scopeType or scopeId semantics. It fails to compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states a specific verb ('Read') and a well-defined resource ('connection-scoped delivery, receipt, consumption, preview, and deferral history'), and distinguishes itself from sibling tools like get_project_sync_status and record_project_sync_decision by emphasizing it does not change the sync cursor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'Use when' and 'Do not use when' conditions, lists prerequisites (authenticated API authority, selected project), and includes detailed failure handling instructions, offering comprehensive guidance on when and how to use this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_project_sync_statusGet Project Sync StatusA
Read-onlyIdempotent

Use when: Read the selected Project contextVersion, monotonic connection/session cursor, pending proposal count, and sync state. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeYes
connectionIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds that the effect is 'read' and human approval is not required. It also details failure scenarios (login_required, project_not_selected, etc.) and appropriate responses, exceeding what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, but it is lengthy and includes many conditional statements (e.g., 'On failure: ...') that could be condensed. Some repetition (e.g., 'Requires:') adds clutter. Could be more concise without sacrificing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 parameters, enums, output schema), the description covers read behavior, failure modes, and post-action steps. However, it lacks parameter-level detail and does not describe the output schema. Output schema exists but is not referenced.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, but the description does not elaborate on individual parameters. It mentions 'connection/session cursor' hinting at connectionId and scopeType, but does not explain values or constraints beyond the schema. More parameter guidance is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb 'Read' and specific resources: 'selected Project contextVersion, monotonic connection/session cursor, pending proposal count, and sync state.' It distinguishes this tool from siblings like get_project_sync_activity by focusing on the current sync status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'Use when' and 'Do not use when' with concrete conditions (e.g., 'narrower tool better matches the intent', 'project scope is unresolved'). It also provides follow-up guidance: 'review pending proposals/drafts before any canonical apply.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_retention_policyGet Retention PolicyA
Read-onlyIdempotent

Use when: Human-account tool: read a team workspace retention policy. Use team_onboarding_status to find teamWorkspaceId. The API enforces team visibility; this tool does not accept actorId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
teamWorkspaceIdYesTeam workspace id returned by team_onboarding_status.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds details beyond annotations: requires authenticated API authority, effect is read, no human approval needed, and failure behavior mapping, which supplements the readOnlyHint and idempotentHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with sections but slightly verbose with multiple lines of failure cases; generally front-loaded and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given presence of output schema, description adequately covers prerequisites, failure handling, and effect, making it complete for agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema covers parameter fully with 100% description coverage. Description adds context that teamWorkspaceId comes from team_onboarding_status, adding value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'read a team workspace retention policy' and clearly differentiates from siblings like update_retention_policy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Includes explicit 'Use when' and 'Do not use when' sections, with examples and conditions for not using, providing clear decision guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_weekly_reviewGet Weekly ReviewA
Read-onlyIdempotent

Use when: Read the deterministic current weekly review and return its Web review handoff URL. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint, idempotentHint, destructiveHint), the description adds behavioral details: effect is read, human approval not required, and specific failure handling steps (e.g., login_required → login). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main purpose and uses a structured format with headings. While it is somewhat lengthy due to comprehensive failure handling, every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description is complete: it covers purpose, usage, behavior, failure modes, and post-action steps without redundancy. No significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and fully documents the projectId parameter with default behavior and error conditions. The description adds no additional semantic value beyond what the schema already provides, meeting baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the deterministic current weekly review and returns its Web review handoff URL. It uses a specific verb and resource, distinguishing it from sibling tools like review_pending or get_project_sync_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides 'Use when' and 'Do not use when' conditions, including alternatives when a narrower tool matches intent or when project scope is unresolved. It also outlines prerequisites and failure handling, giving comprehensive guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_wiki_pageGet Wiki PageA
Read-onlyIdempotent

Use when: Read the current or a historical immutable PageVersion. Current root Brief reads also expose non-canonical live sections and a deterministic empty-Brief bootstrap proposal. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYes
versionNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds detailed behavioral context: effect is read, required authentication, human approval not needed, and explicit failure handling instructions. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Use when, Do not use when, Requires, etc.), but it is somewhat verbose and includes some redundancy (e.g., 'Effect: read' is already evident). Still, it is organized and front-loaded with key guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (versioning, multiple failure modes, dependency on project selection), the description is comprehensive. It covers authentication, human approval, error recovery, and interaction with proposals/drafts. Output schema exists, so return values need no elaboration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, with only projectId having a description. The description mentions reading current or historical versions, implying the role of the 'version' parameter, but does not explicitly define pageId or version semantics beyond the schema. Partial compensation but insufficient detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool reads the current or a historical PageVersion, including specific behavior for root Brief reads. It clearly distinguishes from sibling tools like list_wiki_pages and create_wiki_page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear context for when to invoke this tool, including indications of alternative tools and conditions that preclude use (e.g., unresolved project scope, user declined).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

governance_ir_appendixGovernance Ir AppendixA
Read-onlyIdempotent

Use when: Human governance tool: export an aggregate-only markdown IR appendix for the selected project. Raw conversation, context package body, recall answer text, and evidence quote payloads are excluded by the API. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=true and destructiveHint=false, and the description adds that the effect is read, raw data is excluded, and details failure handling (e.g., login_required, stale_version). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear sections (Use when, Do not use when, etc.) and front-loaded with key information. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low parameter count, presence of output schema, and rich annotations, the description covers purpose, usage, behavior, error handling, and follow-up steps comprehensively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with one optional parameter. The description adds context on default behavior (project selection) and rejection scenario, enhancing understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool exports an aggregate-only markdown IR appendix for the selected project. It uses specific verb 'export' and identifies the resource, though it does not explicitly differentiate from sibling tools beyond the context of being a governance tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides 'Use when' and 'Do not use when' conditions, lists requirements, and specifies human approval needs. This gives clear guidance on when to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

grant_agent_project_accessGrant Agent Project AccessA

Use when: Human project-manager tool: grant or reactivate a project-scoped agent principal. The API generates actorId as agent:: and enforces manageProject. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYesProject role to grant to the agent.
statusNoOptional project access status. Defaults to active when granting access.
agentKeyYesProject-local agent key, for example design-reviewer or codex-worker-1.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (all false). Description adds critical behavioral details: canonical mutation, no idempotency key, human approval required, and retry prohibition. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections and front-loaded usage guidance. Slightly verbose but every sentence adds value. Could be tightened slightly without losing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity (4 params, output schema, many siblings), the description covers purpose, usage, effects, human approval, follow-up, and failure modes. No gaps for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with existing parameter descriptions. The description does not add new parameter-level meaning beyond examples already in schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it grants or reactivates a project-scoped agent principal, with specific verb and resource. It distinguishes from siblings like list_agent_access and update_agent_project_access by focusing on grant/reactivate action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear context and alternatives. Also specifies requirements (authenticated API, selected Project) and failure handling strategies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoffHandoffA

Use when: Create a project-scoped handoff context package for the next AI conversation. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoOptional context package title.
topicNoOptional topic to focus the context package.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
sessionIdNoOptional session id. Defaults to the selected project session.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses behavioral traits beyond annotations, such as 'canonical mutation', lack of idempotency key, and detailed error handling instructions. Annotations indicate readOnlyHint=false, idempotentHint=false, destructiveHint=false, and the description adds context on when to use and the effect on state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections, though somewhat verbose. Every sentence adds value, but could be slightly more concise without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers all necessary aspects: purpose, usage, requirements, error handling, and behavioral traits. An output schema exists, so return values are not needed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 4 parameters are documented in the schema with descriptions, and the description adds value by explaining defaults (e.g., 'Defaults to the project chosen with select_project') and conditions for rejection. This goes beyond the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Create a project-scoped handoff context package for the next AI conversation.' This provides a specific verb and resource, and it distinguishes the tool from siblings like 'list_context_packages' or 'get_context_package'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes explicit 'Use when:' and 'Do not use when:' sections, outlining prerequisites, required authentication, and error handling steps. It also specifies when not to use the tool, such as when a narrower tool matches or the user declines.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_agent_accessList Agent AccessA
Read-onlyIdempotent

Use when: Human project-manager tool: list agent principals that can access the selected project. Requires project management permission in the Abyss API. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint, idempotentHint, and destructiveHint; the description adds that the effect is 'read' and details failure handling and human approval not required, providing valuable context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections, but it is somewhat verbose with detailed failure handling and generic statements that could be trimmed without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with an output schema, the description covers purpose, usage, conditions, and failure modes adequately, though it lacks specifics about output format (mitigated by output schema).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description does not add additional information about the projectId parameter beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and the resource 'agent principals' in the context of a selected project, distinguishing it from siblings like grant_agent_project_access and update_agent_project_access.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear guidance on when to use this tool versus alternatives, and specify requirements like project management permission.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_audit_eventsList Audit EventsA
Read-onlyIdempotent

Use when: Human governance tool: list checkpoint audit events for the selected project. Use this to inspect propose, commit, and discard history before governance/export work. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoOptional pagination offset.
perPageNoOptional page size.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
checkpointIdNoOptional checkpoint id to filter the audit timeline.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds extensive behavioral detail beyond annotations: 'Effect: read', 'Human approval: not required', post-action steps, and error handling scenarios (e.g., login_required, project_not_selected). This provides comprehensive transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections but is verbose (multiple lines). While clear, it could be more concise without losing key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (optional parameters, error handling, governance context), the description covers usage, conditions, effect, post-action, and error recovery. The presence of an output schema reduces the need to explain return values, but the description is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the description does not need to repeat parameter details. The description adds minor context (e.g., projectId defaults) but no new semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists checkpoint audit events for a selected project, with a specific use case: inspecting propose, commit, and discard history before governance/export work. It distinguishes from siblings by focusing on audit events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes explicit 'Use when' and 'Do not use when' sections, providing context for when to use this tool versus alternatives. It mentions 'a narrower tool better matches the intent' but does not name specific sibling tools, slightly reducing specificity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_change_proposalsList Change ProposalsA
Read-onlyIdempotent

Use when: List durable proposals awaiting human review in the selected Project. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly=true and idempotent=true. Description adds 'Effect: read', 'Human approval: not required', and detailed failure scenarios (login_required, project_not_selected, permission_denied, stale_version, projection_pending) with prescribed actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is structured and contains valuable information, but it is somewhat verbose with multiple sections. Could be more concise while maintaining clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool (1 optional param, no nested objects, output schema exists), the description covers purpose, usage, effects, failure modes, and follow-up actions completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers the single parameter with description. Tool description adds context: optional with default from select_project, and rejection behavior if omitted without a selected project. This goes beyond schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List' and resource 'durable proposals awaiting human review in the selected Project'. It provides specific scope but does not explicitly differentiate from sibling tools like 'open_proposal_review' beyond generic 'narrower tool' guidance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections with specific conditions: narrower tool, unresolved project scope, declined action. Also states prerequisites (authenticated API authority, selected Project).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_context_packagesList Context PackagesA
Read-onlyIdempotent

Use when: List reusable context packages in the selected Abyss project. Use this before get_context_package or reuse_context_package when continuing team work. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoOptional pagination offset.
perPageNoOptional page size.
purposeNoOptional package purpose filter.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
sessionIdNoOptional session id filter.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=true, destructiveHint=false) are complemented by description stating 'Effect: read' and 'Human approval: not required for this read or staging action'. The description also details failure modes (login_required, project_not_selected, etc.) and follow-up actions, providing rich behavioral context beyond annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with clear sections ('Use when', 'Do not use when', 'Requires', etc.), making it scannable. Every sentence adds value, but conciseness could be improved by reducing redundancy (e.g., 'read or staging action' vs. annotations).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters, output schema exists, and many siblings, the description covers purpose, usage, effect, error handling, and post-action guidance ('Then: follow typed result state; review pending proposals/drafts'). It is complete and addresses common scenarios. The output schema covers return values, so no need to describe them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description does not add parameter-specific semantics beyond the schema. However, the 'On failure' section indirectly relates to parameter usage (e.g., projectId validation). No extra credit for parameter details, but no deduction either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists reusable context packages in the selected Abyss project. It uses a specific verb ('List') and resource ('context packages'), and distinguishes itself from siblings like 'get_context_package' (retrieve one) and 'reuse_context_package' (action). The purpose is unambiguous and well-scoped.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use ('Use this before get_context_package or reuse_context_package when continuing team work') and when not to use ('a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action'). This is comprehensive and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList ProjectsA
Read-onlyIdempotent

Use when: List the user’s Abyss projects with the API-collapsed safe navigation destination. Call select_project with a returned projectId before project-scoped tools. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoOptional pagination offset.
searchNoOptional search text.
statusNoOptional filter, for example active or archived.
perPageNoOptional page size.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses effect as 'read', no human approval needed, and covers failure modes (login_required, permission_denied, etc.). Annotations already indicate read-only and idempotent, but description adds valuable context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with clear sections (Use when, Requires, Effect, etc.). Slightly lengthy but each part adds value. Could be more concise by removing redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers usage, prerequisites, failure scenarios, and post-list actions (select_project, review pending proposals). Output schema exists, so description does not need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline 3. Description does not add parameter-level details beyond schema, which is acceptable given high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool lists user's Abyss projects and instructs to call select_project afterward. Differentiates from sibling tools like create_project and select_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' sections, and lists requirements. Does not name specific narrower tools, but implies they exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_wiki_pagesList Wiki PagesA
Read-onlyIdempotent

Use when: List the selected Project Wiki root Brief and bounded Page hierarchy. Project names are never authorization evidence; the API authorizes the resolved projectId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description confirms the read-only, idempotent, non-destructive nature consistent with annotations. It details effects, human approval needs, and provides explicit failure mode handling (e.g., login_required, project_not_selected). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Use when, Do not use when, Requires, Effect, Human approval, Then, On failure). Every sentence adds value, and the structure aids quick comprehension despite length. Front-loaded with critical guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and comprehensive annotations, the description covers all necessary aspects: purpose, usage boundaries, behavioral effects, failure modes, and post-action steps. It is complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, projectId, is fully described in the schema (100% coverage). The description repeats the schema information and adds the note about authorization but does not provide additional semantic nuance beyond the schema. Baseline score applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists the selected Project Wiki root Brief and bounded Page hierarchy. It distinguishes from sibling tools like get_wiki_page by focusing on the hierarchy. The special note about project names not being authorization evidence adds clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear context. It specifies when to use (list hierarchy) and when not (if a narrower tool matches or project scope unresolved). It also lists requirements like authenticated authority and selected project, and gives failure handling instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

loginLoginA

Use when: Start Abyss browser device login. After browser approval, call complete_login, then auth_status. Authentication does not prove project scope; verify list_projects and select_project next. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description contradicts annotations: annotations set readOnlyHint=false, indicating not read-only, but the description calls it a 'read or staging action'. This inconsistency undermines transparency despite other behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is quite long but each sentence adds value, with the key use case front-loaded. Slightly verbose, but efficient given the amount of information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters and annotations are present, the description thoroughly covers the login flow, failure scenarios, required subsequent steps, and operational constraints, making it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has no parameters (0 params, 100% coverage), so baseline is 4. The description does not need to add parameter info; it appropriately focuses on process and context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Start Abyss browser device login', specifying the verb (start) and resource (login). It distinguishes from siblings like complete_login and auth_status by indicating the sequential flow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use, with 'Use when:' and 'Do not use when' conditions. Also provides clear next steps (complete_login, auth_status, list_projects, select_project) and failure handling, offering comprehensive guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

logoutLogoutA
Destructive

Use when: Revoke the current Abyss MCP credential and remove the local credential file. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (destructiveHint: true, idempotentHint: false), the description adds crucial behavioral details: 'Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome.' It also outlines failure handling steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections but is slightly verbose. It front-loads the core purpose effectively, though some parts (e.g., 'Then: follow typed result state') could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, an existing output schema, and detailed coverage of usage, failure handling, and safety, the description is fully complete for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters and 100% coverage, so the description naturally adds no parameter information. Baseline 4 applies as no further elaboration is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Revoke the current Abyss MCP credential and remove the local credential file.' This is a specific verb+resource pair that clearly distinguishes it from sibling tools like login, auth_status, and whoami.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes 'Use when:', 'Do not use when:', and 'Requires:' sections that provide explicit guidance on appropriate usage contexts and prerequisites, making it highly informative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_proposal_reviewOpen Proposal ReviewA
Read-onlyIdempotent

Use when: Fetch a proposal and return a Web review handoff URL. Human review/apply remains outside the agent-safe MCP surface. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
proposalIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations. It confirms the effect as 'read', states that human approval is not required, and provides detailed error handling and follow-up instructions (e.g., 'On failure: login_required → login; ...'). Annotations already declare readOnlyHint, idempotentHint, and destructiveHint as true/true/false, and the description aligns with and extends these with actionable guidance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear sections ('Use when', 'Do not use when', 'Requires', etc.) and is front-loaded with the purpose. However, it is verbose, containing detailed error handling and follow-up instructions that could be more concisely presented. It could be trimmed by half without losing essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the parameter semantics gap, the description provides comprehensive context: purpose, usage conditions, prerequisites, effect, approval requirements, error handling, and post-use actions. An output schema exists, so return value details are not required. For a tool with two parameters and moderate complexity, the description covers all necessary aspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50% (only projectId has a description; proposalId does not). The description does not add new semantic information about the parameters beyond what is already in the schema. It mentions the projectId default behavior, but that is already in the schema's description. For proposalId, no additional context is provided. Given the low coverage, the description should compensate but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's core action: 'Fetch a proposal and return a Web review handoff URL.' It uses a specific verb ('fetch' and 'return') and resource ('proposal', 'Web review handoff URL'). The description also distinguishes the tool from siblings by noting that human review/apply remains outside the agent-safe MCP surface, and provides explicit conditions for when not to use it (e.g., 'a narrower tool better matches the intent').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes explicit 'Use when' and 'Do not use when' sections, providing clear context for tool invocation. It lists prerequisites ('authenticated API authority and an API-authorized selected Project'). It mentions alternatives indirectly ('a narrower tool better matches the intent') but does not name specific sibling tools, which would have made it a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_project_resumePrepare Project ResumeA
Idempotent

Use when: Create a task-bounded ContextPackage with an included/excluded/warning manifest. Package creation never advances the sync cursor. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
taskIdYes
scopeIdYes
targetIdNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeYes
toVersionNo
targetTypeNo
connectionIdYes
idempotencyKeyYes
selectedChangeIdsNoOptional subset from review_project_delta. Accepts changes[].id (change-set ids) and resolvedChanges[].changeId (individual resolved change ids). If safety requires a hydrated full resync, the current authorized snapshot can expand beyond the requested subset and the manifest reports selected_changes_expanded_for_full_resync.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond annotations: 'Package creation never advances the sync cursor', 'Effect: canonical mutation', idempotency key behavior, human approval not required, and detailed failure handling. No contradictions with annotations (idempotentHint=true, destructiveHint=false).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections ('Use when', 'Do not use when', 'Requires', 'Effect', 'Human approval', 'Then', 'On failure'). Every sentence provides necessary information without redundancy, appropriate for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 parameters, 6 required, output schema exists), the description covers usage conditions, behavioral effects, failure modes, and idempotency details. It provides sufficient information for an AI agent to correctly select and invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 18%, meaning most parameters have no schema description. The main description does not directly explain individual parameters like 'connectionId', 'scopeType', 'taskId', etc. While the description provides overall context, it fails to add meaning to each parameter beyond what little the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with 'Use when: Create a task-bounded ContextPackage with an included/excluded/warning manifest', which provides a specific verb and resource. It also includes explicit exclusions ('Do not use when...'), clearly distinguishing the tool from siblings like 'review_project_delta' or 'propose_wiki_changes'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use ('Use when') and when not to use ('Do not use when') with specific conditions. It also lists prerequisites: 'Requires: authenticated API authority and an API-authorized selected Project'. This provides comprehensive guidance beyond typical descriptions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_wiki_changesPropose Wiki ChangesA
Idempotent

Use when: Persist Page/Object operations as a digest-bound ChangeProposal. This agent-safe tool cannot apply, reject, or choose a conflict winner. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: proposal. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: required at the later commit/apply boundary. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
summaryYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
operationsYes
connectionIdNo
idempotencyKeyYes
sourceSessionIdNo
sourceExecutionIdNo
baseProjectVersionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds significant context beyond annotations: agent-safe, idempotency behavior, human approval required later, failure handling steps (login_required, project_not_selected, etc.). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with bolded sections (Use when, Do not use when, Effect, Then, On failure). Slightly long but every sentence adds value; front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers usage conditions, idempotency, auth, project selection, human approval, failure handling. Output schema exists so return values not needed. Operations array is not detailed but that is acceptable given complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 11%, yet the description does not explain individual parameters beyond idempotencyKey usage. The schema is detailed but lacks descriptions; description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states verb 'persist' and resource 'Page/Object operations as a digest-bound ChangeProposal'. Distinguishes from siblings by noting it cannot apply/reject/choose conflict winner, but does not explicitly name a narrower alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when:' and 'Do not use when:' sections provide clear conditions for invocation. Mentions alternatives (narrower tool) and user intent rejection. Also lists prerequisites (authentication, project selection).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recallRecallA
Read-onlyIdempotent

Use when: Human-account read-only recall. Requires an active project selected with select_project. Project agent credentials are intentionally limited to checkpoint/resume/handoff flows. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
topKNoOptional bounded recall count, 1-50.
intentYesRecall intent.
subjectNoOptional decision or topic to focus on.
questionNoOptional user-written recall question.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnlyHint=true, destructiveHint=false. The description adds 'Effect: read' and 'Human approval not required,' plus detailed error cases, extending behavioral understanding beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with key usage guidelines, structured with clear sections. It is moderately lengthy but each sentence adds value; could be slightly more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and rich annotations/output schema, the description covers usage, prerequisites, error handling, and effect. It does not detail output schema (not required) but is sufficiently complete for effective agent invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all parameters. The description adds minimal extra parameter-level context, e.g., projectId defaults, but does not elaborate on intent enum values. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is for 'Human-account read-only recall' and lists specific intents (decision_reason, changed_view, etc.). It distinguishes from narrower tools by advising against use when a better match exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes explicit 'Use when' and 'Do not use when' sections, specifies prerequisites (active project, credentials), and provides error handling guidance, offering comprehensive usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recall_decision_reasonRecall Decision ReasonA
Read-onlyIdempotent

Use when: Read-only compatibility alias for recall intent decision_reason. Prefer recall for new clients. Deprecated compatibility alias. Prefer recall; do not select this alias for new workflows. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
topKNoOptional bounded recall count, 1-50.
subjectNoOptional decision or topic to focus on.
questionNoOptional user-written recall question.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true. The description adds effect 'read', failure details, and post-action steps, which provides good context beyond annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with labeled sections, front-loaded with key purpose. Slightly verbose but each sentence adds value; no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage boundaries, effect, required auth, failure modes, and post-action steps. Given zero required params and presence of output schema, the description is fully adequate for correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with all parameters described. The description repeats the parameter info without adding new semantics, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it's a read-only compatibility alias for recall intent decision_reason and explicitly prefers the sibling tool 'recall' for new clients, distinguishing it effectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when', 'Do not use when' conditions, required authentication, effect, human approval requirements, and failure handling, offering comprehensive guidance for selection and invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_context_consumptionRecord Context ConsumptionA

Use when: Record that a delivered ContextPackage was actually used by an AIExecution. Delivery and consumption remain separate facts. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
executionIdYes
impressionIdsNo
actualPromptHashYes
contextPackageIdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses important behavior beyond annotations: it's a canonical mutation with no client idempotency key, advising not to retry automatically. Also clarifies human approval is not required and describes failure modes. This adds significant context not available from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections, but somewhat verbose due to detailed failure enumeration. Could be more concise without losing essential guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description provides comprehensive guidance on usage, failure handling, and behavioral expectations. It covers when to use, when not, and what to do on various failures.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only projectId has a description). The tool description does not elaborate on the meaning of executionId, contextPackageId, impressionIds, or actualPromptHash, leaving the agent with minimal understanding of parameter purposes beyond the sparse schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Record that a delivered ContextPackage was actually used by an AIExecution'—a specific verb+resource. It distinguishes from sibling tool 'confirm_context_delivery' by noting that delivery and consumption are separate facts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when:' and 'Do not use when:' conditions, including when to avoid using (narrower tool, unresolved scope, user declined). Also includes failure handling instructions, guiding correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_project_sync_decisionRecord Project Sync DecisionA
Idempotent

Use when: Record a version-bound preview, one-time skip, or 24-hour snooze decision. This never claims delivery or consumption. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
scopeIdYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeYes
connectionIdYes
idempotencyKeyYes
throughVersionYes
expectedCursorVersionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses key behavioral traits: canonical mutation, idempotency key reuse policy (same key for retry, new key for changed intent), replay behavior, no human approval needed, and detailed failure mappings. This goes well beyond the annotations (idempotentHint=true, readOnlyHint=false, destructiveHint=false) without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections and front-loaded purpose. Some verbosity in details, but every section earns its place. Could be slightly more concise but still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters (7 required) and an output schema, the description covers usage, behavior, error handling, and postconditions (e.g., follow typed result, review proposals). Output schema exists so return values need not be explained. No gaps identified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%; description provides context for 'action' (preview, skip, snooze) and idempotencyKey reuse, but does not explicitly explain other parameters like scopeType, scopeId, connectionId, or throughVersion. Partial compensation but insufficient for low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly states the action (record a version-bound preview, one-time skip, or 24-hour snooze decision) and what it does not do (never claims delivery or consumption), clearly distinguishing it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' conditions, requirements (authenticated API authority, selected Project), and a 'Then' section for post-action steps, offering comprehensive guidance on when to use versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resumeResumeA

Use when: Create a project-scoped resume context package for continuing the current AI chat. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoOptional context package title.
topicNoOptional topic to focus the context package.
taskIdNo
scopeIdNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
scopeTypeNo
sessionIdNoOptional session id. Defaults to the selected project session.
connectionIdNo
syncDecisionNo
idempotencyKeyNo
selectedChangeIdsNoRequired when syncDecision is selected. Accepts changes[].id and resolvedChanges[].changeId from review_project_delta. A safety-driven full resync can expand to the current authorized snapshot and reports selected_changes_expanded_for_full_resync.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the effect ('canonical mutation'), idempotency constraints ('no client idempotency key; do not retry'), human approval (not required for read/staging), and detailed error paths ('On failure: login_required → login;...'). Annotations are sparse but consistent; the description adds essential behavioral context beyond them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with clear headings ('Use when', 'Then', 'On failure', etc.), making it easy to scan. However, it is somewhat verbose, containing repetitive phrases like 'project-scoped' and 'API-authorized'. Could be tightened without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 params, 0 required, has output schema, many siblings), the description is remarkably complete: it covers prerequisites, errors, post-action steps, and references the output schema. Little is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 45%, leaving many parameters undocumented in the schema. The description does not add parameter-level semantics beyond the schema; it focuses on usage and error handling. Without parameter details, tool invocation may be unclear for some parameters. A moderate score reflects minimal added value for parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a 'project-scoped resume context package' for continuing the AI chat, distinguishing it from sibling tools like recall or handoff. The verb 'Create' and resource 'resume context package' are specific. It implies a distinct purpose among many sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear guidance on appropriate use cases and exclusions. Additionally, it lists prerequisites ('authenticated API authority', 'selected Project') and postconditions ('Then: follow typed result state').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reuse_context_packageReuse Context PackageA
Idempotent

Use when: Record that the selected project reused a context package in another prompt, conversation, document, export, publish, or report target. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. For the same logical mutation retry, reuse the same idempotencyKey. If the payload or user intent changes, use a new key. Replay returns the first canonical result ID and version; the key never bypasses authorization or expected-version checks. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoOptional UI or client action, for example copy_context.
metadataNo
targetIdNoOptional external target id.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
sessionIdNoOptional session id. Defaults to the selected project session.
clientTypeNoOptional client type. Defaults to mcp.
targetTypeYesWhere the package was reused.
graphVersionNoOptional graph version from the package.
contextLengthNoOptional reused context character count.
evidenceCountNoOptional reused evidence count.
idempotencyKeyNoOptional idempotency key for mutation replay safety.
contextPackageIdYesContext package id returned by resume, handoff, or list_context_packages.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate idempotent mutation. Description adds extensive behavioral context: canonical mutation, idempotency key replay, authorization checks, expected-version handling, human approval not required, failure modes with actions, and post-action steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections (Use when, Requires, Effect, Human approval, Then, On failure). Front-loaded with core action. Slightly verbose but justified by complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 12 parameters, 2 required, high schema coverage, and existing output schema, the description adds necessary workflow context, failure handling, and post-action guidance. Fully complete for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 92%, so the schema already documents most parameters. The description adds context about idempotency key usage but does not significantly expand on parameter meanings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Record that the selected project reused a context package' with specific target types. It names the verb and resource, but does not differentiate from sibling tools like record_context_consumption.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when' and 'Do not use when' sections provide clear context, including conditions like 'project scope unresolved' and 'user declined'. Mentions that a narrower tool may be better but does not list specific alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_pendingReview PendingA
Read-onlyIdempotent

Use when: List draft checkpoints for the selected project before they are committed or discarded. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoOptional pagination offset.
perPageNoOptional page size.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true, openWorldHint=false, indicating a safe read operation. The description adds value by confirming 'Effect: read', stating human approval is not required, and detailing on-failure behaviors (login_required, project_not_selected, permission_denied, etc.), providing behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, using clear section headings ('Use when:', 'Do not use when:', etc.). Every sentence adds value, and the core purpose is front-loaded. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters with full schema coverage, existing annotations, and an output schema, the description is complete. It covers usage context, prerequisites, failure scenarios, and next steps, leaving no critical gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds context by clarifying that projectId defaults to the selected project and that pagination parameters are optional. This additional meaning justifies a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists draft checkpoints for a selected project before they are committed or discarded. It uses a specific verb ('list') and resource ('draft checkpoints'), distinguishing it from sibling tools like commit_checkpoint and discard_checkpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use ('List draft checkpoints for the selected project') and when-not-to-use instructions ('a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action'). It also lists requirements (authenticated API authority, selected project) and human approval info, giving clear guidance on when to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_pending_connectionsReview Pending ConnectionsA
Read-onlyIdempotent

Use when: Read-only compatibility tool for suggested thoughts or connections that need confirmation. Prefer review_pending for checkpoint drafts. Deprecated compatibility alias. Prefer review_pending or list_change_proposals; do not select this alias for new workflows. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional number of review items to return.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes beyond annotations by stating required authentication and project selection, that it is a read effect, and that no human approval is needed. Detailed failure handling is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections, front-loading key info. Slightly verbose due to exhaustive failure handling, but every section adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage, requirements, behavior, and error handling thoroughly. Output schema exists, so return values are not needed. Annotations complement the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description does not add further detail about parameters beyond what the input schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a read-only compatibility alias for reviewing pending suggestions/connections, and explicitly distinguishes it from siblings like review_pending and list_change_proposals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' and 'Do not use when' conditions, suggests preferred alternatives, and advises against selecting this alias for new workflows.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_project_deltaReview Project DeltaA
Read-onlyIdempotent

Use when: Preview ordered ProjectChangeSets by contextVersion. This read does not create a package or advance a cursor. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
toVersionNo
fromVersionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description adds behavioral details beyond annotations: 'This read does not create a package or advance a cursor,' 'Effect: read,' error handling steps. No contradiction with annotations (readOnlyHint=true, destructiveHint=false).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (use when, do not use, requires, effect, then, on failure). It is comprehensive but slightly verbose; a bit more brevity could improve it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 params, output schema present, annotations present), the description is thorough. It covers purpose, usage, behavior, error handling, and post-call actions ('follow typed result state; review pending proposals/drafts'). No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% with no descriptions for fromVersion and toVersion. The description mentions 'by contextVersion' but does not explain the parameters. It fails to compensate for the low schema coverage, leaving agents uncertain about parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Preview ordered ProjectChangeSets by contextVersion.' It specifies the resource (ProjectChangeSets) and distinguishes from siblings by noting it does not create a package or advance a cursor, and suggests not using when a narrower tool matches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use ('Preview ordered ProjectChangeSets by contextVersion') and when-not-to-use ('a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action'). Also lists prerequisites: authenticated API authority and API-authorized selected project.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_project_contentSearch Project ContentA
Read-onlyIdempotent

Use when: Exact/lexical search across authorized current Wiki Pages and built-in Continuity Object revisions. Projection watermarks are returned separately from canonical commit. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
includeObjectsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and idempotentHint; the description adds valuable behavioral context like effect 'read', error handling scenarios, and the fact that projection watermarks are returned separately. Minor omission: no mention of pagination or limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured into clear sections, front-loads the core purpose, and each sentence adds value. Though somewhat verbose, it remains focused and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and annotations, the description covers purpose, guidelines, requirements, error handling, and follow-up actions. The only gap is parameter documentation, but overall it provides a solid understanding of the tool's role.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only projectId has a description). The description does not elaborate on the meaning or usage of parameters like limit, includeObjects, or query, leaving the agent without sufficient detail to use them correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs exact/lexical search across authorized wiki pages and continuity object revisions, explicitly distinguishing it from narrower tools like get_wiki_page or get_continuity_object.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' and 'Do not use when' conditions, along with requirements (authenticated API authority, selected project), offering clear guidance on when to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_projectSelect ProjectA

Use when: Select an Abyss project and open an API-backed continuity session. If this host exposes conversation identity, title, observation times, or a safe reference, send them in externalConversation. Mark unavailable host fields instead of inventing them. Abyss session identifiers and lifecycle timestamps remain API-owned. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNoOptional session title.
scopeIdNo
projectIdYesA project id returned by list_projects.
scopeTypeNo
sourceRefNoLegacy safe source reference. Prefer externalConversation.sourceRef for V1 capture.
connectionIdNo
externalConversationNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide limited info (non-read-only, non-destructive, non-idempotent). The description adds critical context: 'canonical mutation', 'no client idempotency key', 'do not retry automatically'. It also describes failure modes without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections, but it is lengthy. Every sentence adds value, though some details (e.g., failure cases) could be more compact. Front-loading the main purpose is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (7 params, nested objects, output schema exists), the description covers usage context, failure modes, behavioral notes, and some parameter guidance. It does not detail return values (output schema exists) but is otherwise thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 43% (low). The description explains the externalConversation parameter and its fields in detail, compensating for part of the gap. However, it does not describe other parameters like scopeId, scopeType, or connectionId, which remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Select an Abyss project and open an API-backed continuity session.' This is a specific verb+resource, and the 'Do not use when' section distinguishes it from narrower tools or when the project scope is unresolved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use when:' and 'Do not use when:' sections provide clear guidance. It also lists requirements and failure handling, e.g., 'Requires: authenticated API authority' and 'On failure: login_required → login'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

share_abyss_feedbackShare Abyss FeedbackA

Use when: Offer an optional Abyss satisfaction survey, submit explicitly confirmed product feedback, prepare a user-reviewed email, manage feedback prompt preference, or subscribe to the newsletter with separate explicit consent. You may suggest this once after the tool reports that at least 3 meaningful Abyss calls succeeded, or when the user expresses praise, frustration, a bug, or a feature request. Never interrupt the primary task, repeatedly solicit feedback, submit feedback, subscribe an email, or send an email without the user explicitly asking and confirming. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNo
scoreNo
actionYes
clientNo
consentNo
messageNo
categoryNo
confirmedNo
preferenceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses it is a canonical mutation without idempotency key, requires human approval, and lists failure handling behaviors (login_required, permission_denied, etc.). Adds context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured with clear sections for when to use, not to use, effects, failure handling. However, it is somewhat verbose and could be more concise without losing essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers usage context, failure modes, and human approval, but lacks parameter documentation. With 9 parameters and 0% schema coverage, the description should provide more parameter-level guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not explain individual parameters like email, score, consent, message, category, confirmed, preference, or client. The action enum is described in purpose, but other parameters lack guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool offers a survey, submits feedback, prepares email, manages prompt preference, and subscribes to newsletter. It distinguishes from siblings by listing specific actions not covered by other tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (after 3 successful calls or user praise/frustration/bug/feature request) and when not to use (interrupting primary task, without explicit consent, narrower tool available, project scope unresolved, declined action). Provides clear guidelines.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_ai_executionStart AI ExecutionA

Use when: Start the durable AIExecution that will consume an already delivered ContextPackage. The API, not MCP, owns authorization and may return a request-scoped execution grant; MCP must pass that result through without interpreting billing or payment state. Call this after confirm_context_delivery and before record_context_consumption. Do not retry automatically because this creates a new execution. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
statusNo
taskIdYes
providerNo
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.
sessionIdNo
connectionIdNo
inputSyncReceiptIdsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses the mutation nature (no readOnlyHint), lack of idempotency, authorization delegation to API, and no human approval needed, adding value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear sections, but somewhat verbose; sections like 'Human approval' could be more concise given annotations already cover this.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers failure modes and workflow placement well, but lacks description of return value structure and parameter details, given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 13% schema description coverage, the description does not explain the meaning of critical parameters like taskId or inputSyncReceiptIds, leaving a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Start the durable AIExecution that will consume an already delivered ContextPackage' and positions it between confirm_context_delivery and record_context_consumption, clearly differentiating it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Includes explicit 'Use when' and 'Do not use when' clauses, retry warnings, and a structured 'On failure' section covering common errors, providing comprehensive guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

team_onboarding_statusTeam Onboarding StatusA
Read-onlyIdempotent

Use when: Human-account tool: show the current account’s team onboarding status and recommended next step. It does not accept actorId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds significant behavioral context beyond annotations: specifies it does not accept actorId, requires authenticated API authority, effect is read, human approval not required, and details error handling (login_required, project_not_selected, permission_denied, stale_version, projection_pending) and subsequent steps. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with labeled sections ('Use when', 'Do not use when', etc.) and front-loads key information. However, it is somewhat verbose with details like 'review pending proposals/drafts before any canonical apply' which could be more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, an output schema, and rich annotations, the description covers usage context, authentication, effect, error handling, and post-action steps. It provides complete guidance for the AI to understand and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters, and schema coverage is 100%. The description adds that it does not accept actorId, which is not a parameter in the schema, clarifying what the tool does not take. This extra context justifies a score above baseline 3 for 0-param tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb 'show' and resource 'team onboarding status and recommended next step'. It also specifies 'human-account tool' and that it does not accept actorId. However, it does not explicitly distinguish from sibling tools, though no sibling appears to directly overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Use when' and 'Do not use when' clauses, including conditions about narrower tools, unresolved project scope, or user declination. This clearly guides the AI on when to select this tool over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_agent_project_accessUpdate Agent Project AccessA
Destructive

Use when: Human project-manager tool: update a project-scoped agent principal role or status. Use status=removed to disable access without deleting audit history. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority and an API-authorized selected Project. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied �� stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoProject role to grant to the agent.
statusNoOptional project access status. Defaults to active when granting access.
agentKeyYesProject-local agent key to update.
projectIdNoOptional project id. Defaults to the project chosen with select_project. If omitted and no project is selected, the call is rejected with instructions to select one.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description details idempotency (no idempotency key, do not retry), human approval requirement, error recovery steps (login_required, project_not_selected, etc.), and required post-actions (review proposals/drafts). This adds significant behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections (Use when, Do not use when, Requires, Effect, etc.) and front-loaded key information. While slightly lengthy, every sentence contributes value, so it earns a 4 rather than a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all necessary aspects: usage context, parameter behavior, error handling, and follow-up actions. Given the presence of an output schema, the description is fully complete for an AI agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the description adds extra context like using status='removed' to disable access without deleting audit history and clarifying default behaviors for status and projectId. This enhances understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool updates a project-scoped agent principal's role or status, using the verb 'update' and specifying the resource type. It differentiates from the sibling 'grant_agent_project_access' by focusing on existing access modification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides 'Use when' and 'Do not use when' conditions, including prerequisites like authenticated authority and selected project. It also mentions alternative scopes and user consent, guiding appropriate tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_retention_policyUpdate Retention PolicyA
Destructive

Use when: Human-admin tool: update a team workspace retention policy. The API enforces owner/admin workspace permission; this tool does not accept actorId. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: canonical mutation. This mutation has no client idempotency key in the current contract; do not retry it automatically after an unknown outcome. Human approval: required before this action or its canonical follow-up. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
teamWorkspaceIdYesTeam workspace id returned by team_onboarding_status.
defaultRetentionPolicyYesRetention policy JSON, or null to clear it.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses destructive nature (consistent with destructiveHint true), non-idempotency, no retry policy, and permission requirements. Adds significant context beyond annotations, such as 'canonical mutation' and 'no client idempotency key'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with labeled sections and front-loaded purpose. Slightly verbose but all sentences add value; could be trimmed slightly without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers usage context, prerequisites, effects, failure modes, and follow-up steps. With output schema present, this provides complete guidance for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100% with good descriptions. Description adds no further parameter-level detail beyond what schema provides, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'update a team workspace retention policy' with specific verb and resource. Distinguishes from sibling tool 'get_retention_policy' and others via 'Human-admin tool' qualifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'Use when' and 'Do not use when' conditions, including human approval requirement, permission constraints, and failure handling. Gives clear context for appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoamiWhoamiA
Read-onlyIdempotent

Use when: Return the authenticated Abyss principal classification and team onboarding summary. Do not use when: a narrower tool better matches the intent, the project scope is unresolved, or the user has declined the action. Requires: authenticated API authority. Effect: read. Human approval: not required for this read or staging action. Then: follow typed result state; review pending proposals/drafts before any canonical apply. On failure: login_required → login; project_not_selected → list_projects/select_project; permission_denied → stop; stale_version or conflict → read current state; projection_pending → report canonical success separately and wait.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataYes
statusYes
warningsYes
authorityYes
projectIdYes
reviewUrlYes
idempotencyYes
nextActionsYes
reviewRequiredYes
canonicalVersionYes
projectionStatusYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, non-destructive behavior. The description adds further context: 'Effect: read', human approval not required, post-call steps (follow typed result state, review proposals), and detailed failure handling. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is relatively long (95 words) but well-structured with explicit sections (Use when, Do not use when, Requires, etc.). It is front-loaded with the core purpose. Minor verbosity reduces conciseness, but overall acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, the description covers all necessary aspects: purpose, usage boundaries, prerequisites, effects, post-call actions, and failure modes. An output schema exists, so return values are not needed in description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, so schema coverage is 100%. Baseline is 4 as per rules. Description adds no param info, but none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns 'authenticated Abyss principal classification and team onboarding summary', which is a specific verb and resource. It distinguishes itself from sibling tools like 'auth_status' and 'team_onboarding_status' by specifying the combination of principal classification and onboarding.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sections for 'Use when' and 'Do not use when' provide clear guidance on appropriate contexts. Also specifies requirements ('authenticated API authority'), effects, and even failure handling patterns, leaving no ambiguity about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 55 tool updatesv0.5.9
    • First observedask_if_this_should_be_remembered
    • First observedauth_status
    • First observedcheckpoint
    • First observedcommit_checkpoint
    • First observedcompare_wiki_versions
    • First observedcomplete_login
    • First observedconfirm_context_delivery
    • First observedcontinue_previous_thread
    • First observedcreate_project
    • First observedcreate_wiki_page
    • First observeddiscard_checkpoint
    • First observedexecute_block_command
    • First observedexplain_why_linked
    • First observedexport_wiki
    • First observedget_block_candidates
    • First observedget_context_package
    • First observedget_continuity_object
    • First observedget_continuity_operations
    • First observedget_page_backlinks
    • First observedget_project_sync_activity
    • First observedget_project_sync_status
    • First observedget_retention_policy
    • First observedget_weekly_review
    • First observedget_wiki_page
    • First observedgovernance_ir_appendix
    • First observedgrant_agent_project_access
    • First observedhandoff
    • First observedlist_agent_access
    • First observedlist_audit_events
    • First observedlist_change_proposals
    • First observedlist_context_packages
    • First observedlist_projects
    • First observedlist_wiki_pages
    • First observedlogin
    • First observedlogout
    • First observedopen_proposal_review
    • First observedprepare_project_resume
    • First observedpropose_wiki_changes
    • First observedrecall
    • First observedrecall_decision_reason
    • First observedrecord_context_consumption
    • First observedrecord_project_sync_decision
    • First observedresume
    • First observedreuse_context_package
    • First observedreview_pending
    • First observedreview_pending_connections
    • First observedreview_project_delta
    • First observedsearch_project_content
    • First observedselect_project
    • First observedshare_abyss_feedback
    • First observedstart_ai_execution
    • First observedteam_onboarding_status
    • First observedupdate_agent_project_access
    • First observedupdate_retention_policy
    • First observedwhoami

TDQS

A4.1/5.0
Disambiguation4/5

Most tools have clearly distinct purposes with detailed usage descriptions. However, the presence of deprecated alias tools (e.g., recall_decision_reason, review_pending_connections) adds some confusion and may cause incorrect selection.

Naming Consistency4/5

The majority of tools follow a consistent verb_noun pattern (e.g., list_projects, create_wiki_page). A few tools like 'whoami' and deprecated aliases break the pattern, but overall naming is predictable.

Tool Count2/5

With 55 tools, the surface is very large and likely overwhelming for agents. Many tools handle niche operations (e.g., sync, governance) that could potentially be merged or simplified.

Completeness4/5

The tool set covers authentication, project management, wiki, search, sync, proposals, checkpoints, agent access, and governance. Minor gaps exist (e.g., no delete_wiki_page), but the core workflows are well-covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Official MCP server for PostIdentity - Generate AI-powered social media posts, threads, and replies from any MCP-compatible AI assistant with identity management and refinement capabilities.
    17
    1
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    Official MCP server for accepting crypto payments through OrcaRail. It enables AI agents to create payment intents, manage subscriptions, handle product catalogs, and get exchange rates via natural language.
    25
    14
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/my-abyss-project/abyss-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server