Skip to main content
Glama

gpp (git++)

CI License: MIT Rust 2024

AI 에이전트는 코드를 계속 변경하면서 커밋 사이에 일어난 작업을 잃어버립니다. 그리고 코드가 움직이는 동안 코드베이스에 대해 당신이 남겨둔 메모(CLAUDE.md, 메모리 뱅크, 지식 파일)는 조용히 낡아 버립니다. gpp는 바로 그런 현실을 위해 만들어진 버전 관리 시스템입니다. 변경이 일어나는 순간마다 캡처하고, 의도와 출처가 담긴 선별된 변경셋을 승격(승격)하며, 프로젝트 지식을 저장소 자체의 역사 위에 호스팅합니다. 그래서 어떤 커밋이 코드에 대해 믿고 있던 사실을 무효화하면, gpp는 그 커밋을 정확히 지목할 수 있습니다.

gpp demo

(녹화 영상은 scripts/demo.sh로 생성됩니다 — 결정적이고, 재현 가능하며, 목업이 전혀 없습니다.)

30초 체험하기

cargo install gpp-cli

gpp init --graphex
echo "fn main() {}" > main.rs
gpp timeline                     # captured already — no staging, no commit
gpp promote -m "first cut" --intent feature
gpp diff HEAD                    # semantic diff: renames/moves are one op

그리고 다른 VCS는 하지 못하는 부분 — 코드에 대해 알고 있는 것을 기록해 두고, 이후 역사가 그것을 감독하도록 하는 것:

gpp belief add --claim "token expiry is 24h" --evidence src/auth.rs:7-7
# ...weeks of commits later...
gpp belief bisect "token expiry is 24h"
# INVALIDATED  cs:fhcpef7c  "raise token expiry to 7 days"
#  -     7 | pub const EXPIRY_HOURS: u64 = 24;
#  +     7 | pub const EXPIRY_HOURS: u64 = 168;

결정적(eterministic)입니다 — diff 교차와 blob 해시를 사용하며 LLM 이나 네트워크 호출이 전혀 없습니다. (규모를 감안하면: 저장된 메모가 무효화되었는지 판단하라고 요청했을 때 최고 프론티어 모델조차 STALE 벤치마크에서 55.2% 정확도에 그칩니다 — 단, 여기서는 판단이 아니라 히스토리 조회입니다.) 실제 히스토리를 네 가지 언어의 다섯 저장소 — axum 0.6→0.7, flask 1.1→2.0, clap 3→4, zod 3→4, go-redis 8→9 — 에서 검증했으며, 무효화된 모든 belief는 고정되고 문서화된 원인 커밋으로 이분 탐색(bisect)되었고, 대조군 belief는 그대로 유지됩니다. 전체 매트릭스는 demos/belief-bisect/를 참고하세요.

Related MCP server: ITHZ MCP

다른 점

  • 연속 캡처. 고빈도 타임라인이 모든 파일 변경을 기록합니다(SQLite WAL, 디바운스 감시자). 이 타임라인에서 선별된 변경셋은 명시적 의도, 작성자 유형(사람/에이전트), 비용 기록과 함께 승격됩니다. 커밋 사이에서의 손실이 없습니다.

  • Graphex: 버전 관리되며 암호화되는 프로젝트 지식. 아키텍처, 규칙(관례), 결정, 그리고 beliefs는 저장소 내부의 암호화된 지식 그래프에 담기며, 에이전트 신뢰 수준별 티어로 접근이 제한되고, 모든 읽기에서 감사 기록이 남으며, 그 지식을 떠받치고 있는 히스토리에 대한 낡은 상태(staleness)를 검사합니다.

  • 에이전트 거버넌스. 평판 점수를 포함한 일급 에이전트 정체성, 캡처/승격/동기화 단계에서 강제되는 코드형 규정준수(compliance-as-code) 정책, 이상 감지, 그리고 변경셋별 토큰·비용 배분.

그 외의 모든 것 — Noise 기반 P2P 동기화, 재생(replay), 리뷰/RBAC, 릴레이 노드, TUI — 은 그 밑을 받치는 플랫폼입니다: docs/ARCHITECTURE.md를 참고하세요.

Git은 하부 구조이지 경쟁자가 아닙니다. 그 브릿지(gpp git-import / gpp git-export / gpp git-브릿지)가 실제 Git 커밍을 왕복시키므로 GitHub 과 기존 워크플로우는 계속 동작합니다 — gpp의 지식·출처·거버넌스 계층은 그 곁에서 함께 움직입니다. 위의 axum 데모는 전적으로 이 브로의 교체를 통해 가져온 히스토리 위에서 실행됩니다.

Claude Code 연결(또는 모든 MCP client)

gpp는 MCP 서버를 내장합니다. 에이전트는 지식 그래프를 조회하고(모든 belief에는 신선도 포장 — 앵커, 이후 커밋 수, 그리고 그 belief를 낡게 만든 변경셋 — 이 붙어 있습니다), propose_belief를 사용해 근거가 있는 자신의 belief recording을 작성하며(사람이 승인한 뒤, 이후 히스토리로 감시), changeset을 제안하고 비용을 보고합니다. 저장소 루트의 .mcp.json에 다음을 넣으세요:

{
  "mcpServers": {
    "gpp": { "command": "gpp", "args": ["mcp-server", "--stdio"] }
  }
}

자세한 클라이언트 설정(Claude Desktop, 일반 stdin/stdout) 과 노출된 도구 목록: docs/MCP.md.

현황

로드맵의 9개 단계(0–8)가 모두 구현되었습니다. 단계별 인도물과 기록된 이탈사는 docs/ROADMAP.md, 우선순위가 있는 백로그는 docs/TODO.md, 엔지니어링 작업 로그는 docs/WORKLOG.md에서 확인할 수 있습니다.

2026-08-23 검증 기준: 워크스페이스 테스트 184개 성공, cargo clippy와 cargo fmt 깨끗, 전체 워크스페이스 빌드 완료. 스텁 크레이트는 남아 있지 않습니다 — 모든 크레이트가 실제 구현된 모습입니다. Coverage는 CI에서 측정됩니다(cargo llvm-cov; 기준 67.0.7% line, 계속 높이는 중).

테스트 깊이는 여전히 불균등합니다: 하위 계층은 잘 반영되어 있고(gpp-core, gpp-graphex, gpp-diff, gpp-tui) 80% 이상 line coverage; CLI에는 정책 강제, 비용 보고, 리뷰어 배정, belief 이분 탐색(Bé)에 대한 엔드투엔드 스위트가 있습니다), 여러 통합 크레이트는 스몰(smoke) 수준에 머물러 있습니다(gpp-sdk, gpp-notify, gpp-rbac, gpp-replay). 여기서 "구현"은 각 마일스톤을 기준으로 빌드와 테스트를 통과했음을 뜻하며, 모든 곳에서 완벽하게 완성되었음을 가리키지 않습니다. 이 사이를 좁히는 것이 docs/TODO.md의 최우선 항목입니다.

레이어 전체 표

레이어

크레이트

구현 내용

저장소

gpp-core

컴텐츠 주소 참조 저장소 (BLAKE3 + zstd), Blob / Tree, 검증된 사용 검증 후 raw 프레임 전송

연대기

npm-timeline

SQLite(WAL) 캡처, .gppignore, debounce 와쳐, 정리(pruning)

History

가프-history

Changeset / Intent / Author, branch refs, 승격(promote), DAG 순회

Diff

gpp-diff

행 단위 + tree-sitter 의미적 diff (Rust/Python/TS/Go), 이름 바꾸기/이동 감지

Git bridge

gpp-git-bridge

git-import / git-export / git-bridge, SQLite 해시 매핑

Graphex

gpp-graphex

암호화(age + AES-GCM) 지식 그래프, tier로 제한되는 투영, 질의, 라이프사이클, 감사, belief + stalenes에 대한 엔진

SDK / MCP

gpp-mcp-sdk

AgentSession; gpp mcp serve --stdio (JSON-RPC MCP)

신뢰

gpp-trust

평판 점수, 상태 전환, 오버라이드(override), 이벤트 처리

정책

gpp-policy

.policy TOML rule, 승격 시(차단) + 타임라인(경고) + sync(차단) 강제, 내장 행사 템플릿

비용

gpp-cost

변경단위 token/$ 기록, budget, 효율성, 에이전트 자기보고

비정상

gpp-anomaly

scope/burst/size 탐지, 완화 워크플로우

동기화

gpp-sync

Noise_XX P2P; objects/refs/policies/graphex; fork 보존

재생

gpp-replay

재현 가능한 환경 스냅샷 + drift diff

Review / RBAC / 조

gpp-review, gpp-rbac, gpp-notify

리뷰 라이프사이클, 역할을 정의한 브랜치 보호, 이벤트/인박스/HMAC 웹훅

원격

gpp-remote

GitHub/GitLab/Bitbucket PR 생성, 내용 보강, CI/review 반발(import)

중계

gpp-relay

Always-on sync 허브 바이너리 + 헬스 엔드포인트 + Dockerfile

클라이언트

gpp-cli gpp-tui gpp-deps

full CLI, ratatui TUI(gpp ui), dependency 인텔(gpp deps + OSV)

기타: extensions/gh-gpp, vscode-gpp, neovim-gpp, GitHub Actions + GitLab CI templates, deploy/ Docker 이미지, packaging/ Homebrew 패키지.

문서화된 후속 계획(ROADMAP/TODO에 기록하며, 숨기지 않고 책): 의존성의 registry/license API, PyO3/napi 기본 바인딩, 외부로 나가는 리뷰 동기화, apt/dpkg 패키지.

설치 / 빌드

cargo install gpp-cli                    # the `gpp` binary (crates.io)
cargo install gpp-relay                  # relay node (optional)

# or from a clone
cargo build --release
cargo test --workspace
cargo bench -p gpp-core -p gpp-diff      # criterion perf suite

Linux, macOS(ARM + Intel) 및 Windows용 사전 빌드 바이너는 각 릴리스에 포함되어 있습니다; cargo install --git https://github.com/mahabubul470/gpp gpp-cli는 아직 릴리스되지 않은 개발 버전을 따라갑니다.

더 시도할 것들

# Governance
gpp policy template secrets-scan
gpp trust show
gpp audit --include-cost --include-graphex

# Decentralized: sync two repos over Noise
gpp sync serve 127.0.0.1:9473             # on peer A
gpp sync add a 127.0.0.1:9473 && gpp sync # on peer B

# GitHub-compatible
gpp remote setup --platform github --repository acme/webapp
gpp remote pr-create --base main

프로젝트 컨텍스트는 CLAUDE.md, 상세 스펙(architecture, data model, CLI, protocols, roadmap)은 docs/, 사용자 안내서 및 튜토리얼은 docs/book/ 을 확인하세요.

라이선스

MIT

Available Tools

9 tools
graphex_conventionsB

List applicable coding conventions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only says 'List applicable coding conventions' with no indication of the output format, whether it returns a list, or any side effects (none expected). It does not say what 'applicable' means or what constitutes a convention. This is minimal and not transparent beyond the action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that is front-loaded with the action and object. There is no fluff, and it is appropriately sized for a tool that takes no parameters. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless tool with no output schema, the description is minimally adequate. It tells the agent what it does. However, it leaves open questions about the nature of the output (e.g., plain text vs. structured list) and what 'applicable' means in the current context. Given that there is no output schema or annotations, a bit more detail—such as 'returns a list of convention identifiers'—would improve completeness. Still, it is not severely lacking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters and the schema is empty (100% coverage trivially). According to the scoring guidelines, a baseline of 4 is appropriate for 0-parameter tools. The description does not need to explain parameters because none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and a clear resource ('applicable coding conventions'). It is unambiguous and does not conflate with siblings like graphex_query (query) or graphex_glossary (glossary). However, it does not explicitly differentiate from them, though the resource itself is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It merely states the action without context such as 'use this when you need the list before proposing changes' or any exclusions. An agent would have to infer when it is applicable, which is a gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graphex_glossaryC

Look up domain glossary terms.

ParametersJSON Schema
NameRequiredDescriptionDefault
termNo

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full responsibility for behavioral disclosure. It only states the action with zero additional context about side effects, read-only nature, error behavior, or output characteristics, failing to inform the agent beyond the basic verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the core action without any extraneous words. Every character earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks essential details such as return format, behavior when the optional parameter is omitted, or any constraints. Given the absence of an output schema and the sparse parameter info, the description is inadequate for an agent to confidently invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one optional parameter 'term' with no description, and the tool description does not explain its meaning or behavior (e.g., what happens when omitted). With 0% schema description coverage, the description must compensate but does not, leaving parameter semantics entirely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'look up' and a clear resource 'domain glossary terms', making the tool's purpose unambiguous. It is readily distinguishable from sibling tools like graphex_query or graphex_conventions without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, and no exclusions or prerequisites are mentioned. While the purpose implies its use, there is no explicit context for selection, leaving agents to infer conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graphex_queryD

Project knowledge-graph context (tier-filtered).

ParametersJSON Schema
NameRequiredDescriptionDefault
budgetNo
patternNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It doesn't mention whether the operation is read-only, what kind of output to expect, or any side effects. 'Tier-filtered' hints at a behavior but leaves it unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this is under-specification rather than conciseness. A single vague phrase conveys almost no actionable information and does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a query tool with two optional parameters and no output schema, this description provides nowhere near enough context to invoke it correctly. It fails to explain the purpose, parameters, or expected behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema defines two parameters (budget and pattern) with 0% description coverage, and the tool description says nothing about them. Agents are left with no idea what these parameters control or how to use them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Project knowledge-graph context' but doesn't specify the action or resource. 'Tier-filtered' is ambiguous, and there's no differentiation from siblings like graphex_glossary or graphex_conventions. An agent cannot tell what this tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any of the eight siblings. No context, no exclusions, no alternatives. The description offers zero help in selecting the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graphex_statusA

Graph statistics: node/edge counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. It mentions 'statistics' but does not explicitly state that it is a read-only operation, does not describe output format, potential errors, or any side effects. The description is too thin to give an agent confidence about what to expect beyond the basic counts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is minimal—a single colon-separated phrase—with no wasted words. It is front-loaded with the core idea and is appropriately sized for a tool with no parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no annotations, no output schema), the description is nearly complete. It tells an agent what the tool does and what it returns conceptually (node/edge counts). However, it omits details like whether the counts reflect the entire graph or a current state, and it does not mention if there is any filtering or scope. Still, for a trivial status tool, the coverage is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the baseline is 4 per the rubric. The schema is empty, and the description adds no parameter information, but none is needed. The description adequately communicates that no inputs are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Graph statistics' with the detail 'node/edge counts.' It clearly identifies what the tool does and is distinct from sibling tools like graphex_query or graphex_glossary, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings. The description does not mention alternatives, prerequisites, or context—it simply states what it does. For a tool that provides stats, an agent might need to know when to call it instead of a query, but no such direction is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_beliefB

Record an evidence-anchored belief about the code (lands as Proposed for human approval; staleness-checked against history from the moment it exists). evidence entries are "path:start-end" (1-based, inclusive lines at the current changeset); paths are repo-relative paths or globs; symbols are "path:Name". At least one of evidence/paths/symbols is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNo
claimYes
pathsNo
symbolsNo
evidenceNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose meaningful behavior: the record lands as 'Proposed for human approval' and is 'staleness-checked against history from the moment it exists.' This adds real value. However, it does not state the return/response shape, failure behavior, or any side effects beyond the proposal status — a notable gap for a write operation with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with the core purpose front-loaded, followed by format specs and the constraint. The format specifications are packed tightly and each sentence earns its place. Slightly heavy on the type-notation details, but nothing is wasted — appropriately concise for the information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five parameters, no annotations, and no output schema, the description covers the critical input semantics and constraint well but leaves gaps: 'claim' and 'tier' are unexplained, the response/proposal-approval flow is only hinted at, and the 'staleness' mechanism is mentioned without elaboration. It is workable for a competent agent but not fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it largely does. It explains evidence format ('path:start-end', 1-based inclusive lines), paths ('repo-relative paths or globs'), symbols ('path:Name'), and the mutual-requirement constraint. This adds substantial meaning beyond the bare schema. However, the 'claim' and 'tier' parameters receive no explanation, which keeps this from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Record an evidence-anchored belief about the code') and adds a distinguishing detail: the belief 'lands as Proposed for human approval'. This differentiates it from siblings like reaffirm_belief (which presumably updates an existing belief) and propose_changeset, but it does not explicitly name the sibling it is not, so a small inference gap remains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit when-to-use vs when-not-to guidance and names no alternative tool. While 'At least one of evidence/paths/symbols is required' clarifies a precondition, it never routes the agent to reaffirm_belief for updating existing beliefs or explains when proposal vs. direct record is appropriate. The contrast with reaffirm_belief is left entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_changesetC

Promote pending timeline entries into a changeset.

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNo
messageYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. The verb 'promote' implies a state-changing operation, but the description never states what actually happens to the timeline entries (are they consumed, deleted, marked?), whether changes are reversible, whether authorization is required, or what the tool returns. Consequences of the mutation are entirely undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is grammatically efficient and front-loads the core action with no wasted words. However, it is under-specified: for a tool with an undocumented required parameter, this brevity crosses from conciseness into incompleteness. It reads cleanly but does too little documentation work.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is the only documentation for a state-changing tool. It fails to explain the `message`/`intent` parameters, the side effects of promotion, or what a changeset is. Notably the JSON response wasn't part of this review, but without a schema defining the return, an agent cannot know what to expect after invoking it. Inadequate for a mutation with undocumented parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate at all — it never mentions `message` (required) or `intent` (optional). The agent has no idea what these values should contain (e.g., is `message` a commit note? does `intent` describe the change's purpose?). With zero coverage and zero description help, both parameters are effectively opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Promote pending timeline entries into a changeset.' The verb+resource is clear, but it does not distinguish itself from its 'propose_' siblings (propose_graph_update, propose_belief). Also relies on undefined domain jargon ('timeline entries', 'changeset') which assumes context. Clear but no sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus propose_graph_update or propose_belief. There is no stated condition, prerequisite, or when-not-to-use guidance. The reader can only infer usage from the name, which is not sufficient routing information among three similarly-named siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_graph_updateC

Propose a new graph node (lands as Proposed for human approval).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
tierNo
node_typeYes
descriptionYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It discloses that the proposal lands for human approval (a state-change), which is useful. However, it does not explain what happens after proposal, whether it can be undone, permissions required, or what the response looks like. For a mutating tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no fluff. It front-loads the action and outcome. However, it is almost too brief—it sacrifices needed behavioral context for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, no annotations, and no parameter documentation, this description is insufficient for an agent to call the tool correctly. The tool involves proposing a node with multiple parameters, yet only the high-level purpose is stated. The agent would need to infer parameter semantics from names alone, which is risky. Sibling differentiation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions none of the four parameters (name, tier, node_type, description). The description gives no meaning beyond what the schema types imply (strings). With zero coverage and no parameter details, an agent is left guessing about required values, constraints, or relationships.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (Propose) and resource (graph node), with a distinctive outcome (lands as Proposed for human approval). It does not explicitly differentiate from sibling tools like propose_changeset or propose_belief, though the phrase 'graph node' hints at a different domain. It is not a tautology but could be more specific about what 'graph node' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives. It implies a workflow of proposing for human approval, but does not mention exclusions or alternatives. Sibling tools like propose_changeset, propose_belief, and reaffirm_belief exist in the same namespace, and without any usage guidance an agent might select the wrong one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reaffirm_beliefA

Re-anchor a belief at the current changeset after you have re-verified it against the code (e.g. a stale-candidate whose claim still holds). Optional new evidence spans replace the old ones. Not allowed for invalidated beliefs — their grounds are gone; propose a new belief with current evidence instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
beliefYesbelief id or exact claim
evidenceNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that it re-anchors at the current changeset, that new evidence replaces old ones, and that it rejects invalidated beliefs. It doesn't state side effects like whether it creates a new changeset or overwrites prior anchors, but the core behavior is adequately transparent for a belief-management operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The purpose is front-loaded, followed by a clarifying example and a crisp restriction. Every sentence carries essential information, making it efficient and easily parseable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two simple parameters, no output schema, and no nested objects, the description covers the essential aspects: action, target, optional behavior, and an explicit exclusion. It doesn't address edge cases like invalid belief IDs or failure modes, but those are not critical given the tool's simplicity and the presence of sibling tools for broader context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% (belief has a schema description, evidence does not). The description adds crucial meaning by explaining that 'evidence' is optional and that new evidence spans replace old ones, compensating for the schema's lack of description on that parameter. The belief parameter is reinforced by the overall context, though it doesn't add new syntax-level detail beyond 'id or exact claim'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Re-anchor a belief') on a specific resource ('belief') and gives a concrete use case (stale-candidate whose claim still holds). It also contrasts with an alternative ('propose a new belief... instead'), which distinguishes it from siblings like propose_belief without needing to inspect schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('after you have re-verified it against the code'), when not allowed ('Not allowed for invalidated beliefs'), and points to the alternative ('propose a new belief with current evidence instead'). This gives clear, actionable guidance for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_costA

Attribute your token/compute usage to a changeset (the id returned by propose_changeset). Reports accumulate. cost_microdollars is integer micro-dollars (1 = $0.000001).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
changesetYes
input_tokensNo
cached_tokensNo
output_tokensNo
cost_microdollarsNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that reports accumulate (additive behavior) and clarifies the unit of cost_microdollars. However, it does not mention whether the operation is a write or if it has side effects beyond attribution, nor does it describe error behavior or reversibility. It adds some context but not comprehensive transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the purpose, and includes crucial unit clarification. No extraneous words. It could be more structured (e.g., separating parameter notes), but it is efficient and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple reporting tool, the description provides key facts: the source of the changeset id and the accumulation behavior. However, with no annotations and no output schema, it omits details like what happens on success/failure, whether all token fields are required together, or how to handle partial reports. The absence of any explanation of the remaining numeric fields leaves gaps for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all six parameters. It only explains changeset (source of value) and cost_microdollars (unit). The other parameters (model, input_tokens, cached_tokens, output_tokens) are left to their names, which are not fully self-explanatory (e.g., what counts as cached_tokens?). This falls short of adequate compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Attribute') and a clear resource ('token/compute usage to a changeset'), and precisely identifies the source of the changeset id ('returned by propose_changeset'). This leaves no doubt what the tool does and distinguishes it from sibling tools like propose_changeset or propose_graph_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear prerequisite: the changeset must be the id returned by propose_changeset, which tells the agent when this tool is appropriate. It also states that reports accumulate, implying multiple calls are allowed. However, it does not explicitly list when NOT to use it or mention alternative tools for cost tracking, so it's slightly short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observedgraphex_conventions
    • First observedgraphex_glossary
    • First observedgraphex_query
    • First observedgraphex_status
    • First observedpropose_belief
    • First observedpropose_changeset
    • First observedpropose_graph_update
    • First observedreaffirm_belief
    • First observedreport_cost

TDQS

B3/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a distinct purpose: query, status, glossary, conventions, and then actions for proposing changesets, graph updates, beliefs, reaffirming beliefs, and cost reporting. No two tools overlap in functionality; even the related belief tools are clearly differentiated (new vs. reaffirm).

Naming Consistency4/5

The tools are grouped semantically: 'graphex_' prefix for read-only graph queries, and verb-based names (propose_, reaffirm_, report_) for actions. This is consistent within each group, but there is a mix of naming styles (prefix vs. verb) across the set, making it slightly less uniform than a pure verb_noun convention.

Tool Count5/5

Nine tools is well within the ideal range and each tool serves a clear, necessary function for the server's purpose of knowledge-graph interaction and change proposal. No redundant or missing tools are apparent.

Completeness4/5

The surface covers the core workflows: querying graph context, checking status, looking up glossary/conventions, proposing changes (changeset, graph update, belief), reaffirming beliefs, and reporting cost. Minor gaps like listing existing beliefs or changesets are not directly present, but the design intentionally routes proposals to human approval, so those may be handled externally.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Local-first memory layer for AI coding agents — captures issues, attempts, fixes, and decisions, and warns at git commit before you repeat a mistake.
    17
    850
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Local-first deterministic project memory for AI coding agents, with context packs, decisions, gates, risks, scoped claims and explicit checkpoints in project-owned files.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Gives AI coding agents persistent, branch-aware memory and a dependency-tracked task graph by storing decisions, lessons, and tasks as plain JSON and Markdown committed directly into the repository. Agents can record and fuzzy-search past decisions, dump instant project context, and create, claim, complete, and query tasks whose completion automatically unblocks downstream work.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides AI coding agents with git-native persistent memory and a dependency-aware task graph, letting them record and fuzzy-recall architectural decisions, lessons, and gotchas while creating, claiming, and completing tasks that auto-unblock downstream work. Stores everything as plain JSON and Markdown committed inside the repository, so context stays branch-aware, team-shared, and reviewable in pull requests.
    9 npm
    MIT