viberevert
VibeRevert는 AI 코딩 세션을 기록하고, 결정적 규칙으로 위험한 변경을 표시하며, 커밋하지 않은 작업까지 포함해 프로젝트 파일을 세션이 시작되기 전 상태로 정확히 복원할 수 있습니다.
AI가 AI를 심판하지 않습니다: 리스크 판단 결과는 재현 가능하고 설명 가능하며, 직접 살펴볼 수 있는 규칙에 기반합니다.
사용하던 코딩 에이전트를 계속 쓰세요. VibeRevert는 Claude Code, Cursor 및 다른 코딩 에이전트 워크플로와 함께 세션 주변에 안전 계층을 추가합니다.
작동하는 모습
실제 CLI를 베타 결제 실행 결과에서 축약했습니다. check가 위험을 표시하고, rollback이 --apply 전에 변경 내용을 미리 보여줍니다.
$ viberevert check
risk: CRITICAL · payments
app/api/checkout/route.ts payments (critical)
app/api/webhooks/stripe/route.ts payments (critical)
$ viberevert rollback <session> # preview — nothing changed yet
package.json tracked_restored
app/page.tsx tracked_restored
app/api/checkout/route.ts untracked_deleted
app/api/webhooks/stripe/route.ts untracked_deleted
lib/stripe.ts untracked_deleted
.gitignore README.md (your uncommitted work) skipped_unchangedRelated MCP server: uacos
보호 약속이 아니라 증명
실제 AI 코딩 세션 3개. 프로젝트 파일 복원 3건, 정확하게. 결제 처리, 데이터베이스 마이그레이션, 배포 인프라 — 모두 커밋하지 않은 기존 작업을 보존한 채로 이루어졌습니다. 베타 리포트 읽어보기 →
빠른 시작
Git 저장소 안에서 동작하며 Node.js 22+가 필요합니다. VibeRevert는 기록을 사용자 PC에 보관하며 VibeRevert 계정이나 호스팅 서비스가 필요하지 않습니다.
npm install -g viberevert@betaviberevert init # one-time setup in your project
viberevert run <your-agent> # records the session and saves your files' starting state
viberevert check # see what changed and what looks risky
viberevert rollback <session> # preview putting your files back; add --apply to restore무엇으로부터 보호해 주나요
AI 에이전트가 프로젝트 전반에서 작업하도록 했는데, 예상보다 더 많은 곳을 건드렸을 수 있습니다. — 결제 코드, 데이터베이스 마이그레이션, 배포 파일 — 사용자가 만들어 놓은 미완성 작업은 커밋되지 않은 채 그대로 있었을 것입니다. VibeRevert는 먼저 프로젝트의 시작 파일 상태를 캡처하기 때문에, 만족스럽지 못한 세션이 여러분의 오후를 앗아가지 않습니다.
VibeRevert의 방식
VibeRevert는 각 AI 코딩 세션에 대해 다음과 같이 동작합니다.
프로젝트를 먼저 체크포인트로 저장 — 커밋하지 않은 작업을 포함한 작업 트리, 스테이징된 변경, 추적되지 않는 파일까지.
세션을 기록하고, 세션이 실행되는 동안 변경된 프로젝트 파일을 남깁니다.
위험한 수정사항을 표시 — 인증, 결제, 데이터베이스, 비밀 정보, 의존성 또는 인프라를 건드리는 변경(리스크 분류).
이 결과를 그 다음 반복에서 바로 사용할 수 있는 에이전트용 수정 프롬프트로 정리합니다.
필요할 때 파일을 체크포인트 상태로 복원합니다.
알려줄 뿐, 조용히 작업을 막지는 않습니다. 선택적으로 적용하는 pre-commit 훅이 설정한 위험 임계값을 넘는 커밋을 거부할 수도 있습니다. 다른 로컬 Git 훅처럼 선택 적용, 조절, 우회가 가능합니다.
규칙이 판단하고, 에이전트가 수정합니다
VibeRevert는 다른 언어 모델에게 에이전트의 변경 사항이 위험한지 묻지 않습니다. 판정은 결정적입니다. 동일한 입력과 동일한 설정은 항상 동일한 결과를 만들며, 모든 표시에는 검토 가능한 근거가 있습니다.
주의가 필요한 부분이 있으면, VibeRevert는 그 결과물을 에이전트가 바로 쓸 수 있는 수정 프롬프트로 전환합니다. 무엇을 표시할지는 규칙이 정하고, 나중에 어떻게 할지는 여러분이 정합니다.
롤백이 복원하는 것과 그렇지 않은 것
viberevert rollback은 커밋하지 않은 작업을 포함한 로컬 프로젝트 파일(추적 중, 후보 스테이지, 비추적 파일)을 세션이 시작됐을 때의 상태로 복원합니다. 기본적으로 먼저 미리보기를 보여 주며, --apply는 변경 사항을 적용하기 전에 응급 체크포인트를 저장합니다.
프로젝트 파일 밖의 효과는 되돌리지 않습니다. 배포, 데이터베이스 쓰기, 제3사 API 호출, 결제, 발송된 이메일은 자체적으로 복구해야 합니다. 롤백은 원자적이라기보다 상태 기반이며, AI 세션의 안전망이지 테스트나 리뷰를 대체하지 않습니다. 신뢰하고 사용하기 전에 롤백이 복원할 수 있는 것과 없는 것을 읽어 보세요.
다양한 코딩 에이전트에 빠는 하나의 복구 계층
선호하는 코딩 도구를 계속 사용하세요. VibeRevert는 CLI 에이전트 세션을 비교 범위로 감쌀 수 있고(viberevert run <your-agent>), 설치 프로그램을 통해 Claude Code 및 Cursor 같은 도구와 통합할 수 있습니다.
내장 체크포인트온 에이전트 워크플로우 안에 살아 있습니다. VibeRevert는 프로젝트에 남아 있습니다.
프로젝트 시작 상태(커밋하지 않은 작업 포함)를 잡고, 세션을 기록하며, 위험 변경을 표시하고, 롤백을 미리 보여 주며 이후에도 그 복구 경로는 에이전트 바깥에 보존됩니다.
오늘은 Claude Code를, 내일은 Cursor를 쓰거나 다른 터미널 에이전트를 감쌀 수도 있습니다: 안전 계층은 프로젝트와 함께합니다.
도구에 연결하기
VibeRever트는 비파괴적으로 설치되며 깨끗하게 제거할 수 있습니다. Claude Code and Cursor(MCP server 및 hooks)와 통합되며, 모든 설치 과정은 변경 사항을 미리 보여 줘고 복구 저널을 유지합니다. 시작하기를 참조하세요.
실험적인 터미널 브리지(\--pty)는 대화형 에이전트 셸 안에서 명령을 가로챌 수 있습니다. Best-effort 방식이며, PTY 계약에 그렇게 문서화되어 있습니다.
지원 플랫폼
Linux, macOS, Windows, Node.js 22+ 대상입니다. CI는 세 플랫폼 모두에서 Node 22와 24를 대상으로 전체를 실행합니다. 각 플랫폼이 정확히 수준에서 테스트되므로 호환성을 확인하세요.
더 알아보기
VibeRevert를 후원하세요
VibeRevert는 Apache-2.0 오픈소스입니다. 후원금은 지속적인 데스크톱 플랩폼 테스트, 보안 작업, 롤백 시나리오 테스트, 독립적 검토 비용에 사용됩니다. 후원이 보안 위험 판정이나 출시 결정에 영향을 미치지 않습니다. → VibeRevert 후원하기 · 후원금의 활용
라이선스
Apache-2.0. NOTICE 및 라이선스 감사를 참조하세요.
Available Tools
8 toolscheck_repoB
Run safety checks against the working tree or staged diff and return the resulting ReportFile (or summary when large). Side-effecting (class B per D99.V): always persists the ReportFile under .viberevert/.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| since | No | ||
| staged | No | ||
| threshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly notes that the tool is side-effecting and always persists a ReportFile, which is critical behavioral information, and no annotations are provided to cover this. It also mentions that the return value may be a summary when large, adding transparency. This is helpful but does not go into deeper details like exact file paths or side effects on the environment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the main action and return type. The mention of side effects and persistence is valuable. However, one sentence might be dense, potentially combining multiple ideas, but it remains clear and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters with zero schema descriptions, no output schema, and no annotations, the description is incomplete because it does not explain parameter usage or the full behavior. However, the mention of side effects and large output handling adds some completeness haven't fully addresses the parameter gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema coverage is 0%, so no parameters are described in the schema. The description mentions 'working tree or staged diff' but does not explicitly connect these to parameters like 'since' or 'staged'. The enum for threshold is not explained, and the meaning of 'task' and 'since' remains unclear, so the description fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: run safety checks against working tree or staged diff and return a report. It also distinguishes from siblings like classify_risk or list_risky_files by mentioning the ReportFile output and the scope of checks. However, it omits any detailed scope of what the safety checks actually include.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool is described as running safety checks, which implies it should be used to assess code changes for risks, but there is no explicit guidance on when to use it versus alternatives like classify_risk or list_risky_files. The description does not mention exclusions or alternative tools, so while the context is somewhat clear, it lacks explicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classify_riskA
Project a check report's findings into per-severity counts (critical, high, medium, low). Read-only (class A per D99.V); no threshold input -- returns the full per-level histogram.
| Name | Required | Description | Default |
|---|---|---|---|
| report | No | ||
| session | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Explicitly states 'Read-only' and 'no threshold input', providing key behavioral context. Since no annotations exist, the description carries the burden and addresses safety and side-effect concerns adequately, though other traits (e.g., performance) are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two succinct sentences with no redundancy. The structure flows logically from purpose to behavioral notes, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately describes the return (per-level histogram) and mentions the lack of threshold filtering. It covers core functionality, though parameter details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% coverage for parameters 'report' and 'session'. The description hints that 'report' is a check report but does not explain 'session' or clarify expected formats, leaving significant ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: 'Project a check report's findings into per-severity counts' precisely describes the operation and output. Distinct from siblings by its focus on severity histograms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly indicates when to use (when severity counts are needed) and mentions 'no threshold input' to clarify scope, but does not explicitly contrast with alternative tools like 'list_risky_files'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_checkpointA
Create a standalone VibeRevert checkpoint capturing the current working tree state. Returns the new checkpoint id and creation timestamp. Side-effecting (class B per D99.V); writes a checkpoint directory under .viberevert/checkpoints/. When name is supplied, uses the D22 name-uniqueness lock; nameless checkpoints skip the lock entirely.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden of disclosing side effects. It explicitly states the operation is side-effecting (class B), writes under '.viberevert/checkpoints/', and explains the name-uniqueness lock behavior. It also discloses the return values (id and timestamp). This is a high level of transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, then quickly covers returns, side effects, and conditional behavior. Every word earns its place; there is no redundancy or filler. The structure is logical: purpose → output → side effect → conditional nuance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a single optional parameter and no output schema, the description is complete. It covers the primary action, the return format (id and timestamp), the location of side effects (checkpoint directory), and the lock behavior under a specific condition. No critical information is missing, such as error cases or prerequisites, which are not needed given the simplicity of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no description for the 'name' parameter (0% coverage), but the description compensates by explaining that supplying 'name' triggers the D22 name-uniqueness lock, implying that 'name' is the checkpoint name and affects uniqueness. This adds meaningful semantics beyond the schema's pattern and maxLength constraints, though it could be more explicit about the checkpoint naming purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Create'), the resource ('a standalone VibeRevert checkpoint'), and the specific action ('capturing the current working tree state'). It also mentions the returned values (checkpoint id and creation timestamp), making the purpose unambiguous. The term 'standalone' differentiates it from potential sibling operations like start_session or check_repo, though none are directly similar.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on usage: it's a side-effecting operation (class B) that writes a checkpoint directory, and it details the conditional lock behavior when 'name' is supplied. While it doesn't explicitly state when to use it versus alternatives, the sibling tools are distinct and no direct alternative exists. The lock condition effectively guides parameter usage (when to supply 'name' and when not).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_diffA
Render a CommonMark explanation of a check report alongside the report's metadata. Reads an existing ReportFile (read-only, class A per D99.V); does not run new checks.
| Name | Required | Description | Default |
|---|---|---|---|
| report | No | ||
| session | No | ||
| threshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It explicitly states the tool is read-only and does not run new checks, which covers safety and side effects. It also references a class A security classification (per D99.V), adding context. However, it does not describe potential errors, return format details, or dependencies beyond reading a report.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences that efficiently convey the core function and key behavioral caveats. It is front-loaded with the primary purpose and contains no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% parameter coverage in the description, the tool is incomplete for an agent to use correctly. The description covers the primary action and read-only nature but lacks parameter semantics and output expectations. For a simple read-only tool, it might suffice if parameters were self-explanatory, but they are not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides no explanation of the parameters (report, session, threshold). Schema description coverage is 0%, and the schema itself has no descriptions. The tool requires three parameters with no guidance on their meaning or format, making correct invocation difficult.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Render') and resource ('CommonMark explanation of a check report'), clearly indicating the tool's function. It also distinguishes itself from siblings by stating it does not run new checks, differentiating it from check_repo and similar tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by stating it reads an existing ReportFile and is read-only, which suggests it should be used after checks have been run. However, it does not explicitly name alternative tools or provide a detailed when-to-use vs. when-not-to-use, but the 'does not run new checks' caveat gives clear guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_fix_promptA
Render a fix-prompt from a VibeRevert check report. Returns the rendered prompt text, the absolute path to the persisted sibling fix-prompt.txt, and the source report id. Side-effecting (class B per D99.V); writes the sibling fix-prompt.txt via atomic temp+rename. The on-disk sibling is always the full untruncated text (D81); the MCP wire response's prompt_text is capped at 262144 bytes (D99.U) -- when truncated, the response carries truncated:true + bytes_omitted.
| Name | Required | Description | Default |
|---|---|---|---|
| report | No | ||
| session | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly discloses side effects: 'Side-effecting (class B per D99.V); writes the sibling fix-prompt.txt via atomic temp+rename.' It also transparently explains output truncation behavior: 'the MCP wire response's prompt_text is capped at 262144 bytes (D99.U) -- when truncated, the response carries truncated:true + bytes_omitted.' This goes beyond typical annotations and fully informs about side effects and output limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, starting with the main purpose, then the return values, then side effects, and finally truncation details. Each clause adds necessary information, and while it is somewhat long, it avoids redundancy and is appropriately sized for the complexity of the behavior it describes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers behavioral aspects thoroughly (side effects, truncation) and mentions return values, but it lacks explanation of the input parameters and when to use the tool. Given the simple schema (two string fields) and no annotations, the description is moderately complete but leaves gaps in parameter semantics and usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not relate the input parameters (`report` and `session`) to their meaning or usage. It mentions 'from a VibeRevert check report' but never explicitly states that the `report` parameter is the report ID or that `session` refers to a session. This leaves the schema fields completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Render a fix-prompt from a VibeRevert check report.' It uses a specific verb ('Render') and a specific resource ('fix-prompt'), and distinguishes itself from sibling tools by focusing on fix-prompt generation from check reports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives. It lacks context or prerequisites, such as 'use when a check report is available' or comparisons to other tools like create_checkpoint. The only usage hint is implicit in the tool's purpose, but no explicit direction is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_policyA
Return the project's resolved policy slice from .viberevert.yml: block/warn risk thresholds, check toggles, frameworks, and rollback exclude patterns. Applies M C defaults (D57) before returning. Read-only (class A per D99.V); does not run checks or mutate state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and handles it well: it explicitly declares 'Read-only (class A per D99.V); does not run checks or mutate state' and discloses the defaults-application behavior ('Applies M C defaults (D57) before returning'). This adds safety and resolution context beyond any structured annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste: first sentence states the verb+resource+contents, second explains the resolution behavior, third covers the safety profile. Each sentence earns its place; the description is front-loaded with purpose and contains no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, read-only tool with no output schema, the description is nearly complete: it lists the return content, the defaults-resolution behavior, and the safety profile. The only minor gap is that it doesn't describe the exact output structure/format of the returned policy slice, but for this simple tool the component list suffices.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the rubric baseline is 4. The description compensates by describing the return-value content (the policy components), which gives the agent a clear picture of what it will receive even without an output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource ('Return the project's resolved policy slice from .viberevert.yml') and enumerates the exact contents (risk thresholds, check toggles, frameworks, rollback exclude patterns). It clearly distinguishes from siblings — none of create_checkpoint, check_repo, explain_diff, classify_risk, or list_risky_files describe retrieving a resolved policy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context — it's a policy-lookup tool that's safe because it doesn't run checks or mutate state — but it never explicitly states when to choose this vs alternatives, and no sibling tool is named for comparison. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_risky_filesA
Project a check report into a per-file risky-files list (path, max_severity, finding_count) sorted by [max_severity DESC, finding_count DESC, path ASC]. Read-only (class A per D99.V). Capped at 500 entries per D99.U.
| Name | Required | Description | Default |
|---|---|---|---|
| report | No | ||
| session | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosure. It explicitly states the tool is read-only (class A per D99.V) and capped at 500 entries (per D99.U), which are important behavioral constraints. It also discloses the sorting behavior, adding value beyond what annotations would typically provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences, with the primary function front-loaded in the first sentence and constraints in the second. Every word adds value, no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema or annotations, the description does a good job of specifying the return format (list of fields) and sorting, but it leaves the 'session' parameter unexplained and provides no usage context relative to sibling tools. The references to D99.V and D99.U are opaque without external knowledge, creating additional ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. While 'report' is implicitly clarified as the 'check report', the 'session' parameter is never explained, leaving its purpose and relationship to the report ambiguous. This is a significant gap for an agent trying to invoke the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: projecting a check report into a per-file risky-files list with explicit output fields (path, max_severity, finding_count) and sorting order. This specific verb+resource+output format differentiates it from sibling tools like check_repo or classify_risk.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context implies the tool is used after a check report exists ('Project a check report'), but it does not explicitly state when to use this tool over alternatives, nor does it provide any when-not-to-use guidance. No alternative tools are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_sessionA
Start a new VibeRevert session: create the inner checkpoint, capture pre-session git status, and acquire the active-session lock. Returns the new session id, checkpoint id, and start timestamp. Side-effecting (class B per D99.V); writes session state and uses the D22 start lock during execution.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full weight. It does an excellent job of disclosing side effects: it states it is 'side-effecting (class B per D99.V)', writes session state, and uses the 'D22 start lock' during execution. It also specifies what resources are affected (checkpoint, git status, lock) and provides a reference to a classification system (D99.V) for risk assessment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, three sentences, and front-loads the primary action. It efficiently packs essential details: what it does, what it returns, and what side effects to expect. Every sentence adds value, from the core function to the side-effect warning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple internal steps: checkpoint, git status, lock) and lack of output schema, the description covers the key aspects. It explains the side effects, returns, and even classifies the side-effect level. It could potentially specify what happens if the lock is already held, but that's a minor omission given the details provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one parameter 'task' with only a pattern requiring non-space characters. The description doesn't explicitly define 'task', but it mentions that the session creation is the main action. Since there's only one parameter and it's not described in the schema beyond the pattern, the description could have hinted at what 'task' means. However, given the schema coverage is 0%, the description's silence on 'task' is a minor gap, but the low complexity (1 param) mitigates the impact. The description adds value by explaining the side effects but not the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: starting a VibeRevert session. It specifies the concrete actions (create inner checkpoint, capture git status, acquire lock) and the return values (session id, checkpoint id, start timestamp). While the sibling tools are listed, the description itself includes enough specificity to distinguish it from potential siblings like create_checkpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (start of a session) and notes it acquires a lock, which hints at concurrency management. It doesn't explicitly say when not to use it or name alternative tools, but the context of session initialization and the lock acquisition gives clear usage context. The mention of 'start lock' suggests a prerequisite or constraint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
- First observed
check_repo - First observed
classify_risk - First observed
create_checkpoint - First observed
explain_diff - First observed
generate_fix_prompt - First observed
get_policy - First observed
list_risky_files - First observed
start_session
TDQS
Scored across 8 tools
Most tools have clearly distinct roles (create checkpoint, run checks, configure policy, start session), and the report-consuming tools produce different outputs. However, create_checkpoint and start_session overlap in checkpoint creation, and explain_diff, classify_risk, and list_risky_files all operate on check reports, requiring careful reading to choose the right one.
All tool names follow a consistent verb_noun snake_case pattern: create_checkpoint, check_repo, explain_diff, classify_risk, list_risky_files, get_policy, start_session, generate_fix_prompt. There is no mixed casing, vague single-word verbs, or inconsistent naming style.
Eight tools is a well-scoped surface for a checkpoint and safety-reporting server. Each tool covers a recognizable stage of the workflow, and none feel like duplicates or padding.
The server can create checkpoints, start sessions, run checks, and analyze reports, but there is no end/abort session tool, no revert or rollback operation despite the server being named VibeRevert, and no way to manage checkpoints beyond creating them. These are significant gaps that leave core workflows incomplete.
Maintenance
Related MCP Connectors
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceOpen-source AI code review MCP server for local git diff auditing with deterministic security rules and AI-powered analysis using any OpenAI-compatible model.4MIT
- AlicenseNot gradedqualityBmaintenanceLocal-first code intelligence and safety layer for AI coding agents. MCP server exposes dependency graph, impact analysis, and AST-compressed repo context, backed by typed local memory, patch-scope safety gates, and git-independent transaction rollback.1MIT
- AlicenseCqualityDmaintenanceSecurity scanner and MCP server that catches dangerous patterns in MCP servers and AI agent projects, such as leaked secrets, shell execution, and prompt-injection text. Runs as both a CLI and MCP server with CI-friendly severity gates.21MIT
- AlicenseNot gradedqualityAmaintenanceAI code reviews and git activity digests with machine-readable risk scoring, available as an MCP server for use within an agent session.1MIT