Skip to main content
Glama

VIBE ENGINEERING MCP

An Agentic Software-Engineering Operating System & Control Plane that manages requirements, architecture, tasks, workspaces, execution, verification, integration, and checkpoints for autonomous software development.

CI License TypeScript MCP SDK


πŸ’‘ Core Philosophy

  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                 WORKER INTELLIGENCE (LLM)                    β”‚
  β”‚        (Claude Code, Cursor, Gemini, GPT, Custom Agent)      β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚ JSON-RPC / MCP Protocol
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚              ENGINEERING CONTROL PLANE (MCP)                 β”‚
  β”‚                     Vibe Engineering MCP                     β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚                           β”‚                           β”‚
β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Task DAG Scheduler   β”‚   β”‚  Completion Gate     β”‚   β”‚ Checkpoint & Resumptionβ”‚
β”‚ Cycle Detection      β”‚   β”‚  Observable Harness  β”‚   β”‚ Git Commit Snapshots   β”‚
β”‚ Topological Ordering β”‚   β”‚  Contract Verifier   β”‚   β”‚ Worker Continuation    β”‚
β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
     β”‚                           β”‚                           β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚              ISOLATION & EXECUTION POLICIES                  β”‚
  β”‚   Git Worktrees  Β·  Command Policy  Β·  Path Policy  Β· Observer β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                   DUAL-LAYER STATE STORAGE                   β”‚
  β”‚       SQLite (ACID Operational DB)  Β·  .vibe/ (YAML Specs)   β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • The LLM is NOT the source of truth for project state.

  • The MCP server is the engineering operating system.

  • A task cannot become VERIFIED merely because an LLM claims it is finished.

  • A capability is complete only when: $$\text{Implementation} + \text{Runnable Harness} + \text{Observable Output} + \text{Contract Verification} + \text{Integration Compatibility} + \text{Checkpoint}$$ have all passed with reproducible evidence.


Related MCP server: Orchestrator Python MCP Server

πŸ”„ Task State Machine & Completion Gate

Tasks progress through strict state machine transitions:

$$\text{BLOCKED} \longrightarrow \text{READY} \longrightarrow \text{CLAIMED} \longrightarrow \text{BUILDING} \longrightarrow \text{VERIFYING} \longrightarrow \text{VERIFIED} \longrightarrow \text{INTEGRATING} \longrightarrow \text{VERIFIED}$$

The Completion Gate (task_complete)

When task_complete is invoked by a worker:

  1. Inspects the task and active workspace.

  2. Inspects Git working tree and modified files.

  3. Inspects Git diff.

  4. Executes unit and component tests.

  5. Runs the module's registered Observable Harness (CLI, HTTP, DB, AI, UI).

  6. Captures standard output, standard error, exit code, and execution time.

  7. Evaluates contract rules (exit codes, regex patterns, JSON schema verification, artifact existence).

  8. Checks cross-module integration requirements.

  9. On Pass: Transitions task to VERIFIED, updates DAG readiness for dependent tasks, and records state.

  10. On Fail: Rejects completion, transitions task to FAILED, and automatically schedules an Automated Repair Task in state READY populated with:

    • Original task ID

    • Failure reason

    • Failing command

    • Relevant stdout/stderr

    • Affected files

    • Suggested remediation action


πŸ›‘οΈ Git Worktree Workspace Isolation

Unlike naive implementations that share or pollute the project root directory, Vibe Engineering MCP provisions dedicated Git worktrees for each worker/task:

  • Worktrees are created under .worktrees/ws-<workerId>-<taskId> on isolated Git branches (worker/<workerId>/task-<taskId>).

  • The system never silently falls back to the project root. If worktree provisioning fails, it throws an explicit diagnostic error.

  • Worktree metadata is fully tracked (workspaceId, taskId, branch, worktreePath, repositoryPath, status, createdAt).

  • Worktrees can be safely cleaned up with cleanupWorkspace (git worktree remove --force).


πŸ”’ Security Model

Command Policy (src/security/command-policy.ts)

  • Scopes terminal command execution to the project or designated worktree.

  • Blocks destructive commands (rm -rf /, formatting disks via mkfs, overwrite via dd, fork bombs :(){ :|:& };:, unauthorized sudo).

  • Requires execution inside valid workspace boundaries.

  • Every command execution records: execution ID, command, cwd, timestamp, exit code, stdout, stderr, and duration to SQLite.

Path Policy (src/security/path-policy.ts)

  • Enforces strict containment within project/worktree root.

  • Resolves symlinks and prevents directory traversal attacks (e.g. ../../etc/passwd).


πŸ› οΈ Registered MCP Tools (28 Tools)

Migrated to the current MCP TypeScript SDK V2 (@modelcontextprotocol/server):

Category

Tool

Description

PROJECT

project_initialize

Initializes or connects to a project with SQLite state & .vibe/ directory

project_inspect

Full operational snapshot (DAG counts, git status, active workers)

project_resume

High-density context briefing for cold model resumption

REQUIREMENTS

requirements_compile

Compiles human specifications into structured capabilities & criteria

ARCHITECTURE

architecture_generate

Records modular boundaries, entrypoints, and contracts

architecture_validate

Validates DAG acyclicity and interface consistency

TASKS

task_create

Creates a new task with dependencies and observable contract specification

task_next

Returns next highest priority READY task (prioritizes repair tasks)

task_claim

Assigns a READY task to a worker model

task_complete

Evaluates task through Completion Gate with auto-repair synthesis

task_fail

Explicitly marks a task as FAILED with diagnostic reason

WORKSPACE

workspace_create

Creates isolated Git worktree and branch for a worker

workspace_status

Inspects all active worker workspaces

REPOSITORY

repo_read_file

Safe file read bounded by workspace containment

repo_write_file

Safe atomic file write bounded by workspace root

repo_delete_file

Safe workspace file deletion

TERMINAL

terminal_exec

Policy-checked scoped terminal command execution

GIT

git_status

Returns working tree status, staged, modified, and untracked files

git_diff

Returns working tree or cached diff

git_commit

Stages changes and creates git commit

git_branch

Creates and checks out branch

git_merge

Merges branch with --no-ff

MODULES

module_run

Executes module runnable harness and captures metrics

module_observe

Records reproducible observable evidence

VERIFICATION

contract_verify

Evaluates evidence against contract rules

integration_verify

Executes cross-module integration test harness

CHECKPOINTS

checkpoint_create

Records comprehensive snapshot (SHA, capabilities, DAG, decisions)

checkpoint_load

Loads historic milestone checkpoint


πŸš€ Getting Started

Prerequisites

  • Node.js: v20.0.0 or higher (Tested on Node 22 with built-in node:sqlite)

  • Git: 2.30+ with worktree support

Installation

git clone https://github.com/prabhurajvardhan/Dev-MCP.git
cd Dev-MCP
npm ci

Typecheck & Test Suite

npm run typecheck
npm test

The test suite includes:

  • 4 Unit test suites: Task DAG scheduler, SQLite state store, command security policy, contract engine.

  • 4 Integration test suites: Real MCP client protocol over stdio (@modelcontextprotocol/client), 19-step autonomous engineering loop, task completion gate & repair generation, MCP server tools interface.

Build Project

npm run build

Compiles TypeScript into ./dist.


πŸ”Œ Running the MCP Server (Stdio)

The MCP server uses pure stdio JSON-RPC transport (serveStdio).

  • Standard output (stdout) is strictly reserved for JSON-RPC messages.

  • All logging and diagnostic outputs are routed to stderr.

npm run mcp
# or
npx tsx src/server/index.ts

Configuration in Claude Code / Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "vibe-engineering": {
      "command": "npx",
      "args": ["-y", "tsx", "/path/to/Dev-MCP/src/server/index.ts"]
    }
  }
}

Configuration in Cursor

In Cursor Settings $\to$ Features $\to$ MCP Servers:

  • Name: vibe-engineering

  • Type: command

  • Command: npx -y tsx /path/to/Dev-MCP/src/server/index.ts


πŸ“‚ .vibe/ Inspectable Project State

In addition to SQLite operational ACID storage, human-readable YAML state is maintained under .vibe/:

.vibe/
β”œβ”€β”€ architecture.yaml       # System design, modular boundaries, contracts
β”œβ”€β”€ requirements.yaml       # Product goals & acceptance criteria
β”œβ”€β”€ modules.yaml            # Registered modules & runnable harnesses
β”œβ”€β”€ tasks.yaml              # Task DAG and execution status
β”œβ”€β”€ checkpoints/            # Frozen capability snapshots (latest.yaml + chk-*.yaml)
β”œβ”€β”€ contracts/              # Interface contracts
β”œβ”€β”€ observations/           # Reproducible evidence logs (evd-*.yaml)
└── decisions/              # Architectural decision records

⚠️ Current Scope & Limitations (V0)

  • Local Focus: V0 is optimized for single-machine local development using Git worktrees and SQLite.

  • No Remote Workers: Distributed worker pools or cloud message brokers are deferred to V1.

  • Synchronous Worktree Operations: Worktree allocation is local to the filesystem where the MCP server runs.


πŸ“„ License

Licensed under the Apache License, Version 2.0. See LICENSE for details.

Available Tools

28 tools
architecture_generateC

Synthesizes and records system design, modular boundaries, observable entrypoints, and contracts

ParametersJSON Schema
NameRequiredDescriptionDefault
modulesYesModules list
projectIdYesID of the project
interfacesYesInterfaces list
dependenciesYesInter-module dependency connections
systemDesignYesArchitecture overview and design principles

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It implies a mutation ('records') but says nothing about required permissions, whether it overwrites prior architecture, what happens if validation later fails, or whether the target project must already exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with no filler, front-loading the verb and resource. It packs several concepts but stays compact and readable, with nothing that fails to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, fully-required mutation tool with no annotations and no output schema, the description is minimally adequate: it names the recorded artifacts but omits workflow context, the relationship to architecture_validate, and any return or failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description maps loosely onto systemDesign, modules (boundaries/entrypoints) and interfaces (contracts) but omits the dependencies parameter entirely and adds no format or syntax guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('synthesizes and records') plus the resource and lists the concrete contents it captures (system design, modular boundaries, entrypoints, contracts). This clearly conveys what the tool does, though it does not differentiate itself from the likely sibling architecture_validate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when this tool should be used versus architecture_validate or the other project-lifecycle siblings, no prerequisites, and no exclusion conditions. The agent must infer placement in the workflow entirely on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

architecture_validateB

Validates architecture DAG consistency, cycle prevention, and interface completeness

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesID of the project to validate

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says what is validated but does not disclose whether the operation is read-only, whether it mutates state, what permissions are needed, how failures are reported, or what the return value looks like. For a validation tool with zero annotation coverage, this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It efficiently lists the core validation aspects without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a simple one-parameter tool with no annotations and no output schema, the description identifies what is validated but omits behavioral context an agent may need, such as read-only status, success/failure semantics, or how the validation result is returned. It is minimally adequate but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter 'projectId' is documented in the schema as 'ID of the project to validate.' The description adds no further parameter semantics, so the baseline of 3 is appropriate when the schema already does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Validates') and resource ('architecture DAG') along with the exact validation dimensions: consistency, cycle prevention, and interface completeness. This is much clearer than a tautology, but it does not explicitly distinguish itself from related sibling tools like architecture_generate or contract_verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and description: validate an architecture DAG. However, the description offers no explicit when-to-use guidance, no prerequisites, and no alternatives such as when to use contract_verify or integration_verify instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkpoint_createA

Creates a persistent checkpoint recording complete engineering state (git commit SHA, verified capabilities, tasks, failures, observations) for cold resumption

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory
summaryNoSummary description of this checkpoint
projectIdYesID of the project to checkpoint

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses what the checkpoint records (git SHA, capabilities, tasks, failures, observations) and that it is persistent, but does not cover permissions, overwrite behavior, or side effects beyond creation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single, front-loaded sentence with no filler, and the core purpose and payload are immediately clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with no output schema and no annotations, the description explains what is created and why, but lacks guidance on invocation timing, return behavior, or relationship to checkpoint_load and project_resume. The schema handles parameter documentation, so the remaining gap is moderate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description adds no parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Creates) and resource (persistent checkpoint), and explains the scope of what is recorded. It is clear against siblings like checkpoint_load by contrast, but does not explicitly differentiate itself from alternatives such as project_resume or checkpoint_load.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for cold resumption' implies a use case, but there is no explicit guidance on when to create a checkpoint versus loading one or resuming a project. Alternatives are not mentioned, leaving the agent to infer timing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkpoint_loadB

Loads a specific checkpoint by ID to inspect prior milestone state

ParametersJSON Schema
NameRequiredDescriptionDefault
checkpointIdYesID of the checkpoint to load

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. 'Loads ... to inspect' usefully signals an inspection/read intent, resolving some ambiguity about whether loading restores state, but it never states whether calling this mutates current state, what happens on an unknown ID, or whether it is safe/idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the verb and identifier up front and the intent trailing. It is appropriately sized for a one-parameter tool; nothing is wasted, though it is minimally informative rather than maximally compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-param tool with no output schema, the description covers the core purpose but omits side-effect disclosure and any hint at what is returned or how to obtain a valid checkpoint ID. Adequate but with clear gaps given the lack of annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single checkpointId parameter, so the schema already documents its meaning (ID of the checkpoint to load). The description adds no format, source, or example detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Loads a specific checkpoint by ID') plus the intent ('to inspect prior milestone state'), which cleanly separates it from the sibling checkpoint_create. It doesn't explicitly name or contrast with siblings, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'to inspect prior milestone state' implies the usage context, but there is no explicit when-to-use guidance, no mention of when to prefer this over checkpoint_create or project_resume, and no preconditions (e.g., whether a checkpoint ID must come from a prior create).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

contract_verifyC

Validates collected observable evidence against contract rules (exit codes, output regex, JSON schema, artifacts)

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdNoTask ID to fetch latest evidence for
contractIdYesID of the contract to verify against
evidenceIdNoID of the stored evidence to verify

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full behavioral burden but only says what is validated. It does not disclose whether verification is read-only or records a verdict, what a failing result looks like, or whether failure blocks downstream tasks β€” the most important behavior for a verify step in an orchestration pipeline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight, front-loaded sentence with the rule taxonomy in parentheses β€” no filler. It is efficient, though the parenthetical list slightly compresses rather than expands understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a verification tool with no annotations and no output schema, the definition covers the input side adequately but omits the outcome side: what the verification returns and what a negative result implies. Enough to know the tool's intent, not enough to call it confidently in a pipeline.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The description hints at the evidence selection model ('collected observable evidence') but adds no clarification about the taskId-vs-evidenceId relationship or precedence beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Validates') and resource ('observable evidence against contract rules'), and enumerates the rule categories checked (exit codes, output regex, JSON schema, artifacts). It does not distinguish itself from the similarly named sibling integration_verify, leaving that ambiguity to the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites beyond the required contractId, and never names integration_verify or architecture_validate as alternatives. It states what is checked but not when the agent should reach for this tool versus those siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_branchC

Creates and checks out a new branch

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory
branchNameYesName of the branch to create
startPointNoOptional starting commit or branch

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the key side effect that the working tree HEAD changes ('checks out'), but omits whether it fails on an existing branch, whether uncommitted changes are affected, and whether it returns an error or a new branch ref.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, efficiently front-loaded sentence with no wasted words. It is terse to the point of leaving behavior unexplained, but structurally there is nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating git operation with zero annotations and no output schema, the description is too thin: it says nothing about failure modes, preconditions, or the resulting repository state beyond the checkout. The schema covers parameters, but the behavioral contract for the mutation is largely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so branchName, cwd, and startPoint are already documented. The description mentions none of them, adding nothing beyond the schema, but the baseline of 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Creates and checks out a new branch') that clearly distinguishes it from repo-mutating siblings like git_commit, git_merge, and repo_write_file. It captures both the creation and the checkout side effect. It does not, however, contrast with any branch-adjacent alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any stated preconditions (e.g., whether a branch already exists, whether uncommitted changes block the checkout). Nothing tells the agent when-not to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_commitB

Stages changes and creates a git commit with structured message

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory
filesNoSpecific files to stage and commit (defaults to all modified/untracked files)
messageYesGit commit message

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose one non-obvious behavioral trait: it stages changes before committing (unlike a plain commit that only commits already-staged changes), which is useful context. But it omits failure modes (e.g., nothing to commit), whether the commit is local or pushed, permission requirements, and reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the core action front-loaded. No wasted words, though it is arguably under-specified rather than optimally concise for a mutation tool with no annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter tool with full schema coverage and no output schema, the description is minimally viable: it covers the core purpose and staging behavior. It is incomplete regarding failure conditions, return value (e.g., commit hash), and whether the commit is pushed, but the complete schema reduces the severity of those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (cwd, files, message) are already documented in the schema. The description adds only the vague phrase 'structured message', which does not clarify format or syntax beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Stages changes and creates a git commit'. Clear what it does and clearly distinct from read-only siblings like git_status and git_diff. However, it does not explicitly differentiate itself from other mutation siblings (e.g., git_merge) or name an alternative when committing is not the right action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no alternatives mentioned. The description only says what the tool does, leaving the agent to infer when it is appropriate versus git_status, git_diff, or other workflow tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_diffC

Returns git diff of modified files or staged changes

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory
fileNoSpecific file path to diff
stagedNoWhether to inspect staged changes (--cached)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only operation but says nothing about the return format (unified diff text vs structured), how large diffs are handled, or behavior on a clean tree/empty diff. For a tool with zero annotation coverage this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is appropriately sized, though it is short enough that the omitted usage and behavioral context is a content gap rather than a verbosity problem.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A simple three-parameter read tool with fully documented schema and no output schema, so not much is required. However, with no annotations and no return-value description, an agent gets no signal about the diff output shape or edge-case behavior, leaving it minimally viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents cwd, file, and staged. The description's phrase 'modified files or staged changes' loosely maps to the file and staged parameters but adds no format or syntax detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns git diff of modified files or staged changes'), which is clear enough to distinguish it from git_status and git_commit. It stops short of naming any sibling or clarifying the overlap with git_status, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to reach for this tool versus git_status, git_commit, or repo_read_file, nor any preconditions (e.g., repo must exist, working directory assumptions). Usage has to be inferred entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_mergeB

Merges specified branch into current branch with --no-ff

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory
branchNameYesName of the branch to merge

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the --no-ff flag, telling the agent a merge commit is always created even in fast-forward cases, but it says nothing about conflict behavior, failure/abort semantics, or repository dirty-state requirements for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words; the key flag is included rather than omitted. It is efficient, though slightly terse for a mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with zero annotations and no output schema, the description should cover failure modes, conflict handling, and preconditions. Instead it only states the operation and one flag, leaving the agent without the behavioral context needed to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both cwd and branchName are already documented in the schema, setting the baseline at 3. The description adds no syntax or format detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Merges') and resource ('specified branch into current branch'), which lets an agent distinguish it from repo_read_file or git_commit. It does not name a sibling alternative, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives (e.g., git_branch, git_commit) or when to prefer them. The agent must infer that this is the commit-integration step from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

git_statusB

Returns git working tree status, staged files, modified files, untracked files, and current branch

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional repository working directory

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the return categories, implying a read-only inspection. However, it omits operational context such as behavior outside a git repository, permission requirements, or whether the working directory defaults to the current directory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every clause enumerates a distinct return category. Nothing is wasted or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, zero-required-parameter inspection tool with no output schema, the description sufficiently covers what is returned. The minor gap is the absence of any note about non-repo behavior or cwd defaults, but the essential information for invocation is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole optional 'cwd' parameter is fully described in the schema. The description adds no syntax, default, or format detail beyond that, so the baseline of 3 for a schema-covered single parameter applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns') and resource ('git working tree status') and enumerates the concrete outputs (staged, modified, untracked files, current branch). It is immediately clear what the tool does, though it never names or contrasts itself with siblings like git_diff or git_branch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no conditions or prerequisites, and no mention of alternatives such as git_diff for content-level changes. The agent must infer usage entirely from the purpose statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

integration_verifyC

Executes end-to-end integration harness across multiple modules to verify complete system behavior

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional working directory
nameYesName of the integration test
projectIdYesID of the project
descriptionYesDescription of cross-module interaction being verified
expectedPatternNoRegex pattern expected in output
modulesInvolvedYesList of module IDs involved
integrationCommandYesCommand to run integration test suite

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says the tool 'executes' a command but does not disclose side effects, whether it requires a built workspace, what happens on failure, or whether it is safe to re-run – significant gaps for an execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the verb and scope front-loaded and no filler. It is appropriately sized, though the terseness leaves the behavioral and usage gaps unaddressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param execution tool with no annotations and no output schema, the description is too thin. It omits failure behavior, prerequisites (e.g., that modulesInvolved must already exist), and any sense of what a successful versus failed verification yields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the seven parameters (projectId, name, modulesInvolved, integrationCommand, expectedPattern, etc.) are already fully documented. The description adds no syntax, format, or constraint detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Executes) and resource (end-to-end integration harness across multiple modules) with a clear outcome (verify complete system behavior). The 'across multiple modules' scope implicitly separates it from single-module siblings like contract_verify or module_run, though no sibling is named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternative is named. The 'end-to-end' phrasing hints at its place after modules exist, but the agent must infer when this beats contract_verify or module_run.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

module_observeC

Captures and records reproducible observable evidence from an executed command or system state

ParametersJSON Schema
NameRequiredDescriptionDefault
stderrYesStandard error captured
stdoutYesStandard output captured
taskIdNoAssociated task ID
commandYesThe command or harness that was observed
exitCodeYesExit code resulting from execution
moduleIdNoAssociated module ID
evidenceTypeNoObservation modality (CLI, HTTP_API, DATABASE, etc.)
artifactPathsNoFile paths of generated artifacts to hash and verify
executionTimeMsYesDuration of execution in milliseconds

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says the tool records evidence, which implies a write/mutation operation, but it does not disclose permissions, persistence behavior, idempotency, whether existing records are overwritten, error behavior, or rate limits. The behavioral profile is therefore substantially incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant text. It is appropriately sized, though it does not provide structural cues for complex usage or sequencing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given nine parameters and no output schema, the description is minimally adequate: it conveys the core action but omits return values, storage semantics, and post-capture behavior. With no annotations and no output schema, more context would be expected for a recording/mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all nine parameters, including evidenceType and artifactPaths. The description adds no parameter-level meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Captures and records reproducible observable evidence from an executed command or system state.' It distinguishes the tool from execution-focused siblings like module_run and terminal_exec by focusing on evidence recording rather than execution. It does not explicitly name alternatives, so a 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used after a command has executed or when recording system state, but it gives no explicit when-to-use guidance, prerequisites, or alternatives. It does not clarify when to choose module_observe over contract_verify, integration_verify, or checkpoint_create.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

module_runC

Executes a module through its registered runnable harness, capturing full execution output and telemetry

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional working directory
taskIdNoOptional task ID to associate observation with
moduleIdYesID of the module to execute
inputArgsNoArguments or inputs to pass to the module harness
customHarnessNoAlternative harness command override

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Executes' implies running arbitrary code, which is potentially destructive, yet there is no disclosure of side effects, sandboxing, permission requirements, timeouts, or failure modes. 'Capturing full execution output and telemetry' hints at returns but nothing about safety or behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficiently front-loaded sentence with no filler. It is appropriately sized, though its brevity is arguably under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter execution tool with no annotations and no output schema, the single sentence is too thin. It omits behavioral and safety context an agent needs before executing a module, even though the schema covers the parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters (cwd, taskId, moduleId, inputArgs, customHarness). The description adds no meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (executes) and resource (module) and adds the distinguishing mechanism (registered runnable harness) plus the capture of output and telemetry. An agent can tell it runs something, though it does not explicitly differentiate from the sibling module_observe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the closely related module_observe tool. The agent is left to infer when running a module is appropriate versus merely observing it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_initializeB

Initializes or connects to a Vibe Engineering Project with SQLite operational state and .vibe/ directory

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProject name
rootPathNoProject root directory path (defaults to current working directory)
descriptionYesDetailed project description and purpose
initialPhaseNoInitial development phase (e.g. V0_FOUNDATION)

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full disclosure burden. It mentions creation of SQLite state and a .vibe/ directory, which hints at a write operation, but doesn't state whether it requires permissions, is idempotent, what happens if the project already exists, or what it returns. This is a significant gap for a mutation tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded and wastes no words. It could be more informative, but conciseness itself is good.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a setup/mutation operation with no annotations, no output schema, and four parameters. The description is too thin to guide an agent on when and how to use it, missing details like idempotency, required permissions, or effects on existing projects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description only names the tool and mentions SQLite/.vibe, adding no parameter semantics. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Initializes or connects to a Vibe Engineering Project') and mentions side effects (SQLite operational state, .vibe/ directory). It doesn't differentiate from siblings like project_resume or project_inspect, but the core purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. With siblings like project_resume, project_inspect, and workspace_create, an agent needs to know which one to pick. The description omits any when-to-use or prerequisite information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_inspectB

Returns a full high-density status report of the engineering control plane (tasks, modules, git state, checkpoint)

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesID of the project to inspect

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Returns' implies a read-only operation and the enumerated contents give useful insight into what the report covers, but it does not explicitly state that the tool is non-mutating, what permissions are needed, or whether it has cost/rate implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the action front-loaded and no wasted words. The parenthetical enumeration is compact and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey the return shape β€” it does so by listing the report's domains, which is helpful. However, with no annotations and no usage context, an agent lacks guidance on when this broad inspection call is warranted over narrower siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single required parameter ('projectId', documented as 'ID of the project to inspect'). The schema does the heavy lifting, so a baseline 3 is appropriate; the description adds nothing about the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns') and resource ('full high-density status report of the engineering control plane') and enumerates what it contains (tasks, modules, git state, checkpoint). It clearly conveys the tool's job, though it doesn't differentiate itself from siblings such as workspace_status, git_status, or project_resume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use this tool versus alternatives, no preconditions, and no exclusions. An agent must infer that this is the broad 'inspect everything' call from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_resumeB

Provides complete, high-density continuation context for a fresh worker model resuming development

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesID of the project to resume

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not state whether the operation is read-only or mutates project state, nor permissions required. 'High-density' hints at output size but no other behavioral traits are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, front-loaded, no wasted words. Well-structured for the small amount of information provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should disclose the safety/mutation profile and operation. It only states what context is returned, leaving the agent unsure whether resuming changes persisted state or is a pure read. Incomplete for a state-ambiguous operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so projectId is fully documented in the schema. The description adds no parameter-level detail, but the baseline of 3 applies when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the tool provides continuation context for resuming development, naming the resource and scenario. However, 'provides' is a weak verb and it doesn't distinguish from siblings like project_inspect or checkpoint_load. Clear but no sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an implicit usage scenario ('fresh worker model resuming development') but no explicit when-to-use instruction or alternatives. No mention of project_inspect or checkpoint_load, so an agent cannot tell which sibling to choose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_delete_fileC

Safely deletes a file within the project workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesRelative path to the file within the repository
workspaceRootNoOptional custom workspace root directory

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Safely deletes' is undefined β€” it does not say whether deletion is reversible, whether a confirmation/backup occurs, what permissions are required, or how a missing file is handled. For a destructive mutation tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is efficient, though arguably too terse for a destructive operation β€” brevity here borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A destructive file-removal tool with no annotations and no output schema needs the description to disclose safety semantics and failure behavior, and it does not. The vague 'safely' does not fill that void.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both filePath and workspaceRoot are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (deletes) and resource (a file within the project workspace), and the verb distinguishes it from siblings repo_read_file and repo_write_file. However, it does not explicitly differentiate itself or name scope limits beyond the workspace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is given. The agent must infer that this is the destructive counterpart to repo_write_file without any routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_read_fileB

Safely reads a file within the project workspace. Traversal outside workspace is blocked.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesRelative path to the file within the repository
workspaceRootNoOptional custom workspace root directory

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does contribute one real trait: traversal outside the workspace is blocked, which tells the agent the security posture. However, it omits other relevant behavior such as missing-file handling, binary/encoding limits, or size constraints, leaving meaningful gaps for a zero-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, and the core purpose plus the safety constraint are front-loaded. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with a fully documented two-parameter schema and no output schema, the description covers purpose and the main safety constraint adequately. Error and edge-case behavior is the only notable omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both filePath and workspaceRoot are already documented in the schema, and the description adds no parameter-level detail beyond them. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('reads a file') and scopes it to the project workspace, which cleanly distinguishes it from the sibling write/delete tools. It stops short of naming an alternative explicitly, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to prefer this over siblings like project_inspect or terminal_exec, nor any prerequisites or exclusions. A read is fairly self-evident, but the definition leaves all routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_write_fileB

Safely writes or updates a file within the project workspace, ensuring parent directories exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesFull content to write to the file
filePathYesRelative path to the file within the repository
workspaceRootNoOptional custom workspace root directory

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that parent directories are created automatically and that existing files are updated, but it omits overwrite risk, permission requirements, and what the operation returns. The word 'safely' is asserted without substantiation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with no wasted clauses, and the action is front-loaded. The vague qualifier 'Safely' is the only filler, but it does not bloat the sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description is only partially complete: it covers directory creation but says nothing about overwrite semantics, failure modes, or permissions. Adequate to call the tool, thin for a write operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (content, filePath, workspaceRoot) are already documented in the schema. The description adds no path-format or content constraints beyond what the schema states, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('writes or updates') and resource ('a file within the project workspace'), which clearly separates it from repo_read_file and repo_delete_file by name. It does not explicitly name siblings or scope boundaries, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives (e.g., terminal_exec, or a create-style sibling), nor any prerequisites or when-not-to-use conditions. The agent must infer usage purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

requirements_compileC

Compiles human-level product specifications into structured engineering capabilities and acceptance criteria

ParametersJSON Schema
NameRequiredDescriptionDefault
goalsYesHigh-level engineering goals
titleYesSpecification title
projectIdYesID of the project
capabilitiesYesStructured capabilities list

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether compilation persists to a project, overwrites existing capabilities, requires authentication, or what it returns; only a high-level transformation is described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word contributes to stating the core transformation. The size is appropriate for conveying the tool's essential purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and four required parameters, the description is insufficient for confident invocation. It omits return behavior, side effects, error modes, and how projectId or capabilities interact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema. The description's mention of capabilities and acceptance criteria echoes schema fields but adds no syntax, constraints, or formatting beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Compiles') and the transformation from human-level product specifications to structured engineering capabilities and acceptance criteria. It does not distinguish itself from siblings like architecture_generate or task_create, so sibling differentiation is absent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternative tools are provided. An agent cannot infer whether this should be called before architecture_generate or task_create, nor what inputs are required beyond the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_claimB

Claims a READY task for execution by a specific worker model

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesID of the task to claim
workerIdYesIdentifier of the worker model claiming the task (e.g. claude-3-7, gpt-4o, gemini)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations and no output schema, so the description carries the full behavioral burden. It implies a state mutation (a task becomes claimed by a worker) but omits whether the claim is exclusive/atomic, whether it fails on an already-claimed task, whether the claim expires, and what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the precondition (READY) and the constraint (specific worker model) both land immediately. It is efficient, though it is terse enough that some needed context was trimmed away.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no annotations and no output schema, in a cluster of task-lifecycle siblings (task_next, task_complete, task_fail). An agent still lacks failure semantics, claim exclusivity, and return information, so the definition is under-complete for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented in the schema, including examples of worker identifiers. The description only alludes to the READY constraint and adds no format or semantics beyond what the schema already states, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (claims) plus resource (task) plus the precondition state (READY) and the actor (worker model) β€” an agent knows exactly what this does. It does not explicitly differentiate itself from sibling task_next, which likely also hands out work, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'READY task' implies the precondition for calling it, which is useful implicit guidance. However, it never says when to prefer this over task_next, nor what to do if the task is not READY or already claimed β€” no explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_completeA

Submits a task for verification through the Completion Gate. Runs tests, executes the observable harness, inspects git diff, and enforces contract rules. If verification fails, an automatic repair task is created.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory of the workspace
taskIdYesID of the task to complete
workerIdNoWorker submitting completion
customTestCommandNoOptional specific test command to execute during gate evaluation

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it names the side effects (runs tests, executes the observable harness, inspects git diff, enforces contract rules) and critically discloses the automatic repair-task creation on failure. It stops short of stating anything about idempotency, whether the call blocks until the gate finishes, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: action first, then the concrete mechanics, then the failure consequence. Front-loaded and every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description usefully covers the most important outcome an agent needs to know (failure triggers an automatic repair task). It lacks detail on what a successful verification returns and whether the call is synchronous, leaving a small gap for a gate tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters including customTestCommand are documented in the schema. The description mentions that tests are run but adds no syntax, format, or default detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Submits a task for verification through the Completion Gate') and enumerates the concrete checks performed. It does not explicitly differentiate itself from siblings like task_fail or contract_verify, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Submits a task for verification,' but there is no explicit guidance on when to call this instead of task_fail, contract_verify, or integration_verify, nor any stated prerequisites. The agent must infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_createC

Creates a new engineering task with dependencies and optional observable contract specification

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesTask title
moduleIdNo
priorityNoMEDIUM
projectIdYesID of the project
descriptionYesDetailed task description
contractSpecNo
dependenciesNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It reveals nothing about persistence, required fields, side effects, validation of dependencies, or permissions for a mutation tool in a multi-agent workflow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. Its brevity is a virtue structurally, though it comes at the cost of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, a nested contractSpec object, and only 43% schema coverage, the description is far too thin for an agent to call this tool correctly without guessing at required inputs and contract semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43% across 7 parameters including a nested contractSpec object. The description names 'dependencies' and 'contractSpec' but adds no semantics for the required fields, the priority enum, defaults, or the evidence-type options, so it barely compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Creates a new engineering task') and names two notable capabilities (dependencies, optional observable contract specification). It is clear what the tool does, though it does not explicitly distinguish itself from siblings like task_claim or task_next.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer for alternative task operations. The agent is left to infer that this is the creation entry point in the task lifecycle.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_failC

Explicitly marks a task as FAILED and records reason

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesExplanation of why the task failed
taskIdYesID of the task that failed

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the FAILED state and that a reason is recorded, but not whether this is terminal, whether the task can be re-claimed or retried afterwards, whether it requires prior ownership via task_claim, or whether the operation is reversible. For a state-mutating tool on a task workflow, these are meaningful gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tightly written sentence with the essential state change front-loaded after the 'explicitly' qualifier. Efficient, though the leading adverb adds little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-param tool with full schema coverage and no output schema, the description covers the core action. It is incomplete regarding lifecycle consequences (terminal state, retry semantics, ordering relative to task_claim) that an agent orchestrating a task workflow would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters (taskId, reason) are self-documented in the schema. The description's phrase 'records reason' merely restates the schema, adding no format, length, or constraint detail. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'marks a task as FAILED and records reason.' An agent immediately understands the state transition and the side effect. It doesn't explicitly contrast with sibling task_complete, so it falls short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of the obvious alternative task_complete (or what to do if a task is merely blocked, not failed). The agent must infer that this is the terminal failure path rather than the success path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_nextA

Returns the next READY task in the dependency DAG for a worker to claim (prioritizing repair tasks and high priority tasks)

ParametersJSON Schema
NameRequiredDescriptionDefault
workerIdNoWorker identifier requesting task
projectIdYesID of the project

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a real behavioral trait beyond the schema: prioritization of repair tasks and high-priority tasks, plus READY-state gating. It does not say whether calling it reserves or claims the task, whether it is side-effect free, or what happens when no ready task exists β€” meaningful omissions for an assignment tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the core action first and the prioritization rule in a compact parenthetical. Every clause earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should sketch the return shape, yet it only says it returns 'the next READY task' without indicating what that object contains or the empty/no-ready-task case. Annotation coverage is absent, so a bit more behavioral context would be warranted for a 2-param orchestration tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are fully described in the schema (100% coverage), so the baseline is 3. The description only marginally reinforces workerId semantics via 'for a worker to claim' and adds nothing about projectId or the relationship between the two.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Returns) and resource (the next READY task) and adds scoping detail: it comes from the dependency DAG and is intended for a worker to claim. It is distinguishable from related siblings by implying retrieval rather than the explicit claim of task_claim, though it never names that sibling, so full sibling differentiation is only inferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied by 'for a worker to claim' and the READY-state framing, so an agent can infer this is the polling/assignment entry point after task_create. However, there is no explicit statement of when to prefer this over task_claim, nor any exclusion or prerequisite (e.g., must the worker be registered first).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminal_execA

Executes a terminal command safely scoped to the project workspace. Evaluates security policies, logs execution history, and captures stdout, stderr, exit code, and execution time.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoRelative sub-directory or workspace path (must be inside project root)
commandYesThe shell command to execute
projectIdNoOptional project ID to associate command logs with
timeoutMsNoExecution timeout in milliseconds (defaults to 30000)

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It adds useful behavioral context: security policy evaluation, logging, stdout/stderr/exit code/time capture, and workspace scoping. However, it doesn't disclose permission requirements, what happens on timeouts, whether commands are sandboxed, or error handling beyond capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scoping constraint, then lists key outputs. Every phrase adds value without redundancy. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description covers core behavior but leaves gaps: no mention of permissions, error conditions, timeout behavior, or result format beyond listing captured fields. It's minimally complete for a tool of this complexity, but an agent would need more to predict outcomes in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no parameter-specific details (e.g., timeout behavior, cwd restrictions beyond 'inside project root', projectId association). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Executes') and resource ('terminal command') with scoping modifier ('safely scoped to the project workspace'). This distinguishes it from repo_read_file, repo_write_file, and other siblings. However, it doesn't explicitly contrast with module_run, which may also execute code, leaving some ambiguity about the primary execution mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use for executing shell commands within the project workspace, but doesn't specify when to choose this over module_run or other execution tools, nor does it mention prerequisites like workspace initialization. Usage context is only partially clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_createB

Creates an isolated workspace and branch for a worker to implement a specific task safely

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesID of the task to be worked on
workerIdYesWorker identifier

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses meaningful behavioral traits β€” the workspace is 'isolated' and creation is 'safe', implying the branch won't affect the main tree β€” but says nothing about idempotency (what happens if one exists), side effects on the repo, permissions, or failure modes for a creation operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the core action front-loaded and no filler. It is appropriately sized, though it could carry one more clause of routing/behavioral detail at no real cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no annotations and no output schema, the description is minimally adequate but omits error behavior, what is returned, and how it relates to the many sibling tools. The agent knows what it does but not the full contract for invoking it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (workerId, taskId) are documented in the schema, so the baseline is 3. The description adds no format, constraint, or relationship detail beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Creates an isolated workspace and branch') and adds scope ('for a worker to implement a specific task'). This clearly distinguishes it from sibling workspace_status, though it doesn't explicitly name or contrast with adjacent tools like git_branch or workspace_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a worker to implement a specific task safely' implies the usage context (call before implementing a task), but there is no explicit when-to-use, when-not, or named alternative such as git_branch or workspace_status. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_statusB

Returns status of all active worker workspaces

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies a read but never states whether it requires authentication, whether it is scoped to a project/session, what 'active' means, or anything about rate limits or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the resource front-loaded and no filler. It is appropriately sized, though it may be too terse for what it omits.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and no return-value documentation, an agent cannot know what fields 'status' returns or how to interpret them. For a zero-param inspection tool the description should at least sketch the returned status shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is nothing for the description to disambiguate beyond the implicit scoping that 'all active' conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Returns') and resource ('status of all active worker workspaces'), which is clearer than the bare name. However, it does not differentiate itself from the sibling workspace_create or clarify what 'status' actually contains, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives are named. The agent is left to infer that this is a read-only inspection call from context alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 28 tool updatesv0.1.0
    • First observedarchitecture_generate
    • First observedarchitecture_validate
    • First observedcheckpoint_create
    • First observedcheckpoint_load
    • First observedcontract_verify
    • First observedgit_branch
    • First observedgit_commit
    • First observedgit_diff
    • First observedgit_merge
    • First observedgit_status
    • First observedintegration_verify
    • First observedmodule_observe
    • First observedmodule_run
    • First observedproject_initialize
    • First observedproject_inspect
    • First observedproject_resume
    • First observedrepo_delete_file
    • First observedrepo_read_file
    • First observedrepo_write_file
    • First observedrequirements_compile
    • First observedtask_claim
    • First observedtask_complete
    • First observedtask_create
    • First observedtask_fail
    • First observedtask_next
    • First observedterminal_exec
    • First observedworkspace_create
    • First observedworkspace_status

TDQS

B3.2/5.0

Scored across 28 tools

Disambiguation4/5

Most tools have clearly distinct resource-action purposes, with prefixes and descriptions that separate tasks, git, repo, workspace, architecture, and verification concerns. Some overlap exists among status/context tools such as project_inspect, project_resume, workspace_status, and checkpoint_load, but their descriptions make the intended use reasonably clear.

Naming Consistency5/5

All tool names use consistent snake_case with predictable domain prefixes such as task_, git_, repo_, workspace_, architecture_, and checkpoint_. The verb/noun ordering is not perfectly uniform, but the convention is highly readable and consistent across the set.

Tool Count3/5

28 tools is heavy and exceeds the ideal 3-15 range for a single MCP server. The complex engineering lifecycle justifies many of them, but there is likely room to consolidate status/context tools, making the surface borderline rather than well-scoped.

Completeness4/5

The set covers the major lifecycle stages: project setup, requirements, architecture, task DAGs, workspaces, file operations, terminal execution, git, modules, contracts, integration verification, and checkpoints. Minor gaps include no explicit task_update/task_list, checkpoint_list, or workspace_cleanup, but agents can work around these through project_inspect and terminal_exec.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a specification-driven workflow layer for AI-assisted coding, enabling agents to follow an explicit 11-phase feature workflow with checkpoints, artifacts, and quality gates.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables pipeline-driven task management for AI coding agents, with stage-gated workflows, dependency tracking, artifact versioning, and multi-agent collaboration.
    112 npm
    20
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables creative software agents to operate within bounded, temporary workcells, handling scored objectives through autonomous inspect, execute, capture, and correct loops while producing evidence-backed receipts for human review.
    Apache 2.0