Skip to main content
Glama

Solumbe

Models generate the change. Solumbe proves whether it is safe to merge.

Formerly Òtítọ́ (@bashbop/otito), renamed in 4.0.0; the CHANGELOG maps every old name to its new one.

CI npm license: MIT node Listed on mcpservers.org Solumbe MCP server – quality and maintenance score on Glama

solumbe demo

Solumbe MCP server – quality and maintenance score on Glama

Solumbe is a local-first, deterministic, model-agnostic trust layer for AI-assisted development. It builds task-aware repository context before an agent edits, scores how much a change actually touches, and gates merge readiness against the exact staged tree, with no server, no account, and no code leaving the machine.

It does not replace Claude Code, Codex, Cursor, Gemini, or any native agent harness. It runs beside them and keeps working as the models change underneath.

A passing local gate is never an automatic merge approval: hosted CI, GitHub review, CODEOWNERS, and the human release decision remain separate authorities.

Install

npm install -g @bashbop/solumbe
solumbe doctor

Or without installing: npx -y @bashbop/solumbe doctor.

Related MCP server: Graft

What it does

Rank what a change actually touches, from the request alone, with no model, no embeddings, and no network:

$ solumbe impact . "add refund handling to checkout" --top 3

  concepts: money flow

   1.  src/payment/checkout.service.ts    score 155
       |- role: required
       |- path matches: checkout
       |- symbol matches: checkout, refund
       |- concept match: money flow (×1.4 + 6.0)
       |- risk: money flow

   2.  src/payment/stripe.webhook.ts    score 6
       |- role: advisory
       |- concept match: money flow (×1.4 + 6.0)

Gate the exact staged tree, and say precisely why:

$ solumbe gate . --staged --base origin/main --request "add refund handling to checkout" --min-convergence 80

  [OK]    Changed files          1 changed file found.
  [OK]    Staged snapshot        Changed-file scope and convergence evidence are captured from the exact staged Git tree.
     |- Tree: 7d0b513dd256877d6bd309a6406468a9b1eaeedc
     |- Base: b776f04fcfbd02c8ce8d63e1a0c2450be58de039
  [OK]    Secret safety          No secret file paths and no credential values found in the changed content.
  [WARN]  Risk review            Risk-sensitive files changed; maintainer review should be explicit.
     |- src/payment/checkout.service.ts
  [OK]    Convergence            Task and diff satisfy the convergence requirement.
     |- Score: 100/100 (aligned)
     |- Receipt handle: rcpt_a8d31efae06e
     |- Subject tree: 7d0b513dd256877d6bd309a6406468a9b1eaeedc

  VERDICT  WARN

The convergence score is the part a model cannot grade for itself: it compares the stated intent against the files the diff actually changed, and produces a receipt bound to the exact base, parent, and staged-tree identity.

The core

Three commands are the product. They call no model, open no socket, and read nothing outside the repository.

Step

Command

Before the edit

solumbe context "add refund handling" --path .

Before the merge

solumbe gate . --staged --base origin/main --request "add refund handling"

For the reviewer

solumbe pr . --base origin/main --out .solumbe/pr-review.md

Nothing hosted is required, and nothing hosted can change a verdict. Local Core, Optional Hosted lists exactly what the optional pieces send, and where.

Supporting commands

Goal

Command

Rank change blast radius

solumbe impact . "add refund handling" --top 12

Score intent vs. execution

solumbe converge "add refunds" --path . --base HEAD --staged

Score a committed change exactly

solumbe converge "add refunds" --path . --base HEAD~1 --head HEAD

Gate a product change across repos

solumbe workspace-gate ../web ../api --request "ship change"

Grade risk flags against history

solumbe calibrate . --window 30

Grade route tiers against history

solumbe regret . --window 30 --offline

Regrade a saved run, same answers

solumbe regret --rescore run.json

Score Agent Experience

solumbe ax . "add a new MCP tool"

Recommend a model tier

solumbe route . "add a new MCP tool"

Sharpen context with a model read

solumbe context "add a new MCP tool" --path . --online

Inspect one repo

solumbe repo . --json

Build a code map

solumbe map . --json

Index and search local projects

solumbe index ~/projects --discover then solumbe search "events controller"

Generate an agent harness

solumbe harness . --out .solumbe/harness.md

Run the MCP server

solumbe mcp

Every command takes --json, and solumbe help lists the full set with flags.

MCP

Solumbe ships a stdio MCP server exposing 14 tools: repo_inspect, repo_map, repo_index, repo_search, context_pack, change_impact, agent_experience, model_route, convergence_score, review_context, review_gate, review_verdict, workspace_report, and repo_harness.

{
  "mcpServers": {
    "solumbe": {
      "command": "npx",
      "args": ["-y", "@bashbop/solumbe", "mcp"]
    }
  }
}

Published in the MCP Registry as io.github.BASHBOP/solumbe. Repo-map lookups use an external per-user cache and never write into the inspected repository. Host-specific setup for Claude Code, Claude Desktop, Codex, Cursor, VS Code, Gemini CLI, and Kimi Code is in MCP and Agent Workflows.

How it compares

Approach

Strengths

Where solumbe differs

Sourcegraph / Cody context

Powerful hosted code search and embedding-based context across an org

Local-first and deterministic: no server, no account, no code leaves the machine, and the same query always yields the same packet

Hand-written CLAUDE.md / rules files

Curated, intent-rich guidance

Hand-written context goes stale; solumbe regenerates context from the actual code (symbols, imports, routes, tests) on every run and complements a short CLAUDE.md

grep / ripgrep

Fast, universal text matching

solumbe ranks whole files by task intent across paths, symbols, exports, and tests, then adds patterns and validation commands, producing a context packet rather than a list of matching lines

Documentation

Full command reference, agent workflows, release process, evaluation method, and the design theses behind the trust layer:

bashbop.github.io/solumbe

Contributing

git clone https://github.com/BASHBOP/solumbe.git
cd solumbe && npm ci && npm run ci

npm run ci is the full gate: format, lint, typecheck, version check, tests, coverage floors (70% lines / 60% branches / 75% functions), three evaluation corpora, dependency audit, and a packaged-tarball smoke test. Run it before requesting review.

Start with CONTRIBUTING.md and the Code of Conduct. All changes need maintainer review; main requires passing gates and resolved conversations. Solumbe follows Semantic Versioning, so say whether a PR is no-impact, patch, minor, or major.


Part of the toolchain

solumbe is one of four tools that form a deterministic trust layer for AI-assisted development. Each uses static analysis to answer a question people keep handing to an LLM.

  • solumbe (this tool), for context: what does this change actually touch?

  • tieline, for contracts: did the front end and back end quietly stop agreeing?

  • bouncer, for compliance: could you defend this to Ofcom?

  • aiglare, for governance: where can the model do something you can't undo?

More at segunolumbe.com. static analysis, never the model.

Available Tools

14 tools
agent_experienceAgent Experience (AX)A
Read-only

Score Agent Experience (AX): a single 0–100 number answering "how cheap and safe is it for an agent to make this change here?", blending Changeability (token cost), Containment (blast radius), Guardrails (tests / validation / CODEOWNERS / CI), and Clarity (task groundedness). Deterministic and composed from the change_impact, token-estimate, and CODEOWNERS engines — no new analysis. Returns sub-scores, drivers, and concrete recommendations for raising AX. Uses a per-user external cache and leaves the target repository unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoNumber of impact files to consider when scoring blast radius. Defaults to 8.
pathNoRepository path. Defaults to current working directory.
queryYesPlain-English change request to score, e.g. "add a new MCP tool".
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint annotation by explicitly stating that it 'leaves the target repository unchanged' and disclosing a real side effect: it 'uses a per-user external cache.' It also adds determinism and composition guarantees, which are important behavioral traits not visible in the schema or annotations. There is no contradiction: the cache note qualifies the read-only guarantee.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the first defines the metric in plain terms, the second explains composition and determinism, the third states the return payload, and the fourth discloses side effects and safety. The core purpose is front-loaded)Skip, and the language is dense without being padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description compensates by specifying the return shape at a useful level: 'sub-scores, drivers, and concrete recommendations for raising AX,' plus the markdown option. It could go slightly further on cache behavior or failure semantics, but for a deterministic read-only scoring tool with a fully documented input schema, this is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description does not add significant new meaning about parameter formats or edge cases; it only reinforces the spirit of the query parameter by describing the change request as 'how cheap and safe is it for an agent to make this change here?' This meets the baseline but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Score Agent Experience (AX): a single 0–100 number.' It then defines what AX measures (cheapness/safety of a change) and lists its four components, making the tool's purpose unmistakable and clearly distinct from the raw engines it composes, such as change_impact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: when an agent needs an aggregate AX score and concrete recommendations rather than raw engine output. It states that the tool is 'composed from the change_impact, token-estimate, and CODEOWNERS engines — no new analysis,' which implicitly positions it as the higher-level or synthetic choice among sibling tools, though it does not explicitly list exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

change_impactChange ImpactA
Read-only

Given a plain-English change request, rank the files most likely to own the change, with risk flags, suggested tests, and an implementation plan. Optional diff base surfaces exact changed-file evidence alongside the heuristic. Uses a per-user external cache and leaves the target repository unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoNumber of files to return. Defaults to 10.
pathNoRepository path. Defaults to current working directory.
queryYesPlain-English change request.
diffBaseNoOptional git ref to validate predictions against (e.g. origin/main, HEAD).
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds useful behavioral context: it leaves the target repository unchanged, uses a per-user external cache, and can combine heuristic ranking with exact diff evidence. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with the main action and outputs front-loaded. Every sentence adds information: what the tool does, what it returns, when the optional diff base matters, and its cache/read-only behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description explains the return payload components: ranked files, risk flags, suggested tests, and an implementation plan. It also notes the optional exact-evidence mode, external cache, and non-mutation, which is enough for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds meaning by characterizing query as a plain-English change request and diffBase as providing exact changed-file evidence alongside the heuristic, going beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Given a plain-English change request, rank the files most likely to own the change' and enumerates the outputs (risk flags, suggested tests, implementation plan). This clearly distinguishes it from sibling repo inspection/search tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly frames when to use the tool: when there is a plain-English change request, and when to use the optional diff base to surface exact changed-file evidence. It does not explicitly name alternatives or exclusions, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_packContext PackA
Read-only

Generate a task-aware local context packet with primary files, related files, tests, and validation commands. Uses a per-user external cache and leaves the target repository unchanged. Deterministic by default; pass online:true to also ask a System One model (needs TYPESAFE_API_KEY) what kind of work the request is and which candidate files it needs, which relabels a confidently read intent and demotes files it judges irrelevant, reported under modelRead.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository path. Defaults to current working directory.
limitNoMaximum primary, related, and test files per section. Defaults to 8.
pathsNoRepository paths for a multi-repo context packet. Omitted, `path` plus the `companions` listed in its .solumberc.json.
queryYesTask or question to gather context for.
onlineNoAsk TypeSafe's Jev to read the request (intent, and relevance of each candidate file) in one call, and apply answers that clear their confidence gate. Needs TYPESAFE_API_KEY; without one the pack is returned unchanged with modelRead.source offline. Defaults to false.
includeEvidenceNoInclude per-file imports, exports, and symbol slices as source evidence. Defaults to false to keep the packet compact.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON packet. Defaults to false.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint; the description adds substantive behavior beyond that: a per-user external cache, determinism by default, a network/auth dependency (TYPESAFE_API_KEY) for online mode, and the concrete side effects of that mode (relabeling a confidently read intent, demoting files judged irrelevant, reported under modelRead). The 'leaves the target repository unchanged' line overlaps the readOnlyHint but usefully confirms write safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose before the optional online mode, with no filler. The third sentence is dense and somewhat run-on, but every clause carries information about behavior the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, cache-backed generation tool with no output schema, the description covers what is produced, the default behavior, the opt-in network path and its prerequisites, and a named output field (modelRead). An agent has enough to invoke it correctly without opening the schema, and return-value detail is minimal since the packet structure is described inline.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's online explanation largely restates the schema's own description of that flag (intent reading, confidence gate, TYPESAFE_API_KEY, offline fallback), adding little beyond the schema, and it says nothing about path, paths, limit, includeEvidence, or includeMarkdown.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ('Generate a task-aware local context packet') and enumerates the packet's contents (primary files, related files, tests, validation commands), which distinguishes it from simpler retrieval siblings like repo_search or repo_map. It stops short of naming which sibling to prefer when, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real usage conditions for the optional mode ('pass online:true to also ask a System One model... needs TYPESAFE_API_KEY') and notes the deterministic default, so the intended workflow is implied. However, it never states when to choose context_pack over alternatives such as review_context, repo_inspect, or change_impact, leaving sibling selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

convergence_scoreConvergence ScoreA

Score convergence: a deterministic 0–100 measure of the distance between a stated task (intent) and the actual git diff (execution), with sub-scores for Coverage (did the intent happen?), Scope (did only the intent happen?), and Risk alignment (did unrequested drift land on risk-sensitive paths?). Composed from the change_impact diff comparison and the shared risk vocabulary — no model, no new analysis. New files beside a confirmed owner and tests of confirmed files count as in scope, listed under drivers.inferredRelated. Emits a recomputable receipt (a timestamp-free hash anyone can regenerate and verify) as durable evidence. Requires a base git ref to diff against; scores the working tree's tracked changes unless head or staged selects an exact subject.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoNumber of predicted owner files to consider. Defaults to 10.
baseYesGit ref to diff against (e.g. origin/main, HEAD~1). Required.
headNoScore exactly base..head (a direct tree diff, no merge base) instead of the working tree, and bind the receipt to the head commit and its tree SHA. Uncommitted and untracked files play no part. Cannot be combined with staged.
pathNoRepository path. Defaults to current working directory.
queryYesThe stated task / intent to measure against the diff, e.g. "add Stripe refunds".
stagedNoMeasure only the exact Git index tree and bind the receipt to its tree SHA. May create Git object or index-cache metadata; source files are unchanged.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.
includeUntrackedNoWorking-tree mode only: also score untracked, non-ignored files. Defaults to false; untracked files are otherwise listed under untracked and not scored.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse readOnlyHint=false annotation, the description discloses determinism, the absence of model analysis, the recomputable timestamp-free receipt, the default working-tree subject, and the side effect that staged mode 'may create Git object or index-cache metadata' while source files remain unchanged. This is rich behavioral context that helps an agent anticipate side effects and reproducibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: definition, sub-scores, composition, scope rules, receipt, and mode selection. It is front-loaded with the core purpose and avoids filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—8 parameters, no output schema, and a rich sibling set—the description covers the essential invocation context: what is measured, how scope is determined, how modes alter the subject, side effects, and the durable receipt. An agent has enough to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful cross-parameter semantics: it explains that head and staged select an exact subject instead of the working tree, that includeUntracked only applies in working-tree mode, and that staged can create metadata. This goes beyond the individual parameter descriptions in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Score convergence: a deterministic 0–100 measure of the distance between a stated task (intent) and the actual git diff (execution).' It names the three sub-scores and explicitly distinguishes itself from the sibling change_impact tool by stating it is 'composed from the change_impact diff comparison' rather than being a raw diff tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for invocation: it requires a base git ref, defaults to the working tree, and explains the head/staged modes. However, it never explicitly states when to choose this tool over siblings like change_impact, review_gate, or model_route, nor does it provide exclusions or alternative conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

model_routeModel RouteA
Read-only

Recommend a model tier (cheap, mid or premium) for a coding request before any work starts. Combines solumbe's deterministic repository half (AX, containment, top-severity risk paths) with a System One model's calibrated read of the request (specificity, blast radius, novelty). The model is called only when TYPESAFE_API_KEY is set in this server's environment and offline is not true; otherwise the read is a labelled offline estimate, and model.source says which. Advisory: it recommends a tier and never feeds review_gate or review_verdict. Pass host to map the tier to that host's model id.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoNumber of impact files to consider. Defaults to 8.
hostNoOptional host whose model map resolves the tier to a model id (built in: claude-code; add others in .solumbe/model-route.json).
pathNoRepository path. Defaults to current working directory.
queryYesPlain-English change request to route, e.g. "fix the typo in the README".
offlineNoNever call the model, even when a key is set. Defaults to false.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses behavior well beyond the annotations: the model is only called when TYPESAFE_API_KEY is set and offline is not true, otherwise a labelled offline estimate is returned, and model.source indicates which path ran. It also explains the deterministic+System One blend and the advisory-only contract. This is rich context an agent could not infer from readOnlyHint/openWorldHint alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then progressively adds the auth/offline condition, the advisory contract, and the host mapping. Dense but every sentence carries information; a touch long, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter advisory tool with no output schema but a markdown-report option, the description covers the return signal (model.source), the fallback path, and the host mapping. An agent has enough to call it and interpret the result, though it never mentions the 'top' impact-file parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description still adds meaning by explaining what host does (maps the tier to that host's model id) and tying offline to the labelled-estimate fallback and model.source. It reinforces, though does not fully replace, the schema's parameter docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Recommend a model tier (cheap, mid or premium) for a coding request before any work starts.' It also separates itself from siblings by noting it is advisory and 'never feeds review_gate or review_verdict', so an agent can tell it apart from the review tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context (use before any work starts, for routing a change request) and names the review_gate/review_verdict alternative it is not. It stops short of stating explicit when-not-to-use cases beyond the advisory disclaimer, but the routing intent is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_harnessRepository HarnessA
Read-only

Generate setup, validation, runtime, and context commands for an agent or CI harness.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository path. Defaults to current working directory.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations declare readOnlyHint: true, so the description does not need to repeat that. It adds a behavioral detail: it can generate different types of commands (setup, validation, runtime, context) and can optionally return a markdown report (via includeMarkdown) or full JSON. This goes beyond what annotations provide, clarifying the output format options and the scope of generated content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence, concise and clear, with the verb first ('Generate') and the main part stated upfront. It efficiently conveys the purpose without unnecessary words. It could be slightly improved by mentioning the markdown option in the description for completeness, but it is already moderately concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no required parameters and a 100% schema description coverage, the description covers the essential purpose and basic behavior. It does not mention return values, but since there's no output schema and the schema covers parameters, the description is sufficient for an agent to understand what the tool does. The lack of detail on how to use the output (e.g., how to interpret the generated commands) is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters adequately. The description does not add semantic detail beyond what the schema provides—e.g., it doesn't explain how 'path' might affect command generation or what 'includeMarkdown' changes in detail. Given the high coverage, a score of 3 is appropriate as a baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and a set of resources ('setup, validation, runtime, and context commands'), which clearly identifies the tool's purpose. However, it does not explicitly differentiate it from sibling tools like 'context_pack' or 'workspace_report', which might also generate context-related outputs. The phrase 'for an agent or CI harness' adds useful context but does not distinguish between the tool and its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage in the context of an agent or CI harness (when you need to generate commands), but it does not explicitly state when to use this tool versus alternatives like 'repo_inspect' or 'context_pack'. There is no mention of conditions or exclusions. The phrasing 'Generate ... commands' gives a hint that it's for command generation, but without explicit guidance, an agent might struggle to select the right tool among siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_indexIndex RepositoriesA

Index local repositories: generate per-user external indexes and add them to the local catalog. Pass discover:true to find repository roots under the given paths first. Pass dryRun:true to discover and report without writing any indexes or catalog (read-only); this replaces the old repo_discover tool. Without dryRun this mutates the persistent local catalog, so it is not a pure read.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoSingle repository path.
depthNoMaximum discovery depth.
limitNoMaximum discovered repositories.
pathsNoRepository paths, or roots when discover is true.
dryRunNoDiscover and report repositories without writing indexes or mutating the catalog. Read-only. Defaults to false.
catalogNoOptional catalog JSON path.
discoverNoDiscover repositories under the provided paths before indexing.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint is false, and the description aligns with that by stating 'Without dryRun this mutates the persistent local catalog, so it is not a pure read.' It also discloses that dryRun is read-only. This goes beyond the annotation by specifying the exact mutation target (catalog) and the read-only mode. No contradictions detected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise—three sentences total. It front-loads the primary purpose, then covers key options (discover, dryRun) and a critical caveat about mutation. Every sentence carries useful information with no filler. Excellent structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 parameters (none required), no output schema, and no nested objects, the description covers the essential operational modes: discovery and dryRun. It explains the mutation side effect and the replacement of an older tool. It does not describe the return value or the format of the generated indexes, but since there is no output schema and the description is already focused on invocation behavior, this is adequate. A 4 is appropriate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds value by explaining the functional interplay of dryRun and discover, e.g., 'Pass discover:true to find repository roots under the given paths first' and 'Pass dryRun:true to discover and report without writing any indexes.' This is beyond the individual parameter descriptions, providing usage semantics that help the agent decide how to set them. Baseline is 3 due to high coverage, but the added context justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: 'Index local repositories: generate per-user external indexes and add them to the local catalog.' It identifies the primary action and resource. It does not explicitly differentiate from sibling tools like repo_inspect or repo_map, but the verb 'index' and the mention of generating indexes is sufficiently distinct. A 4 is warranted because while the purpose is clear, it lacks direct sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on using the dryRun parameter: 'Pass dryRun:true to discover and report without writing any indexes or catalog (read-only)' and explains that without dryRun it mutates the persistent catalog. It also mentions that dryRun replaces the old repo_discover tool. However, it does not compare this tool to alternatives like repo_inspect or repo_map, so the guidance is mostly about parameter usage rather than when to choose this tool over others. This is a solid 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_inspectInspect RepositoryA
Read-only

Inspect repository shape, languages, package managers, script names, entrypoints, and git metadata. Returns up to 200 representative file paths; pass includeScripts:true to get full script command bodies instead of just names.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoRepository path. Defaults to current working directory.
includeScriptsNoInclude full package.json script command bodies, not just their names. Defaults to false.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint:true, and the description adds valuable behavioral details: it returns up to 200 representative file paths and explains that includeScripts:true yields full script bodies instead of names. These specifics go beyond the annotation and help the agent predict output shape and parameter effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero redundancy. It front-loads the core purpose, then adds a key output limit and the conditional includeScripts behavior. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two optional parameters and no output schema, the description provides the essential purpose, key output characteristics (up to 200 paths), and a notable behavioral switch. It lacks details on return structure (e.g., exact schema of metadata), but that is minor given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters. The description adds little beyond the schema—it mentions includeScripts behavior, but that is already captured in the schema. No additional parameter context is provided, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Inspect' with a clear resource 'repository' and enumerates concrete aspects: shape, languages, package managers, script names, entrypoints, and git metadata. This distinguishes it from sibling tools like repo_map (mapping) and repo_search (searching) by specifying the inspection scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by naming the tool's function, but it does not explicitly state when to use it over alternatives like repo_map or repo_index. There is no mention of exclusions or alternative routing, leaving the agent to infer from the name and sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_mapMap RepositoryA
Read-only

Map a repository into a compact JSON code map, optionally narrowed by domain, file kind, or controller route. Pass domain to find files for a feature, kind to find files of a type (route, controller, service, apiClient, component, test), or route to match Nest controller routes by substring/regex. Replaces the old find_domain, find_file_kind, find_backend_route, and find_frontend_api_client tools. Uses a per-user external cache and leaves the target repository unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOptional file kind filter such as route, controller, service, apiClient, component, test.
pathNoRepository path. Defaults to current working directory.
limitNoMaximum files to include. Defaults to 100.
routeNoOptional route filter. Substring or regex matched against each controller's combined route (controllerBasePath + httpMethods paths) and file path.
domainNoOptional domain filter such as booking, payment, email, events.
includeFilesNoInclude matching files in the response. Defaults to false.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description explicitly states 'Uses a per-user external cache and leaves the target repository unchanged,' which discloses side effects and safety. It also reveals that the output is a compact JSON code map and that this tool replaces four older tools, adding useful behavioral context beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is composed of four concise sentences, each carrying distinct information: core function, filter usage, legacy replacement, and caching/non-destructive behavior. It is front-loaded with the primary purpose and contains no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, filter semantics, caching, and non-destructive behavior, and the schema fully documents all six optional parameters. However, it does not describe the structure or contents of the resulting 'code map' beyond calling it compact JSON, which is a notable gap given there is no output schema. An agent might not know what fields or sections to expect in the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds semantic meaning for the main filters: domain for feature files, kind for file types, and route for Nest controller routes. This clarifies intent beyond the schema's terse examples, especially by linking 'kind' to a specific list of file types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Map a repository into a compact JSON code map' with optional narrowing by domain, file kind, or route. It also names the legacy tools it replaces, clarifying that this is the consolidated successor for those search/filter operations. This makes the tool's purpose unambiguous and distinct from the sibling repo_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear guidance on when to use each filter ('Pass domain to find files for a feature, kind to find files of a type...'), which is useful for parameter selection. However, it does not explicitly compare against current siblings like repo_search, repo_inspect, or repo_index; it only mentions replaced legacy tools. Thus, when-to-use versus current alternatives is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_contextReview ContextA

Produce PR review context from local git diff metadata, optionally enriched with GitHub PR comments. Use review_context when you want the raw diff/comment context for a change, not a verdict. Use review_verdict instead for the full impact + review + gate composite, or review_gate for the gate verdict alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoBase ref. Defaults to PR base, upstream, origin/main, or main.
headNoHead ref. Defaults to HEAD.
pathNoRepository path. Defaults to current working directory.
githubNoAsk gh to infer the current branch PR.
numberNoOptional GitHub PR number for gh enrichment.
commentNoCreate or update a sticky GitHub PR comment using gh. This writes to GitHub.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation readOnlyHint=false already signals possible mutationais, and the description does not contradict that. However, the description does not disclose that setting comment=true creates or updates a GitHub comment; that critical side effect is only visible in the schema's parameter description. The description adds scope context but misses the write path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences with zero filler. The core purpose and scope come first, then explicit routing to alternatives. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has seven parameters and no output schema, but the high-coverage schema plus this description cover the main invocation needs. It could add a note about GitHub auth/network requirements for the gh enrichment, but nothing essential is missing for basic tool selection and use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already fully documented. The description adds the high-level concept of 'raw diff/comment context' but does not enrich individual parameter meanings beyond the schema, which matches the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Produce PR review context from local git diff metadata, optionally enriched with GitHub PR comments,' which names a specific verb, resource, and optional data source. It explicitly differentiates from siblings by saying this is raw diff/comment context, not a verdict, and names review_verdict and review_gate as the alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives direct when-to-use guidance: 'Use review_context when you want the raw diff/comment context for a change, not a verdict.' It also names the exact sibling tools for alternative scenarios, leaving little room for an agent to select the wrong tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_gateReview GateA

Gate a change for merge and return a PASS / WARN / FAIL verdict. Omit pr to run the local, no-GitHub gate against a base ref (changed files, secret safety, risk-sensitive paths, release discipline, validation commands, dependency audit, policy profile). Set pr to gate an open GitHub PR via gh (PR state, review decision, CODEOWNERS approvals, unresolved conversations, branch protection, status checks). Merges the old merge_readiness and pr_merge_readiness tools. Use review_verdict instead when you want the full impact + review + gate composite, or review_context when you want diff/comment context with no verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
prNoPR selector (number, URL, or branch). Set, runs the GitHub gate on that PR; omitted or blank, runs the local gate.
baseNoBase ref for the local gate. Defaults to origin/main, then HEAD. Ignored in PR mode.
headNoGate exactly base..head: changed-path, risk, secret, and convergence evidence come from the head commit's tree, so uncommitted and untracked files play no part. Release, validation-command, and optional analyzer checks still inspect the working tree. Cannot be combined with staged. Ignored in GitHub PR mode.
pathNoRepository path. Defaults to current working directory.
policyNoPolicy profile: standard, company, or high-risk. Omitted, the repository's .solumberc.json or user config decides, else standard.
stagedNoUse the exact Git index for changed-path, risk, secret, and convergence evidence. Release, validation-command, and optional analyzer checks still inspect the working tree. May create Git object or index-cache metadata; source files are unchanged. Ignored in GitHub PR mode.
receiptNoOptional convergence inputs hash, JSON object, or path to a JSON receipt artifact. Exact-subject v2 receipts require the full hash; legacy v1 display IDs remain accepted. A receipt bound to a head commit or the staged index verifies only in the same mode (head / staged); a JSON receipt names the mode it needs when they differ.
requestNoOptional change request for context evidence output.
governanceNoGovernance: team or solo. Omitted, the repository's .solumberc.json or user config decides, else team.
minConvergenceNoOptional minimum convergence score (0–100). Enables a failing gate when the task/diff score is below this floor.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, so the description must carry behavioral weight. It usefully describes what each mode inspects and notes the tool merges two prior tools, but it does not disclose mutation side effects, permission/auth needs, or what the verdict contains. The side-effect hint about Git object/index-cache metadata lives only in the schema, not the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then modes, then alternatives in three well-ordered sentences. The long parenthetical enumerations of checks are dense but informative; minor legacy-tool reference is the only slightly expendable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a ten-parameter, no-output-schema, mutation-flagged tool, the description covers mode selection and alternatives thoroughly. It leaves return-value detail and mutation/permission behavior to the schema, which is a modest but real gap given the absent output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all ten parameters in detail. The description reinforces the pr on/off branching but adds no syntax, format, or constraint detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (gate), resource (a change for merge), and the concrete output (PASS/WARN/FAIL verdict). It explicitly separates the two operating modes (omit pr = local gate; set pr = GitHub PR gate) and names the legacy tools it consolidates, so an agent can distinguish it from siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing rules: omit pr for the local gate, set pr for the GitHub PR gate, and names two alternatives with the conditions that select them (review_verdict for the full composite, review_context for diff/comment context with no verdict). Nothing about when-to-use is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_verdictReview VerdictA
Read-only

Run the full review pipeline in one shot: change_impact plus review_context plus review_gate, returning a unified verdict with a derived confidence score. Use review_verdict when you want the complete picture of a change in a single call. Use review_context instead for diff metadata only (no verdict), or review_gate for the gate verdict alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
prNoPR selector (number, URL, or branch). Set, gates that GitHub PR, as review_gate does; omitted or blank, runs the local gate.
baseNoBase ref for local diff. Defaults to origin/main, then HEAD.
pathNoRepository path. Defaults to current working directory.
policyNoPolicy profile: standard, company, or high-risk. Omitted, the repository's .solumberc.json or user config decides, else standard.
receiptNoOptional convergence inputs hash, JSON object, or path to a JSON receipt artifact. Exact-subject v2 receipts require the full hash; legacy v1 display IDs remain accepted.
requestNoPlain-English change request for impact scoring.
impactTopNoNumber of impact files. Defaults to 8.
governanceNoGovernance: team or solo. Omitted, the repository's .solumberc.json or user config decides, else team.
minConvergenceNoOptional minimum convergence score (0–100) enforced by the merge gate.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already covers the safety profile, and the description adds genuinely new behavior: this call composes three other tools and synthesizes a confidence score, which explains latency and result shape. It doesn't discuss permissions, rate limits, or how partial failures in the composed stages are surfaced, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences: capability first, then routing guidance. Every clause earns its place, with no restatement of the name or padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description states what comes back (a unified verdict plus a derived confidence score), which is the essential return-value information for a read-only aggregator. It does not sketch the verdict's internal structure or how the composed stages' findings are merged, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and every parameter carries a rich inline description, including defaults and PR-vs-local gating behavior, so the schema does the heavy lifting. The description adds no parameter-level meaning beyond what is already documented, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and composite resource: it runs the full review pipeline (change_impact plus review_context plus review_gate) and returns a unified verdict with a derived confidence score. It explicitly distinguishes itself from the sibling tools it composes, so an agent can select it without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use condition ('when you want the complete picture of a change in a single call') and names two alternatives with the conditions that select them: review_context for diff metadata only, review_gate for the gate verdict alone. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workspace_reportWorkspace ReportB
Read-only

Generate a product-level report across multiple related repositories.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYesRepository paths to inspect together.
includeMarkdownNoReturn a compact human-readable markdown report instead of the full JSON. Defaults to false.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, and the description adds no behavioral context beyond saying it 'generates' a report. It does not describe output format, aggregation behavior, scope limitations, or any side effects, so it relies entirely on the annotation for safety transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence with no filler. The core action, scope, and subject are front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and read-only, and the schema covers parameters well, but the description does not clarify what a 'product-level report' contains or how the repositories relate. Since there is no output schema, a little more detail about the report's scope or contents would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters (paths, includeMarkdown) are adequately described in the schema. The description itself adds no parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Generate') and resource ('product-level report') and adds the scoping phrase 'across multiple related repositories,' which distinguishes it from single-repo sibling tools like repo_inspect or repo_map. However, it does not explicitly name a sibling, so differentiation is slightly implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'product-level report across multiple related repositories' clearly implies a multi-repo aggregation context, but the description never states when to prefer this tool over alternatives or when not to use it. With 12 sibling tools, explicit selection guidance would meaningfully improve the definition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv3.5.1
    • Changedcontext_pack1 field changed
      • changedInput schema / properties / paths / description
        Previous value: -"Repository paths for a multi-repo context packet. Omitted, `path` plus the `companions` listed in its .otitorc.json."New value: +"Repository paths for a multi-repo context packet. Omitted, `path` plus the `companions` listed in its .solumberc.json."
    • Changedmodel_route1 field changed
      • changedInput schema / properties / host / description
        Previous value: -"Optional host whose model map resolves the tier to a model id (built in: claude-code; add others in .otito/model-route.json)."New value: +"Optional host whose model map resolves the tier to a model id (built in: claude-code; add others in .solumbe/model-route.json)."
    • Changedreview_gate2 fields changed
      • changedInput schema / properties / governance / description
        Previous value: -"Governance: team or solo. Omitted, the repository's .otitorc.json or user config decides, else team."New value: +"Governance: team or solo. Omitted, the repository's .solumberc.json or user config decides, else team."
      • changedInput schema / properties / policy / description
        Previous value: -"Policy profile: standard, company, or high-risk. Omitted, the repository's .otitorc.json or user config decides, else standard."New value: +"Policy profile: standard, company, or high-risk. Omitted, the repository's .solumberc.json or user config decides, else standard."
    • Changedreview_verdict2 fields changed
      • changedInput schema / properties / governance / description
        Previous value: -"Governance: team or solo. Omitted, the repository's .otitorc.json or user config decides, else team."New value: +"Governance: team or solo. Omitted, the repository's .solumberc.json or user config decides, else team."
      • changedInput schema / properties / policy / description
        Previous value: -"Policy profile: standard, company, or high-risk. Omitted, the repository's .otitorc.json or user config decides, else standard."New value: +"Policy profile: standard, company, or high-risk. Omitted, the repository's .solumberc.json or user config decides, else standard."
  2. 1 tool updatev3.5.0
    • Changedcontext_pack1 field changed
      • changedInput schema / properties / paths / description
        Previous value: -"Repository paths for a multi-repo context packet."New value: +"Repository paths for a multi-repo context packet. Omitted, `path` plus the `companions` listed in its .otitorc.json."
  3. 2 tool updatesv3.2.0
    • Changedreview_gate3 fields changed
      • changedInput schema / properties / governance / description
        Previous value: -"Governance: team (default) or solo."New value: +"Governance: team or solo. Omitted, the repository's .otitorc.json or user config decides, else team."
      • changedInput schema / properties / policy / description
        Previous value: -"Policy profile: standard (default), company, or high-risk."New value: +"Policy profile: standard, company, or high-risk. Omitted, the repository's .otitorc.json or user config decides, else standard."
      • changedInput schema / properties / pr / description
        Previous value: -"Optional PR selector (number, URL, or branch). When set, runs the GitHub gate; when absent, runs the local gate."New value: +"PR selector (number, URL, or branch). Set, runs the GitHub gate on that PR; omitted or blank, runs the local gate."
    • Changedreview_verdict3 fields changed
      • changedInput schema / properties / governance / description
        Previous value: -"Governance: team (default) or solo."New value: +"Governance: team or solo. Omitted, the repository's .otitorc.json or user config decides, else team."
      • changedInput schema / properties / policy / description
        Previous value: -"Policy profile: standard (default), company, or high-risk."New value: +"Policy profile: standard, company, or high-risk. Omitted, the repository's .otitorc.json or user config decides, else standard."
      • changedInput schema / properties / pr / description
        Previous value: -"Optional PR selector. When set, pass-pr runs against GitHub instead of local mode."New value: +"PR selector (number, URL, or branch). Set, gates that GitHub PR, as review_gate does; omitted or blank, runs the local gate."
  4. 2 tool updatesv1.15.2
    • Changedconvergence_score2 fields changed
      • addedInput schema / properties / head
        Added value: +{
        +  "description": "Score exactly base..head (a direct tree diff, no merge base) instead of the working tree, and bind the receipt to the head commit and its tree SHA. Uncommitted and untracked files play no part. Cannot be combined with staged.",
        +  "type": "string"
        +}
      • addedInput schema / properties / includeUntracked
        Added value: +{
        +  "description": "Working-tree mode only: also score untracked, non-ignored files. Defaults to false; untracked files are otherwise listed under untracked and not scored.",
        +  "type": "boolean"
        +}
    • Changedreview_gate2 fields changed
      • addedInput schema / properties / head
        Added value: +{
        +  "description": "Gate exactly base..head: changed-path, risk, secret, and convergence evidence come from the head commit's tree, so uncommitted and untracked files play no part. Release, validation-command, and optional analyzer checks still inspect the working tree. Cannot be combined with staged. Ignored in GitHub PR mode.",
        +  "type": "string"
        +}
      • changedInput schema / properties / receipt / description
        Previous value: -"Optional convergence inputs hash, JSON object, or path to a JSON receipt artifact. Exact-subject v2 receipts require the full hash; legacy v1 display IDs remain accepted."New value: +"Optional convergence inputs hash, JSON object, or path to a JSON receipt artifact. Exact-subject v2 receipts require the full hash; legacy v1 display IDs remain accepted. A receipt bound to a head commit or the staged index verifies only in the same mode (head / staged); a JSON receipt names the mode it needs when they differ."
  5. 2 tool updatesv1.14.1
    • Changedcontext_pack1 field changed
      • addedInput schema / properties / online
        Added value: +{
        +  "description": "Ask TypeSafe's Jev to read the request (intent, and relevance of each candidate file) in one call, and apply answers that clear their confidence gate. Needs TYPESAFE_API_KEY; without one the pack is returned unchanged with modelRead.source offline. Defaults to false.",
        +  "type": "boolean"
        +}
    • Addedmodel_route
  6. 13 tool updates
    • First observedagent_experience
    • First observedchange_impact
    • First observedcontext_pack
    • First observedconvergence_score
    • First observedrepo_harness
    • First observedrepo_index
    • First observedrepo_inspect
    • First observedrepo_map
    • First observedrepo_search
    • First observedreview_context
    • First observedreview_gate
    • First observedreview_verdict
    • First observedworkspace_report

TDQS

A3.9/5.0

Scored across 14 tools

Disambiguation4/5

Most tools carve out distinct roles, and descriptions explicitly cross-reference siblings (review_context vs review_verdict vs review_gate; repo_index vs repo_map vs repo_search). Some inherent overlap remains among the analysis engines (change_impact, context_pack, agent_experience, convergence_score), which all consume a task/repo and could be confused by a new agent despite the guidance.

Naming Consistency4/5

All names are snake_case and grouped under clear prefixes (repo_*, review_*), which aids navigation. However the verb/noun form is mixed, with verb-style names (repo_inspect, repo_search, change_impact, model_route) alongside noun-style names (agent_experience, context_pack, convergence_score, workspace_report).

Tool Count4/5

14 tools is within a healthy range for a repo-intelligence suite and each covers a defensible capability. There is mild redundancy since composite tools (review_verdict) are composed from others (change_impact, review_context, review_gate), so a few could be consolidated.

Completeness4/5

The surface covers indexing, search, mapping, impact analysis, review/gating, convergence scoring, context packing, model routing, and workspace reporting — a broad lifecycle for repo understanding. Gaps are minor: it is entirely read-only/advisory with no apply/edit or catalog-mutation management beyond repo_index.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Local-first code intelligence MCP server with hybrid BM25 + ONNX vector search, symbol-level impact analysis, diff-aware PR review with risk scoring, and persistent memory tied to git state.
    36
    132 npm
    80
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Local-first codebase context engine that parses code into a ranked dependency graph and serves it to AI tools via MCP for deep structural understanding.
    5
    19 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Turn your codebase into AI context — entirely on your machine. Single-binary MCP server with AST parsing, call graph, and local embeddings.
    26
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first MCP server that indexes all projects on your machine and provides coding agents with briefings about each project's stack, services, and run instructions, so they can start sessions with context across all repos. No cloud, no accounts, no telemetry.
    25 npm
    MIT