Skip to main content
Glama

Codebase Doctor

npm version npm downloads CI

It finds the thing, and it never guesses.

Models build. Codebase Doctor verifies.

Most repository scanners fail in one of two ways: they flood you with false positives, or they silently skip the hard case and report "clean." Codebase Doctor does neither. Every finding carries evidence, every audit reports what it couldn't analyze, and a clean run means the scope was actually checked.

npx -y codebase-doctor audit . --changed --format brief
codebase-doctor brief
scope=full findings=1 shown=1 coverage=incomplete
[high] security/secrets/provider-token test/unit/audits/ai/agent-surface.test.ts:83 —
  Have an authorized human or external coding agent remove the value and rotate it,
  then rerun the audit.
coverage-limitations: validation: skipped, database: skipped, security: partial,
  performance: unsupported

That last line is the point. Most tools print findings and stop. This one tells you what it didn't check, every single run.


Why this exists

Coding agents ship code faster than review can keep up. The failure isn't that agents can't write code — it's that nothing verifies what they wrote before it merges.

Codebase Doctor is that verification step. It builds a source graph, scans for committed secrets, dependency drift, unsafe Dockerfiles, workflow injection, RLS mistakes, and accessibility regressions — then tells you exactly what it could not verify.

Related MCP server: opzyai

Install

npx -y codebase-doctor audit .          # no install
npm install -g codebase-doctor         # global

Quick start

codebase-doctor audit . --json                   # full audit
codebase-doctor audit . --changed --json         # just my diff
codebase-doctor audit . --changed --base main --json    # PR review
codebase-doctor audit . --format sarif           # GitHub code scanning
codebase-doctor verify . --baseline before.json  # confirm fixes landed

Codebase Doctor terminal preview

Options

--run-checks          Permit configured validation commands
--changed             Audit Git changes and their affected scope (implied by review)
--base <ref>          Compare changed scope from the merge base with this ref
--json                Emit schema-versioned JSON
--format <format>     Output format: text, json, sarif, or brief (review adds markdown, github)
--all-findings        Review only: include findings outside the changed lines
--output <file>       Review only: write the report to a file as well as stdout
--exclude <glob>      Exclude a repository-relative path glob; repeatable
--baseline <path>     Compare with a prior Codebase Doctor JSON report
--timeout <ms>        Per-command timeout (default: 120000)
--fail-on <severity>  info|low|medium|high|critical|none (default: high)
--require-complete    Exit 2 when audit coverage is incomplete
--max-findings <n>    Cap brief output (default: 100)
--with-database       Permit live PostgreSQL catalog access
--with-advisories     Opt-in OSV advisory lookup over lockfile packages
--database-schema     Schema to inspect; repeatable (default: public)
--database-timeout    Catalog statement timeout in ms (default: 10000)

Code review

review is the pull-request command. It always audits changed scope, narrows findings to added diff lines, and prints a verdict — APPROVE, COMMENT, or REQUEST_CHANGES — so unrelated old issues never fail a PR:

codebase-doctor review . --base origin/main --format markdown > review.md
codebase-doctor review . --base origin/main --format github
codebase-doctor review . --format json  # includes a machine-readable review envelope
  • --format markdown renders a PR-comment-ready body with the verdict, findings, source impact, and coverage limitations.

  • --format github emits ::error / ::warning / ::notice workflow commands that annotate pull-request diffs inline from Actions logs, with no network access.

  • A finding on an unchanged line is out of scope for the verdict and counted as omitted; the full audit still reports it. --all-findings disables narrowing, and --output <file> writes the report to a file as well as stdout.

  • With --baseline, only new findings in the diff gate the verdict.

  • Exit 1 means the review requests changes; exit 2 is an operational failure, never a clean result. Inspect coverage before calling the reviewed diff verified or clean.

GitHub Action

permissions:
  contents: read
  security-events: write

steps:
  - uses: actions/checkout@v4
  - uses: subhajitlucky/codebase-doctor@v0.1.10
    with:
      format: sarif
      upload: "true"
      fail-on: high

Findings appear in the Security tab. See docs/github-action.md.


What it checks

Secrets — working tree and history

Precision-first and not exhaustive: it detects private-key material, provider-token shapes, paired AWS credentials, credential-bearing URLs, and high-confidence sensitive assignments. A Git-ignored .env is normal storage and is not a finding; a tracked one containing a real credential is.

security/secrets-history catches the case that matters most: a secret committed and later deleted from the working tree, so a rotated-looking repo doesn't hide an exposure. It inspects the most recent 200 commits across all branches without ever checking out, rewriting, or executing repository content.

Matched values are withheld from every finding, fingerprint, error, and report. Codebase Doctor never prints your secrets — including in its own SARIF. An external authorized human or agent must remediate the shareable content and rotate or revoke the credential, then rerun the same audit.

Dependencies

security/dependencies is read-only and offline. Lockfile-aware for npm lockfile versions 2 and 3, pnpm (v5+), Yarn (classic and Berry), and Bun. Python and other ecosystems remain explicitly unsupported rather than receiving guessed findings.

Rule families: security/dependencies/missing-lockfile, security/dependencies/manifest-lock-drift, security/dependencies/insecure-source, security/dependencies/mutable-git-source, security/dependencies/missing-integrity, security/dependencies/workspace-registry-resolution, security/dependencies/competing-npm-lockfiles, and security/dependencies/competing-lockfiles.

A normal semver range such as ^5.0.0 is not a finding when the lock agrees. Raw dependency specifications and resolved URLs are withheld from reports and never enter a fingerprint. An external authorized human or agent must correct the metadata and rerun the same scope. Inspect coverage before calling the dependency graph clean or verified.

It never invokes npm, another package manager, a shell, an installer, or a lifecycle script, and makes no network request. It makes no CVE or advisory claim on its own.

--with-advisories performs one bounded OSV lookup against resolved packages. Only names, versions, and ecosystems leave the machine.

Source impact — what breaks if I change this file

repository/source-graph uses a real syntax parser (never executes your code) to build a static import graph across JS/TS, Python, Go, Java, and Rust. In changed mode it walks reverse edges and reports a deterministic shortest impact path from each changed file.

Changed mode is mixed-scope per doctor, not a universal file filter: Project Doctor structural rules run with the full repository snapshot and may report findings outside changed paths for manifests, lockfiles, workspaces, and test visibility. Configured validation check plans are built from full project topology and then filtered to affectedProjectIds. Static SQL selects affected migration streams and replays full current history for every selected stream. Live database remains a full observed schema-set audit only with separately requested --with-database. Zero changed findings is not a full clean result.

repository/source-graph recognizes static import, re-export, type-only import, literal require, and literal dynamic import edges across JavaScript and TypeScript (plus Python, Go, Java, and Rust) with a real syntax parser that never executes repository code. Cycles are valid topology, not findings, and this module is finding-free by design.

A separate precision-first repository/source-integrity Doctor emits only the source/import-target-missing rule, keeping topology limitations from becoming guessed bugs. It diagnoses only four proof classes: an explicit relative target with a supported source extension; a single deterministic alias whose configured target explicitly names a supported source file; a unique workspace package whose explicit entry names a supported source file; and an internal Go package under a module path without a replace directive.

Extensionless, JSON, custom-loader, conditional, ambiguous, external, and dynamic references and cycles are not findings. It does not check named exports or validate that a referenced export name exists.

Full mode examines all qualifying edges; changed mode examines changed importers and complete reverse-impacted importers. A deleted or renamed target selects its unchanged importer.

It emits at most 1,000 findings per audit and reports partial coverage whenever that ceiling or any upstream graph limitation applies. Partial coverage is not a clean source-integrity result. Raw import specifiers and source text are withheld from findings, which expose only normalized paths, import kind, proof class, and safe location.

An external authorized human or agent must correct or restore the intended target and rerun the same scope.

codebase-doctor audit . --changed --base main
src/db/schema.ts → src/repositories/user.ts → src/api/users/[id]/route.ts
→ src/app/dashboard/page.tsx → tests/integration/user.test.ts

Cycles are valid topology, not findings. The graph module intentionally emits no bug findings — repository/source-graph is finding-free by design, and the separate precision-first repository/source-integrity Doctor reports only provably missing import targets, so topology limits never become guessed bugs.

Schema-1 reports may include sourceImpact (schema 1). Changed mode walks reverse internal edges, adds impacted projects to affectedProjectIds, and reports a deterministic shortest impact path per changed source root. Reports preserve full impacted-file counts while serializing only bounded impact records. A path proves only the static edge chain, not a bug in the dependant. Raw import specifiers and source text are withheld; the module uses no plugins, network requests, or writes.

Local tsconfig and jsconfig files contribute a deterministic subset of relative aliases; this is not complete Node or TypeScript module resolution. Dynamic non-literal imports, ambiguous targets, unsupported configuration or syntax, unreadable input, and graph ceilings are coverage limitations, not findings.

Workflow and infrastructure

script-injection (attacker-controlled ${{ github.event.* }} in a run: step), pull-request-target-checkout, write-all-permissions, unpinned action refs, unpinned Docker base images, pipe-to-shell, root-user, and world-writable files. Workflows are never dispatched and images are never built.

PostgreSQL and Supabase RLS

Offline database/sql-rls runs automatically when a supported PostgreSQL migration stream is discovered: it requires no credentials, makes no network request, and never executes migration SQL, reconstructing expected table, policy, RLS, and grant state from supported migrations. Partial coverage is not a clean static SQL result.

Live database/rls inspects observed database state through a read-only catalog of policies, privileges, roles, memberships, enforcement, and bypass paths, permissioned separately with --with-database using environment credentials and a read-only, repeatable-read transaction.

database/rls-drift compares the two — expected migration state against observed live state — and reports table-missing-live, rls-disabled-live, force-rls-disabled-live, policy-missing-live, grant-missing-live, rls-enabled-live-only, and policy-unmanaged-live.

Database and Drizzle hazards

The read-only, offline database/drizzle module and its database/drizzle/raw-sql-date-parameter rule catch a runtime boundary: a JavaScript Date interpolated into a raw Drizzle sql template can bypass the column's timestamp encoder, so postgres-js may throw ERR_INVALID_ARG_TYPE, while equivalent SQL can still work in psql.

// Before: raw interpolation can bypass the timestamp column encoder.
const rows = await db.execute(sql`select * from jobs where run_at <= ${date}`);

// After: guidance for a human or separately authorized external coding agent.
const rows = await db.select().from(jobs).where(lte(jobs.runAt, date));

Applicability requires an exact drizzle-orm/postgres-js adapter import, or scoped owning/workspace evidence for both drizzle-orm and postgres. It reports only statically proven Date flows and never infers from a variable name. Findings are medium severity, high confidence.

Not findings: Date(), Date.now(), an encoded toISOString() string, typed comparisons such as lte(column, date), and a fresh inline encoder object passed directly to sql.param(value, encoder). Encoder identifiers and aliases are not statically proven safe even when declared const, because their objects may be mutated elsewhere; those interpolations and unresolved flows become partial coverage limitations rather than guessed findings. Partial coverage is not a clean Drizzle audit. Raw SQL and parameter values are withheld from findings, fingerprints, and reports. An external authorized human or agent must make the repair and rerun the same scope.

Agent surface — the newest attack target

Audits the agent configuration surface without executing or contacting any of it:

  • MCP client configs: unpinned package runners, shell commands, inline credentials, broad filesystem grants

  • SKILL.md: unscoped allowed-tools grants (Bash(*), bare Write)

  • Permission bypass: bypassPermissions, --dangerously-skip-permissions, --yolo, yes-always, broad permissions.allow

Frontend

JSX/TSX and static HTML accessibility (img-missing-alt, iframe-missing-title, html-missing-lang, positive-tabindex) and static SEO (missing-title, missing-meta-description). No browser, no build.

Backend and auth

backend/auth is read-only and offline over JavaScript and TypeScript sources. It never starts a server, sends a request, or issues a token — it reads source text only.

Rules: cors-wildcard-origin-with-credentials (wildcard origin: "*", reflected origin: true, or an allowlist containing *, with credentials enabled), session-cookie-security-disabled (cookie secure or httpOnly explicitly false), jwt-decode-without-verify (a decode call in a file containing no verify call), and jwt-verify-algorithm-unrestricted (no algorithms allowlist).

A rule fires only when the callee provably resolves to the audited package through an import declaration or CommonJS require, so an unrelated local helper named cors or decode is never reported. Configuration that cannot be resolved statically — a non-literal options expression, a computed cookie flag, or a spread property that could supply the value — is a coverage limitation, never a guessed finding, so inspect backend coverage before calling a codebase clean. The decode and algorithm rules are file-scoped: a verify call in middleware in another file does not suppress them. Configured origin and secret literals are withheld from reports and never enter a fingerprint.

Not covered: API shape, worker, webhook, cron, and rate-limit analysis. An external authorized human or agent corrects the configuration, then reruns the same scope.


Current coverage versus north star

This is the part most scanners skip.

There is one unified auditor — one doctor for the whole codebase, not a collection of separate products. Framework- and domain-specific knowledge lives inside it as built-in internal audit modules.

A full audit examines the full requested repository scope for applicable implemented modules. It is not complete or universal — it is not every-domain analyzer coverage. Inspect coverage before calling a codebase verified or clean.

Every report includes domainCoverage — a checklist of nine domains separating applicability from status, so not-detected differs from detected-but-unsupported, skipped, failed, or not-selected, with module-level status details, evidence, and limitations. coverageComplete does not mean the code is bug-free or correct.

That means:

  • A clean run means the scope was actually checked.

  • --require-complete exits 2 rather than letting a skipped area report as clean.

  • A truncated or bounded scan says so in coverageSummary with exact total / emitted / omitted counts and deterministic sample paths.

Domain

Current source coverage

North star

Repository structure

Inventory, framework detection, manifests, workspaces, lockfiles, test visibility, JS/TS + Python + Go + Java + Rust impact graph

Cross-language dependency and behavioral topology

Configured validation

JS/TS and Python command planning; execution only with --run-checks

Sandboxed validation across ecosystems

Database

Offline migration RLS, Drizzle Date hazards, live RLS, static-to-live drift

Schemas, queries, permissions, more engines

Frontend

JSX/HTML a11y and static-HTML SEO

React, Next.js, bundle analysis, broader a11y

Backend and authz

Read-only, offline backend/auth analysis of CORS, session-cookie, and JWT hazards in JS/TS; NestJS detection

API, worker, webhook, cron, rate-limit analysis

Security

Secrets (tree + history), dependency rules, opt-in OSV

Secrets, permission, vulnerability, supply chain

Infrastructure

Dockerfile and GitHub Actions

Hosting and deployment analysis

Performance

No semantic analyzer

Cache, query, memory, profiling

AI systems

Agent-surface audit: MCP configs, SKILL.md grants, permission settings

Prompt, token, grounding analysis

North-star entries are planned modules, not shipped behavior. Built-in source-impact graph, secrets analysis, and dependency analysis ship together in 0.1.4 and are not part of the historical 0.1.3 package.

Precision and bounded-report contract

Workspace publication entries, generated targets, and fixture-controlled paths are coverage limitations unless independently proven broken; they are not missing-target findings by themselves. Detected pnpm, Yarn, and Bun scopes never receive npm-specific findings. Only a cryptographic match to an inventoried localhost-only certificate can classify a private key as an intentional local test key; every other matched private key remains high severity.

Schema-1 reports bound repeated evidence without hiding its size: coverageSummary preserves exact total, emitted, and omitted counts, and limitationGroups preserve each reason, deterministic sample paths, and omitted path counts.

Inspect coverage before calling a codebase verified or clean. Read docs/architecture.md for the full contract.

Read-only by design

Codebase Doctor reports. It exposes no direct target-file write API, has no direct filesystem-write capability, and includes no remediation executor. It can never be granted direct target-write or remediation authority, and never modifies, fixes, or repairs target files. A human or a separately authorized agent makes changes, then reruns the same scope to verify.

Separately authorized --run-checks launches repository-owned validation subprocesses; they are not filesystem- or network-isolated and may have side effects. That is validation execution, not Doctor repair authority.

  • --changed grants no command execution, network, or database access

  • validation commands need --run-checks; live database needs --with-database; OSV lookup needs --with-advisories

  • database credentials come from DATABASE_URL or SUPABASE_DB_URL, never a connection-string flag

  • source analysis parses syntax and never executes source; lockfile analysis never invokes a package manager

  • apart from an explicitly requested OSV lookup, it makes no external network calls

Exit codes

Code

Meaning

0

Requested audits completed and no finding met the threshold

1

Requested audits completed and at least one finding met the threshold

2

An audit could not complete, or coverage was incomplete under --require-complete

Exit 2 is an operational failure, not a clean result.

Baselines and SARIF

--baseline classifies fingerprints as new, unchanged, or resolved, and applies --fail-on only to new findings. After an external fix, confirm it:

codebase-doctor audit . --json > before.json
# ... fix happens elsewhere ...
codebase-doctor verify . --baseline before.json

verify reports each fingerprint as resolved, unchanged, unresolved, or new, and exits 1 unless everything is verifiably resolved. unresolved means absent under incomplete coverage — never a repair.

MCP server and agents

claude mcp add codebase-doctor -- npx -y codebase-doctor mcp

Read-only tools: audit_codebase, verify_changes, explain_finding, describe_capabilities. Responses are bounded at roughly 50 KB; the server never enables --run-checks or live database access.

Registry metadata ships in server.json (io.github.subhajitlucky/codebase-doctor); publishing steps for the official MCP registry, Smithery, and Glama are in docs/mcp-registries.md.

Live listings: Glama · official MCP registry (io.github.subhajitlucky/codebase-doctor).

A Claude Code plugin ships in this repository (.claude-plugin/ + skills/):

/plugin marketplace add subhajitlucky/codebase-doctor
/plugin install codebase-doctor

codebase-doctor instructions prints ready-to-paste snippets for AGENTS.md, CLAUDE.md, .cursor/rules/, .windsurfrules/, .clinerules/, and copilot instructions. It only prints — it never writes files.

Workflow: audit . --changed --format brief after edits, review . --base main --format brief for a PR verdict, full audit . at trust boundaries, verify after an external fix.

Roadmap

  • Built-in backend, performance, and AI semantic audit coverage

  • Per-domain coverage guarantees beyond the global --require-complete gate

  • Pull-request annotations, hooks, and agent plugins on the same report schema

  • Approved validation in read-only mounts or disposable copies

  • Cross-model benchmarks: defects found, false positives, verification success, token cost

Development

npm install && npm run build && npm run typecheck && npm test

Releases are checked with npm pack, installed into a clean temporary project, and run through the generated binary.

License

MIT

Available Tools

4 tools
audit_codebaseA
Read-onlyIdempotent

Run the full built-in Codebase Doctor audit on a repository and return the evidence-backed report. Read-only and offline by default; it never enables validation commands (--run-checks) or live database access (--with-database).

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoGit ref to compare against from the merge base; requires changed. Mirrors the CLI --base option.
pathNoRepository path to audit. Defaults to the server working directory.
formatNoReport rendering: "json" returns the schema-version-1 JSON report; "summary" returns the deterministic text report.
changedNoAudit Git changes (staged, unstaged, untracked, and branch work) and their selected scope instead of the full repository.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly/idempotent/non-destructive/closed-world, so the description earns credit by adding what those hints cannot: the audit is offline by default and will never enable validation commands (--run-checks) or live database access (--with-database). It stops short of disclosing runtime cost, repo-size limits, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero padding. The primary action is front-loaded and the safety constraints follow immediately, and every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema the description still conveys the return artifact and the two rendering options, and all four optional parameters are documented in the schema. It omits what the audit actually checks and how large a repo it can handle, minor gaps for a read-only tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already explained in the schema (base ref, path default, format enum, changed scope). The description adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run the full built-in Codebase Doctor audit on a repository') and names the artifact returned ('the evidence-backed report'). It is easy to distinguish from siblings like verify_changes or explain_finding, though the description never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The default-scope statement ('Read-only and offline by default') and the two explicitly disabled modes give implicit guidance on how the tool behaves, but there is no explicit when-to-use statement or routing advice versus verify_changes/describe_capabilities. Usage must be inferred from the schema's 'changed' flag.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_capabilitiesA
Read-onlyIdempotent

Describe this MCP server: available tools, the nine audit domains in the domainCoverage inventory, the Doctor capability vocabulary, and the permissions this server never grants.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and a closed-world posture, so the safety profile is covered. The description adds genuinely extra context by promising disclosure of 'the permissions this server never grants,' which tells the agent about hard capability boundaries rather than just restating read-only behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; each listed item (tools, audit domains, capability vocabulary, never-granted permissions) is load-bearing information about what the caller receives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description carries the burden of describing the return content and does so by enumerating the four sections the response contains. It is nearly complete for a low-complexity read-only meta tool; only the guidance on when to invoke it is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters and 100% schema coverage, so there is nothing for the description to clarify at the input level. Baseline 4 applies for a no-argument tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Describe') and resource ('this MCP server'), then enumerates the exact content returned (tools, nine audit domains, Doctor capability vocabulary, never-granted permissions). It is clearly distinct from the audit/verify/explain siblings, though it never names them or explicitly contrasts itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement, no mention that this is the recommended first call for orientation, and no exclusions. Usage is only implied by the self-descriptive nature of the tool (an agent wanting an inventory of the server would infer this is the right call).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_findingA
Read-onlyIdempotent

Return the full evidence, remediation guidance, and verification command for one finding, selected by fingerprint or rule id. Read-only and offline.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoGit ref to compare against from the merge base; requires changed.
pathNoRepository path to audit. Defaults to the server working directory.
ruleIdNoRule id, for example database/drizzle/raw-sql-date-parameter. Returns the highest-severity match.
changedNoAudit Git changes and their selected scope instead of the full repository.
fingerprintNoExact finding fingerprint from a prior JSON report.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered; the description adds that it is 'offline' and enumerates the returned artifacts (evidence, remediation, verification command), which is useful since no output schema exists. It doesn't cover the no-match or ambiguity case (e.g., what happens when a rule id matches multiple findings).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler; the primary purpose and the selection criterion are both front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does carry the return-value burden and handles it by naming the three artifacts returned. It remains silent on failure modes (unknown fingerprint, rule id with no match) and whether results are truncated, which leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including the 'highest-severity match' behavior for ruleId. The description only reinforces the fingerprint/ruleId selection axis and adds nothing the schema lacks, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (return), resource (evidence, remediation guidance, verification command for one finding), and selector (fingerprint or rule id). This clearly distinguishes it from sibling tools audit_codebase and verify_changes, which perform audits rather than explain a single result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The selection mechanism ('by fingerprint or rule id') and the provenance of the input ('from a prior JSON report' per the schema) make the intended usage clear. However, it never explicitly contrasts itself with audit_codebase or verify_changes or states when *not* to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_changesA
Read-onlyIdempotent

Verify that findings from a prior schema-1 JSON report are repaired. Runs a fresh read-only audit, compares fingerprints, and reports each baseline finding as resolved, unchanged, or unresolved. Absence under incomplete coverage is unresolved, never resolved. Read-only and offline by default; it never enables validation commands or live database access.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoGit ref to compare against from the merge base; requires changed.
pathNoRepository path to verify. Defaults to the server working directory.
formatNoReport rendering: "json" returns the structured verification result; "summary" returns the deterministic text report.
changedNoVerify against Git changes and their selected scope instead of the full repository.
baselineYesPath to a prior codebase-doctor schema-1 JSON report on this machine.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world; the description goes further by disclosing fingerprint comparison, the freshness of the audit, the conservative coverage rule (absence under incomplete coverage is unresolved, never resolved), and that it never enables validation commands or live DB access.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four front-loaded sentences, each carrying distinct load: purpose, mechanism and outputs, the interpretive rule, and the safety posture. No filler or redundant restatement of annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the result vocabulary (resolved, unchanged, unresolved) and the coverage caveat, so an agent knows what comes back and how to read it. Nothing needed to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline 3 applies, but the description adds real meaning: baseline is a prior codebase-doctor schema-1 JSON report on the machine, and the coverage caveat implicitly frames how results should be interpreted. This is slightly more than the schema conveys, though it does not detail base/changed/path interactions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (findings from a prior schema-1 JSON report being repaired), plus the mechanism (fresh read-only audit, fingerprint comparison). This is clearly distinguishable from audit_codebase's fresh audit because it is explicitly baseline-relative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The requirement of a prior baseline report plus the resolved/unchanged/unresolved framing makes the post-repair verification context clear. It stops short of explicitly naming audit_codebase or state a when-not condition, so it lands just below a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedaudit_codebase
    • First observeddescribe_capabilities
    • First observedexplain_finding
    • First observedverify_changes

TDQS

A4.1/5.0

Scored across 4 tools

Disambiguation4/5

Each tool has a fairly distinct purpose: audit_codebase runs the full audit, verify_changes compares fingerprints against a baseline, explain_finding drills into one finding, and describe_capabilities is meta. The only mild overlap is that audit_codebase and verify_changes both execute a read-only audit, but their descriptions make the baseline-vs-verification distinction clear.

Naming Consistency5/5

All four tools follow a clean verb_noun snake_case pattern (audit_codebase, describe_capabilities, verify_changes, explain_finding), with no mixed conventions or vague verbs.

Tool Count4/5

Four tools is a tightly scoped, coherent set for a read-only audit workflow (run, explain, verify, describe). It is on the lean side but each tool earns its place, so it lands slightly below ideal breadth rather than being bloated.

Completeness4/5

The surface covers the core audit lifecycle: run an audit, inspect individual findings, verify repairs, and introspect the server's capabilities. A dedicated findings-listing tool is absent, but findings are surfaced through the report and explain_finding, so agents can work around it; read-only design legitimately omits write operations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Live codebase intelligence for AI agents. Import graph PageRank for file importance, git forensics for co-change coupling and fragile code, convention detection across 16 domains, and blast radius analysis.
    47 npm
    3
    Business Source 1.1
  • A
    license
    Not graded
    quality
    D
    maintenance
    Local-first security check for AI coding agents — finds hardcoded secrets, exposed .env files, git-history leaks and vulnerable dependencies (OSV), entirely on your machine. Ask your agent "is this safe to ship?" and get a Launch Readiness score with a fix for every finding.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Exposes repository reconnaissance tools — file listing and verbatim reading, regex search, dependency-tree parsing across Python/Node/Rust/Go, and test-coverage ingestion — so agents can gather line-numbered evidence from real source files. This lets an audit pipeline physically re-verify every claim and drop findings that cannot be located in the code.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local code-graph engine for AI coding workflows: static analysis of Rust, Cangjie, ArkTS, TypeScript/JavaScript, C/C++, Python, and Shell projects into symbols, imports, call graphs, and quality gates, with pre-commit change review and impact analysis. Read-only by design — never executes build scripts or uploads code.
    6
    1
    MIT