Skip to main content
Glama

Code Guardian

Build Status License: MIT Node.js >= 18 Tests MCP Ready

Stop shipping code that isn't ready for production. Code Guardian audits your project against industry standards, generates production-ready patterns, and enforces quality gates — across every AI coding agent.

What is this?

You're building an app. You ask an AI agent to write code. It delivers something that works — but doesn't have tests, error handling, CI/CD, security headers, or proper logging. You ship it, and three days later production burns down.

Code Guardian exists to fix that.

It's an MCP plugin that gives every AI coding agent (Claude Code, Cursor, Windsurf, Devin, Codex, Gemini, Antigravity) built-in knowledge of what production-ready code actually looks like. Before any feature is implemented, it audits the project. After any change, it checks whether it meets industry standards. And when you need to start fresh, it generates complete project scaffolds with zero guesswork.

Think of it as a senior engineer who reviews every line of code the AI writes — but automated, instant, and consistent.

Related MCP server: Inspectra

Supported Agents

Works with every major AI coding agent out of the box:

Agent

Integration

How it works

Claude Code

Plugin / MCP Server

Auto-installed via plugin marketplace or manual stdio config

Cursor

MCP Server

Add to ~/.cursor/mcp.json

Windsurf

MCP Server

Add to ~/.windsurf/mcp.json

Devin

Context

Use output directly — paste audit results into Devin's context

Codex

CLI

Run standalone and pipe output into Codex sessions

Gemini

Context

Set expectations in system prompt with audit findings

Antigravity

Context

Define requirements explicitly; run audit as pre-step

Quick Start

Claude Code — Two ways to install

Option 1: npm link (Recommended)

# Clone the repo
git clone https://github.com/justin-coders/code-guardian.git
cd code-guardian

# Build (no dependencies, but validate the entry point)
node --check src/stdio-server.js

# Install as a global CLI so Claude Code can find it
npm link

# In Claude Code, enable the plugin:
#   /plugin marketplace add justin-coders/code-guardian
#   /plugin install

Option 2: Manual MCP config

Add this to your ~/.claude/settings.json under mcpServers:

{
  "code-guardian": {
    "type": "stdio",
    "command": "node",
    "args": ["/absolute/path/to/code-guardian/src/stdio-server.js"]
  }
}

Then restart Claude Code.

After Installation

The plugin loads automatically. Just ask:

"Audit this project for production readiness"
"Generate a production REST API with NestJS and Zod validation"
"Show me the industry standard for JWT authentication"
"Create a GitHub Actions CI/CD pipeline for my Express app"

Cursor — Quick add

Add one entry to your ~/.cursor/mcp.json:

{
  "mcpServers": {
    "code-guardian": {
      "command": "node",
      "args": ["/absolute/path/to/code-guardian/src/stdio-server.js"]
    }
  }
}

Restart Cursor. The tools are now available in your prompts.

Windsurf — Same as Cursor

Add the same MCP config to ~/.windsurf/mcp.json.

Other agents (Devin, Codex, Gemini, Antigravity)

These agents don't support MCP plugins directly. Use Code Guardian as a standalone tool — clone the repo and run it, then paste the output into your agent's context, or follow the agent-specific guidance in INSTALLATION.md.

Updating an existing install

cd /path/to/code-guardian
git pull origin dev    # or your branch

# If using npm link:
npm unlink && npm link

# If using manual MCP config: no action needed — path still points to repo

# Restart your agent
# Claude Code: close and reopen
# Cursor/Windsurf: restart the editor

What It Does

1. Audit Everything

Run a full production-readiness audit on any project in seconds:

> "Audit this project"

Covers 10 dimensions and gives you an A–F grade with specific remediation steps for every failure:

Dimension

What it checks

package.json

Scripts, engines, name, version

Lock file

Dependency pinning for reproducible builds

README

Documentation quality

.gitignore

Sensitive files excluded

ESLint

Code quality enforcement

TypeScript strict mode

Type safety settings

Tests

Coverage and framework config

CI/CD

Pipeline detection

.env safety

Secret management

Build script

Deployment readiness

2. Generate Production-Ready Code

Stop copying Stack Overflow snippets. Get templates that follow industry conventions:

> "Generate a production REST API with NestJS"
> "Create a JWT auth module for Express"
> "Generate a Dockerfile for my NestJS app"

Available templates:

Feature

Frameworks

What's included

REST API

NestJS, Express, Fastify

Controllers, DTOs, validation, Swagger docs, pagination

Authentication

NestJS, Express

JWT, refresh tokens, RBAC, MFA, brute-force protection

Database

Prisma, TypeORM, Knex

Migrations, connection pooling, soft deletes, indexing

Error Handling

All frameworks

Global handlers, custom error classes, structured logging

Logging

Winston, Pino

JSON structured logs, correlation IDs, log levels

Security

All frameworks

OWASP Top 10, helmet, CORS, rate limiting, input sanitization

CI/CD

GitHub Actions

Lint → test → build → deploy pipeline with caching

Docker

Docker

Multi-stage builds, non-root user, healthchecks, compose

Testing

Jest, Vitest

80% coverage thresholds, Arrange-Act-Assert, mock strategies

Branch Strategy

Git

Feature/bugfix/hotfix/release patterns, PR requirements

3. Enforce Standards

Every tool includes remediation guidance — not just "you failed X", but "here's exactly how to fix it":

❌ Missing ESLint
→ Add eslint.config.js with: strict mode, no-unused-vars, prefer-const, no-console in prod

❌ No TypeScript strict mode
→ Enable: strict: true, noUncheckedIndexedAccess: true, exactOptionalPropertyTypes: true

❌ No CI pipeline
→ Add .github/workflows/ci.yml with: lint → test → type-check on PR, deploy on main merge

4. Know Your Agent

Auto-detects which AI agent you're using and surfaces the right best practices:

> "What's the best way to use Code Guardian with Cursor?"
> "Detect which agent I'm using"

Each agent gets tailored guidance on plugin structure, configuration files, and workflows.

10 Industry Pattern Categories

Built-in knowledge of production standards across every major concern:

  1. REST API — Correct HTTP status codes, schema validation, pagination, OpenAPI/Swagger docs

  2. Authentication — JWT with short expiry, RBAC, MFA, token rotation, CSRF protection, brute-force mitigation

  3. Database — Versioned migrations, connection pooling, N+1 query prevention, soft deletes, read replicas

  4. Testing — 80%+ coverage targets, test pyramid (unit/integration/E2E), Arrange-Act-Assert, mock strategies

  5. Error Handling — Centralized global handlers, custom error classes, consistent response shape, Sentry integration

  6. Logging — Structured JSON logs, correlation/request IDs, log levels per environment, centralized aggregation

  7. Security — OWASP Top 10 compliance, helmet.js, CORS whitelists, rate limiting, parameterized queries only

  8. CI/CD — Parallel jobs, node_modules caching, blue-green deployment, rollback scripts, failure notifications

  9. Docker — Multi-stage builds, Alpine/Distroless bases, non-root user, healthchecks, resource limits

  10. Branch Strategy — Git Flow / trunk-based conventions, PR requirements, semantic versioning, squash vs merge rules

File Structure

code-guardian/
├── src/
│   ├── tools.js              # 15 tools + 10 industry patterns + code templates
│   ├── stdio-server.js       # stdio MCP server (Claude Code, Cursor, Windsurf)
│   └── http-server.js        # HTTP/SSE server (port 8765)
├── tests/
│   ├── tools.test.js         # 57 unit tests
│   └── integration.test.js   # 26 integration tests
├── .claude-plugin/
│   └── plugin.json           # Plugin manifest
├── .mcp.json                 # Dual transport config (stdio + HTTP)
├── package.json              # Zero dependencies
├── LICENSE                   # MIT
├── README.md
├── INSTALLATION.md
├── CONTRIBUTING.md
├── UPCOMING_FEATURES.md
└── CHANGELOG.md

Zero dependencies. Pure ESM Node.js. No npm install required.

Complete Tool Reference

Tool

Command

Description

audit_codebase

audit_codebase({ cwd })

Full 10-dimension audit with A-F grade and remediation

check_branch

check_branch({ cwd })

Verify git branch naming conventions

check_tests

check_tests({ cwd })

Discover test framework and run execution check

check_cicd

check_cicd({ cwd })

Detect CI/CD configs and package scripts

check_linting

check_linting({ cwd })

Find linters and run ESLint if present

check_security

check_security({ cwd })

Scan for secrets + run npm audit

check_architecture

check_architecture({ cwd, depth })

Review directory structure and monorepo setup

production_readiness

production_readiness({ cwd })

A-F scorecard with per-item guidance

generate_production_code

generate_production_code({ feature, stack })

Production-ready code templates

get_industry_patterns

get_industry_patterns({ category })

Full checklist for any architectural concern

generate_security_checklist

generate_security_checklist({ cwd })

OWASP Top 10 tailored to your project

generate_github_workflow

generate_github_workflow({ stack, deployTarget })

CI/CD YAML generation

detect_agent

detect_agent({ cwd })

Auto-detect your AI coding agent

get_agent_guidance

get_agent_guidance({ agent })

Agent-specific best practices

generate_starter_repo

generate_starter_repo({ framework })

Complete project scaffold with production defaults

Production Readiness Grading

Grade

Score

What it means

A

90–100%

Production-ready. Ship it.

B

75–89%

Good. Address the warnings before deploying.

C

60–74%

Partial. Significant gaps need fixing.

D

40–59%

Below standard. Don't ship without major work.

F

0–39%

Not production-ready. Comprehensive remediation needed.

Development

# Run all tests
node --test tests/**/*.test.js

# Run with coverage report
node --test --experimental-test-coverage tests/**/*.test.js

# Syntax check
node --check src/tools.js
node --check src/stdio-server.js
node --check src/http-server.js

83 tests passing. 57 unit tests + 26 integration tests covering all 15 tools.

Adding a New Tool

  1. Implement in src/tools.js with defensive args (args = {})

  2. Add to TOOLS array in both src/stdio-server.js and src/http-server.js

  3. Add switch case in handleRequest() / dispatchTool()

  4. Write unit tests in tests/tools.test.js and integration tests in tests/integration.test.js

  5. Document in README.md tools table

See CONTRIBUTING.md for the full guide.

License

MIT — use it freely in personal and commercial projects.

Built With


Made with care by justin-coders. Questions? Open an issue.

Copyright (c) 2026 justin-coders

Available Tools

15 tools
audit_codebaseB

Full audit of a codebase against industry standards: linting, TypeScript, testing, CI/CD, docs, dependencies. Includes actionable guidance.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoProject root directory (defaults to current working directory)
depthNoDirectory scan depth (default: 3)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses scope and that output includes actionable guidance, but says nothing about side effects, whether files are modified, runtime cost, or what a full audit actually returns. This is thin for a tool without annotation safety signals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The verb and scope come first, the coverage areas are enumerated compactly, and the final clause about actionable guidance earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a broad audit tool: it names the audit dimensions and hints at output quality. However, with no output schema and no annotations, an agent still lacks detail on return format, execution behavior, and when to choose this over the more specific sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents both parameters with 100% coverage, including defaults and intended meaning. The description adds no parameter-level detail, but the high schema coverage justifies the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states a specific verb and resource: 'Full audit of a codebase' with explicit coverage areas (linting, TypeScript, testing, CI/CD, docs, dependencies). It does not explicitly name or contrast the granular check_* siblings, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call audit_codebase versus check_tests, check_linting, check_cicd, or production_readiness. The word 'Full' implies a comprehensive sweep, but no explicit when-to-use, exclusions, or alternatives are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_architectureC

Review directory structure, monorepo indicators, and architectural conventions.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
depthNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden of explaining behavior. It only says 'Review', which suggests a read-only inspection, but it does not disclose what is returned, whether any analysis is performed, whether it scans the filesystem, or what side effects might occur. Beyond the verb 'review', it adds little behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler words. It is not redundant with the schema, and every word adds meaning. However, it is slightly under-specified for the number of parameters and behavioral context needed, so it earns a high but not perfect conciseness score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a minimal tool with optional parameters and no output schema, so the description must do more to be self-sufficient. It tells the agent what is reviewed but not how to use cwd/depth, what the output format is, or what happens on failure. For an agent to invoke this tool correctly with confidence, the context is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the undocumented cwd and depth parameters. It does not mention either parameter, nor does it clarify how 'depth' interacts with directory traversal or what cwd defaults to. The parameter names are somewhat self-explanatory, but the description provides no direct semantic support.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb, 'Review', and names concrete targets: 'directory structure, monorepo indicators, and architectural conventions'. This clearly separates it from siblings like check_tests or check_linting, though 'architectural conventions' remains somewhat vague and no explicit mention of alternatives is made.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as audit_codebase or check_branch. The generic 'Review ... architecture' wording implies an architecture-focused use case, but there are no exclusions, prerequisites, or stated conditions for choosing it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_branchB

Check that the current git branch follows conventional naming (feature/, bugfix/, hotfix/, release/, main, master, develop).

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoGit repository root

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state whether the check is read-only, what side effects may occur, how failures are reported, or what the return value looks like. The only behavioral detail is the list of accepted naming patterns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that includes the action, the target resource, and the complete set of naming rules. It is front-loaded and contains no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, but with no annotations and no output schema, the description should explain what a successful or failed check looks like. It specifies naming conventions but omits the result or behavior after invocation, leaving the agent unable to predict the tool's response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter cwd is fully documented in the schema with the description 'Git repository root', giving 100% schema coverage. The tool description adds no additional parameter semantics, but the baseline holds because the schema already provides the necessary meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: check the current git branch, and it enumerates the naming patterns that are accepted. This distinguishes it from sibling tools like check_tests, check_linting, and check_security by scoping validation to branch naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the intended use case: validating branch naming conventions, but it does not explicitly state when to use this tool over sibling check tools or when not to use it. There are no exclusions, alternatives, or workflow context provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_cicdB

Detect CI/CD configuration files (GitHub Actions, GitLab CI, Jenkins, CircleCI, Docker, pre-commit).

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoProject root

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral disclosure burden. It states what it detects but does not say whether it returns a list of found files, a boolean, whether it parses file contents, or what exit/status behavior to expect. For a tool invoked by an agent, this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, lists the relevant file types, and contains no filler or redundancy. It is lean and every word contributes to understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple: one optional parameter and no output schema. The description covers the purpose and scope, but omits what the return value looks like and whether detection means existence only or content inspection. It is usable but leaves the agent to guess at the result shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter 'cwd' is fully documented in the input schema as 'Project root', giving 100% schema description coverage. The tool description adds no additional parameter meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Detect') and resource ('CI/CD configuration files') and enumerates the concrete types covered (GitHub Actions, GitLab CI, Jenkins, CircleCI, Docker, pre-commit). This makes the tool's purpose immediately clear and distinguishes it from sibling check tools that target tests, branches, or linting.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over alternatives, no preconditions, and no exclusions. An agent must infer that it is for checking the presence of CI/CD configs, but no explicit when-to-use or alternative-selection context is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_lintingA

Detect installed linters (ESLint, Prettier, Biome, oxlint, stylelint) and run ESLint if available.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoProject root

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits. It states it will run ESLint if available, which is a side-effectful action, but does not clarify whether it modifies files (e.g., --fix) or only reports issues. It also does not specify what happens if no linters are installed or what the return format is. The description is honest but incomplete regarding side effects and outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the core action (detect linters) and then adds the conditional ESLint execution. There is no redundant or unnecessary information, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and no annotations, so the description must explain both the action and the expected result. It explains the action but not what the tool returns (e.g., linting errors, a summary) or how an agent should interpret the output. It also lacks details on failure scenarios (e.g., no linters found). This leaves significant gaps for an agent deciding how to proceed after calling the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single parameter 'cwd' fully with description 'Project root' (100% coverage). The tool description does not add any additional meaning about this parameter, so it does not go beyond the schema. Baseline 3 is appropriate as the schema carries the semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: detect installed linters (explicitly naming five) and run ESLint if available. It uses a specific verb (detect, run) and resource (linters, ESLint), and distinguishes itself from sibling tools like check_tests or check_security by focusing solely on linting.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when linting checks are needed) but does not explicitly contrast with alternatives or provide exclusions. It gives clear context that this tool is for linting, which is sufficient to differentiate from siblings, but lacks explicit routing guidance like 'use check_tests for tests'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_securityB

Scan for security concerns: .env files at root, dependency vulnerabilities via npm audit.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoProject root

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It lists the specific checks performed, but does not disclose whether the tool is read-only, whether it requires network access (for npm audit), what it returns, or if it can modify files. This is a significant gap for a tool that likely runs external commands.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the action and then specifies the scope. Every word earns its place, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter, no output schema, and no annotations, the description is insufficient. It explains what is scanned but not what the tool returns, whether it is safe to run without side effects, or how failures are reported. An agent could call it but would lack critical information to interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100% for the single parameter 'cwd', which is described as 'Project root'. The description adds no extra meaning beyond what the schema already provides, so the baseline of 3 is appropriate. It does not elaborate on how cwd affects the scan or any default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Scan' and the resource 'security concerns', and specifies two concrete checks: .env files at root and dependency vulnerabilities via npm audit. This makes it easily distinguishable from siblings like check_tests or check_linting, which target different concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives like audit_codebase or production_readiness, nor does it mention any prerequisites or exclusions. The purpose is clear, but the context for when to choose it over similar security-related siblings is only implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_testsA

Discover test files, detect test framework, and run a quick test execution check.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoProject root

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral disclosure burden. It transparently states that it discovers/detects files and runs a test execution, but it does not mention potential side effects of running tests, prerequisites, or what 'quick' means. This is adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with three clear actions and no filler words. Every part earns its place, and the structure makes the tool's purpose immediately readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should clarify what the agent can expect back from the check, such as pass/fail results or test output. It also lacks usage guidance and alternative routing, though the single optional parameter keeps the overall context relatively simple.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description says nothing about the cwd parameter, but the schema provides 100% coverage for it with 'Project root'. Since the schema already handles the parameter meaning, the description adds no additional semantic value, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb-resource pair: discovering test files, detecting the test framework, and running a quick test execution check. This clearly distinguishes check_tests from sibling tools like check_linting, check_security, and check_architecture, which target different aspects of the codebase.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when you need to inspect or execute tests. However, there is no explicit guidance about when to use this tool versus alternatives such as auditing the codebase or checking CI/CD, and it offers no exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_agentB

Detect which AI coding agent is in use (Claude Code, Cursor, Windsurf, Devin, Codex, Gemini, Antigravity) by scanning for agent-specific config files (.cursorrules, .windsurfrules, CLAUDE.md, AGENTS.md). Returns detected agents and all supported agents with best practices.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the scanning mechanism and return payload, which is useful, but it omits behavioral details such as read-only safety, handling of missing config files, and whether the cwd parameter changes what is scanned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose before giving the agent list and return summary. Every phrase earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description needs to carry more context. It covers the return value and supported agents reasonably well, but leaves gaps around cwd semantics, empty-result behavior, and the tool's side-effect profile, making it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, cwd, has zero schema description coverage, and the description never mentions it. While the schema indicates it is an optional string, the agent cannot infer its intended meaning or default behavior without additional guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Detect') and resource ('which AI coding agent is in use'), enumerates supported agents, and states the detection mechanism. It is clear but does not explicitly distinguish itself from sibling tools like get_agent_guidance, so it misses the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context for use is implied: use this when you need to identify which AI coding agent is configured in a project. There is no explicit guidance on when not to use it or which sibling alternatives might be preferable, leaving the agent to infer the appropriate selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_github_workflowB

Generate a production-ready GitHub Actions CI/CD workflow YAML. Supports custom name, stack, and deploy target (docker, k8s, vercel, aws).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoWorkflow name (default: ci)
stackNoTech stack name
deployTargetNoDeployment target: docker (default), k8s, vercel, aws

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It states capabilities (supports custom name, stack, deploy target) but does not disclose the output format (e.g., returns YAML string), side effects (if any), or prerequisites (e.g., GitHub repo existence). The term 'production-ready' hints at quality but not at behavior. This leaves the agent guessing about what happens when the tool is invoked.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loads the primary action, and includes the supported options without any fluff. Every word contributes to the purpose. It is appropriately concise and structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description should clarify what the tool returns (likely YAML content) and any usage constraints. It mentions deploy targets but does not describe the output or any side effects. The tool is simple, so a complete description would mention that it returns the workflow YAML. This omission prevents it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters, each with a description, so the baseline is 3. The description lists 'custom name, stack, and deploy target' but this merely restates the parameters without adding semantic depth. It does not explain parameter interactions, defaults beyond what schema states, or formatting requirements. The description adds marginal value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Generate a production-ready GitHub Actions CI/CD workflow YAML.' The verb 'Generate' and resource 'GitHub Actions CI/CD workflow YAML' are specific, and the mention of 'Supports custom name, stack, and deploy target' distinguishes it from sibling tools like check_cicd (which checks rather than generates) and generate_starter_repo (which generates a broader repo).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives. It does not mention any conditions, exclusions, or contrast with sibling tools such as check_cicd or generate_starter_repo. Usage is only implied by the name and purpose, which is insufficient for an agent to make routing decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_production_codeB

Generate production-ready code templates with industry-standard patterns. Supports: api, auth, database, testing, error_handling, logging, security, ci_cd, docker, branch_strategy. Specify stack (nestjs, express, fastify) and feature type.

ParametersJSON Schema
NameRequiredDescriptionDefault
stackNoFramework: nestjs, express, fastify, default
featureNoFeature type: api, auth, database, testing, error_handling, logging, security, ci_cd, docker
languageNoLanguage: typescript (default), javascript

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool 'generates' templates, implying a non-destructive action, but does not describe any side effects, such as whether files are written to the workspace, whether it requires an existing project, or any permissions needed. Without annotations, this is a significant gap for a code generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences with no fluff. It front-loads the purpose and then lists supported feature types and stack options. The only minor issue is that the list of supported values is also largely present in the schema, but the description's compact enumeration is still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 parameters, no output schema), the description covers the input semantics well but lacks details on the output or usage context. It does not explain what the generated templates look like, how they are returned (e.g., files vs. inline), or how this relates to the overall workflow. With no output schema, the description should clarify the return format, which it doesn't.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all three parameters (stack, feature, language), each with clear descriptions and defaults. The tool description adds value by listing the supported values for feature types and stacks, which supplements the schema. With high schema coverage, the baseline is 3, and the description adds marginal context (e.g., 'default' for stack), so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'generate' and the resource 'production-ready code templates' with specific features and stacks. However, it does not differentiate from siblings like 'generate_starter_repo' or other generation tools, so it is clear but lacks explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by listing the supported stacks and feature types, giving context for when to use it (when generating production code). However, it does not explicitly state when not to use it or mention alternatives such as 'generate_starter_repo' for scaffold-level generation, so guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_security_checklistB

Generate a security checklist based on current project findings + OWASP Top 10 standards. Detects stack and tailors recommendations.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo
stackNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions stack detection and tailoring, but does not state whether the tool performs read-only analysis, modifies any files, or requires specific permissions, nor does it describe the output format or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with the primary purpose front-loaded. No filler or redundancy, though it could include a bit more parameter detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is relatively simple with two parameters and no output schema. The description gives the gist but misses essential details: how to use cwd, whether stack is auto-detected or must be provided, and what the checklist contains or how it is returned. An agent would likely need to probe the tool to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining both parameters. It partially explains 'stack' (used for tailoring) but gives no meaning for 'cwd' and does not clarify if either parameter is optional or required. This leaves the agent guessing about parameter roles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action (generate) and resource (security checklist), and adds context (based on OWASP Top 10 and current project findings). It does not explicitly contrast with siblings like check_security or audit_codebase, but the purpose is distinct enough to avoid confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for generating a checklist based on project context, but it does not specify when to prefer it over check_security or audit_codebase, nor does it mention any prerequisites or exclusions. Usage is inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_starter_repoB

Generate a complete starter project structure with production-ready defaults: directory layout, essential packages, eslint/prettier/tsconfig/jest configs. Supports nestjs, express, fastify frameworks.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNoLanguage: typescript (default), javascript
frameworkNoFramework: nestjs (default), express, fastify

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It says 'Generate' and lists outputs, but does not state whether files are written to disk, whether existing files may be overwritten, what preconditions are required, or what side effects occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight, front-loaded sentences convey the core purpose and concrete outputs without filler. The supported framework list is appended cleanly at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, with two optional well-documented parameters and a clear summary of what it generates. However, without an output schema or annotations, the agent is left guessing about the delivery/side-effect model and when this tool is preferred over its code-generation sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both language and framework already document their allowed values and defaults. The description repeats framework support but adds no meaningful semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Generate') and resource ('complete starter project structure'), and names concrete artifacts (directory layout, packages, eslint/prettier/tsconfig/jest configs). It is clear enough, but it does not explicitly distinguish itself from siblings like generate_production_code.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'starter' implies use for initial project scaffolding, and the framework list gives selection options. However, there is no explicit statement of when to use this tool versus alternatives such as generate_production_code, or any when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_agent_guidanceC

Get agent-specific best practices and configuration guidance. Supported agents: claude-code, cursor, windsurf, devin, codex, gemini, antigravity.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoAgent name: claude-code, cursor, windsurf, devin, codex, gemini, antigravity

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, what the response format is, how errors are handled (e.g., unsupported agent), or any side effects. The listing of supported agents adds minimal behavioral context, but essential traits remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences) and front-loaded with the purpose. The supported agents list is useful but duplicates the schema's parameter description. While efficient, it could have been slightly shorter by omitting the redundant list, but overall it is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional parameter, no output schema), the description covers the core purpose and supported inputs. However, it lacks information about the response structure, behavior for unsupported agents, and the implication of the optional parameter. These gaps could lead to incorrect usage or assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% – the schema already describes the 'agent' parameter with the same list of supported agents. The description repeats this information without adding new meaning (e.g., default behavior when the parameter is omitted, case sensitivity, or format). Since the schema fully documents the parameter, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'agent-specific best practices and configuration guidance', making the purpose specific. It also enumerates supported agents, which adds precision. However, it does not explicitly differentiate from sibling tools like 'get_industry_patterns' or 'generate_production_code', though the resource is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites, scenarios, or exclusions. The only contextual hint is the list of supported agents, which implies a selection criterion but does not explain why one would choose this over related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_industry_patternsA

Get full industry-standard patterns and checklists for any architectural concern: api, auth, database, testing, error_handling, logging, security, ci_cd, docker, branch_strategy.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoPattern category (omit for all patterns)

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. The 'Get' verb and 'patterns and checklists' wording signal a read-only informational operation, but the description does not explicitly state the absence of side effects, external dependencies, or any output format details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence with no filler: the resource is stated first, then the category list is appended in a clear colon structure. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description is mostly complete: it states what is returned and which categories are supported. The main gap is the lack of usage guidance and alternative routing, but the tool's complexity is low enough that this is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the optional category parameter, but the description adds real value by listing the valid category values since the schema defines no enum. This lets the agent know what strings are meaningful without opening any external reference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get') and a concrete resource ('full industry-standard patterns and checklists'), then enumerates relevant categories. It is clear in isolation, but it does not explicitly contrast itself with sibling tools like check_architecture or generate_security_checklist, leaving some differentiation to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: the agent should call this when the user needs standard reference patterns or checklists for an architectural concern. There is no explicit when-to-use vs. when-not-to-use guidance, and no mention of alternatives such as checking the current codebase.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

production_readinessA

Compute a production-readiness scorecard (A–F grade) across 10 dimensions: package.json, lock file, README, .gitignore, ESLint, TypeScript strict mode, tests, CI/CD, .env safety, build script. Includes remediation guidance per item.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the full burden. 'Compute' suggests a read-only analysis and the mention of 'remediation guidance' indicates the output includes advice, but the description does not explicitly state whether the tool modifies any files, requires existing configuration files, or has any side effects. It gives a basic behavioral profile but omits details about auth, rate limits, or destructive actions, which are not relevant here but still lacks explicit 'read-only' or 'no modifications' language.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence with no filler. It leads with the core purpose, enumerates the 10 evaluated dimensions for clarity, and ends with the remediation guidance feature. Every phrase earns its place, and the structure front-loads the action and resource before detailing scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description adequately conveys the output (scorecard with grade and remediation guidance) and the scope (10 dimensions). However, with no output schema or annotations, it does not explain the return format (text vs. structured JSON), the meaning of `cwd` and its default behavior, or any prerequisites like the project being a Node.js repository. These are meaningful gaps for a tool that presumably inspects a filesystem, but the description still covers the main outputs and scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single optional string parameter `cwd`. The description does not mention `cwd` at all, so it fails to compensate for the lack of schema documentation. While `cwd` is a common parameter name meaning 'current working directory', its specific role in this tool (e.g., the directory to analyze) is left to inference. The description adds no semantic value beyond what the parameter name already suggests.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool computes a production-readiness scorecard with an A–F grade and enumerates 10 specific dimensions (package.json, lock file, README, etc.). It distinguishes itself from sibling check tools like check_tests or check_cicd by covering a holistic assessment rather than a single check. The verb 'Compute' and resource 'scorecard' are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (when you want an overall production-readiness score) but does not explicitly state when to prefer this over sibling tools like audit_codebase or check_security, nor does it give exclusions such as 'for individual checks use check_tests'. The broad scope is clear, but there is no direct guidance on tool selection or conditions for invoking this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 15 tool updatesv2.0.2
    • First observedaudit_codebase
    • First observedcheck_architecture
    • First observedcheck_branch
    • First observedcheck_cicd
    • First observedcheck_linting
    • First observedcheck_security
    • First observedcheck_tests
    • First observeddetect_agent
    • First observedgenerate_github_workflow
    • First observedgenerate_production_code
    • First observedgenerate_security_checklist
    • First observedgenerate_starter_repo
    • First observedget_agent_guidance
    • First observedget_industry_patterns
    • First observedproduction_readiness

TDQS

B3.3/5.0

Scored across 15 tools

Disambiguation3/5

The check_* family (tests, linting, security, architecture, CI/CD, branch) is well-separated by suffix, but audit_codebase and production_readiness heavily overlap as both assess overall codebase health. generate_starter_repo and generate_production_code also blur boundaries (full scaffold vs. feature templates), though descriptions partially clarify the distinction.

Naming Consistency4/5

Most tools follow a clear verb_noun snake_case convention (check_tests, generate_starter_repo, audit_codebase, detect_agent). The single outlier is production_readiness, which is a noun phrase rather than a verb action, and the mix of check_/generate_/get_ verbs is still predictable. Overall the pattern is consistent with one minor deviation.

Tool Count4/5

At 15 tools, the server sits at the upper boundary of a well-scoped set and is reasonable for a code-quality/readiness domain. A few tools feel tangential (detect_agent, get_agent_guidance) and add bulk without directly serving code auditing, but the count is not excessive.

Completeness4/5

The surface covers detection (checks), broad assessment (audit, scorecard), generation (starter repo, code templates, workflows, checklists), and reference guidance (industry patterns, agent practices). Notable gaps include a remediation/fix tool and test coverage analysis, but core workflows such as audit, generate, and recommend are well represented.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    A local-first MCP server that scores your codebase's Build Readiness by reading code and running tests on your machine, outputting a diligence-grade score and risk register without uploading your source.
    5
    38 npm
    Apache 2.0