Skip to main content
Glama

code-audit-mcp

CI License: MIT Python 3.11+

Audit your whole stack with a single question to your agent.

An MCP server (stdio) that gives Claude Code (and any MCP-compatible agent) a complete audit of your project: dead code, dependencies, security, CVEs, complexity and lint. It covers Laravel/PHP, Python (FastAPI, Flask), TypeScript/React/Next.js and Flutter/Dart, including monorepos (backend/ + frontend/, Laravel with React, a Flutter app next to its backend).

"Audit this repo" -> within a minute you get a report prioritized by severity, with false positives already flagged and the next step for every finding.

Why it is different

  • Zero installation in the project. vendor/, node_modules, the venv and .dart_tool stay untouched. Tools run through uvx/npx or from ~/.cache/code-audit-mcp, and the first run downloads them on its own.

  • Detects stacks by itself. It walks the repo, finds every project (even when several coexist) and audits each one with its native tools.

  • Built to be used by an LLM. Findings are normalized (file, line, rule, severity, message and how to fix it), sorted by severity and free of the noise of ten different tools.

  • Flags false positives and explains why. Routes registered by decorator, Laravel and Livewire hooks, Pydantic fields or the HEAD strings that bandit mistakes for passwords do not count in the summary. This way the agent never proposes deleting code that is in fact used.

  • Checks git history, not just the current code. A secret committed a year ago and deleted later is still compromised, and it finds it.

  • Secrets never appear in the output. They are used to group findings and to tell whether they are still in the code, but the value never leaves the server.

Related MCP server: ARGUS

Two tools

Tool

Purpose

audit(path=".")

Overview: runs the 6 checks in parallel and returns a summary per check and per stack, totals by severity and the top findings

check(name, path=".", severity, file, offset, limit, false_positives)

Detail for one aspect: the findings of a check sorted by severity, filterable and paginated

Checks: dead_code · dependencies · security · vulnerabilities · complexity · lint

Context savings. Responses are a single compact JSON (no structuredContent copy and no indentation). Messages that repeat identically across several findings of a rule are emitted only once in rules. check returns 50 findings per engine (limit) and leaves out likely false positives (false_positives=true to see them). Narrow results with severity (minimum) and file (glob or prefix, e.g. app/Http/**) and paginate with offset; the response tells you in next how to continue.

What it covers in each stack

Check

Laravel / PHP

Python (FastAPI, Flask)

TypeScript / React / Next.js

Flutter / Dart

dead_code

PHPStan + shipmonk (understands Eloquent, Blade and Livewire)

vulture

knip

—

dependencies

composer-unused + composer-require-checker

deptry

knip

—

vulnerabilities

composer audit

pip-audit (venv, uv.lock or requirements)

npm / pnpm / bun audit; yarn.lock in OSV.dev

pubspec.lock in OSV.dev

lint

the project's PHPStan (with its Larastan)

ruff

the project's tsc + eslint

dart analyze with its analysis_options.yaml

complexity

lizard

lizard

lizard

—

security

semgrep

bandit + semgrep

semgrep

semgrep

In every stack, security adds gitleaks over the whole git history and sensitive files that are tracked, not ignored, or deleted but still in history.

Flutter/Dart. dart analyze needs .dart_tool/ (flutter pub get); without it, it is skipped, because generating it would write to the project. The dart from the Flutter SDK is used if available. The Pub ecosystem has few advisories published in OSV: 0 CVEs does not guarantee there are none.

Security: what generic scanners miss

On top of the public semgrep packs, it ships its own rules (code_audit_mcp/rules/stack.yml) for the problems of these stacks:

  • Postgres/SQL: SQL with interpolation or concatenation in DB::raw/select, whereRaw/orderByRaw…, cursor.execute(f"…"), text(f"…"), pool.query(\…${x}`), $queryRawUnsafe`; connection URLs with a password in the code.

  • Redis: KEYS (blocks the server; use SCAN) and FLUSHALL/FLUSHDB in application code.

  • Laravel: env() outside config/ (breaks with config:cache), $guarded = [], {!! !!} in Blade, eval.

  • FastAPI/Flask: debug=True, CORS * with credentials, SECRET_KEY in the code.

  • React: dangerouslySetInnerHTML, tokens in localStorage.

  • Flutter/Android/iOS: badCertificateCallback accepting any certificate, tokens and passwords in SharedPreferences, usesCleartextTraffic/debuggable in the release AndroidManifest.xml and NSAllowsArbitraryLoads in Info.plist.

Secrets and sensitive files:

  • Files: .env*, private keys and certificates (.pem, .key, .p8, id_rsa, .p12…), credentials (.pypirc, .netrc, .pgpass, .git-credentials, .htpasswd, Android key.properties; .npmrc/.yarnrc.yml and Composer's auth.json only if they carry tokens), data dumps (.sqlite, .db, *dump*.sql, .sql.gz, .bak) and logs. They are reported when tracked in git, one git add . away from being committed (not ignored) and deleted but still present in history, with commit, author and date so you know what to rotate.

  • Secrets in history (gitleaks): tokens and keys that were ever committed on branches, tags or remotes, even if they are no longer in the code. One finding per secret, with the oldest commit and whether it is still in the code.

False positives

Dead-code detectors cannot see implicit usages. Every likely false positive of this kind comes out with likely_false_positive and the reason, and does not count in the severity summary:

  • Python: functions registered by decorator (@app.get, @router.post, @bp.route, @command…), Pydantic/SQLAlchemy/dataclass/Enum model fields, Alembic migrations, framework hooks. In dependencies, packages that are used without being imported (asyncpg/psycopg through the connection URL, python-multipart, uvicorn, alembic…).

  • Laravel: Laravel/Livewire/Filament hooks, Eloquent accessors and scopes, public members of Livewire/Filament components and names referenced from Blade views.

  • knip: laravel-vite-plugin entries (passed to knip as entry) and files of plugins that could not be loaded.

  • bandit: import subprocess (B404), commands with literal arguments (["git", "status"], B603/B607) and "passwords" that are not (HEAD, refresh_token:, uppercase constants; B105-B107).

Tests count as usage, but their findings are not reported.

Optional configuration: .code-audit.toml at the project root

exclude = ["legacy/**"]                         # paths to skip
ignore = ["B404", "DEP002:ruff", "unused-method:*Resource::form"]  # rule or rule:name (globs)

[dead_code]
ignore_names = ["on_*"]
ignore_decorators = ["@command"]   # Python
min_confidence = 60                # Python (vulture)

[complexity]
threshold = 11                     # minimum cyclomatic complexity (C = 11-20, D = 21-30…)

[security]
registry = true                    # public semgrep packs (needs network)

Requirements

  • uv (always).

  • PHP 8.2+ and Composer for PHP projects; Node.js for JS/TS projects; bun if the project uses bun.lock; Flutter or Dart SDK for linting Dart projects.

  • gitleaks downloads itself (official GitHub binary, sha256 verified) to ~/.cache/code-audit-mcp/bin, unless it is already in the PATH. CVEs for yarn.lock and pubspec.lock are queried against the OSV.dev API (network).

  • For complete results, the project must have its dependencies installed (composer install, npm ci, venv). Without them, each affected engine says so in skipped or notes.

The first run downloads the tools (a few minutes). On a large Laravel project, dead code takes 1-2 minutes; the rest of the checks, seconds.

Installation

Globally in Claude Code:

claude mcp add -s user code-audit /path/to/code-audit-mcp/bin/code-audit-mcp-server

Per project (.mcp.json):

{
  "mcpServers": {
    "code-audit": {
      "command": "/path/to/code-audit-mcp/bin/code-audit-mcp-server",
      "args": []
    }
  }
}

Environment variables

Variable

Default

Purpose

CODE_AUDIT_TIMEOUT

900

maximum seconds per tool

CODE_AUDIT_JOBS

3

simultaneous heavy processes

CODE_AUDIT_CACHE

~/.cache/code-audit-mcp

cache for PHP tools, PHPStan and gitleaks

CODE_AUDIT_UVX, CODE_AUDIT_UV, CODE_AUDIT_NPX, CODE_AUDIT_NPM, CODE_AUDIT_PHP, CODE_AUDIT_COMPOSER, CODE_AUDIT_BUN, CODE_AUDIT_DART, CODE_AUDIT_FLUTTER, CODE_AUDIT_GITLEAKS

from PATH

alternative binaries

Development

uv sync
uv run pytest
uvx ruff check .

Layout: core.py (processes, cache, finding format), detect.py (stacks), config.py, engines/{python,php,node,dart,shared}.py, checks.py (which engine resolves each check), server.py. The own semgrep rules are tested against tests/fixtures/rules, where every line that must be detected carries BAD:<rule>.

Contributing

Issues and pull requests are welcome; see CONTRIBUTING.md. Security reports go through SECURITY.md. Release notes live in CHANGELOG.md and planned work in ROADMAP.md. Licensed under the MIT License.


code-audit-mcp: the quality and security report your agent can read, prioritize and fix.

Available Tools

2 tools
auditA
Read-only

Summary of all checks (dead code, dependencies, security, CVEs, complexity, lint) per stack, with totals by severity and the top 5 findings of each engine. Downloads tools on first run.

Args: path: project folder (defaults to the working directory).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds a genuinely useful trait beyond them: 'Downloads tools on first run', which flags network activity and a first-run side effect (consistent with openWorldHint=true). It also discloses the return shape. It stops short of noting runtime cost or caching behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the result shape, then the notable side effect, then the argument in a compact Args block. Two sentences and an arg line with little waste; only minor redundancy in restating the default already present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully characterizes the return value (per-stack summary, severity totals, top 5 findings per engine) and covers the single parameter and first-run download. Adequate for a one-parameter, read-only tool; missing only guidance on cost/runtime and sibling selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (the lone parameter has only a title and a default), so the description must compensate, and it does: it explains that 'path' is the project folder and that it defaults to the working directory. It does not clarify relative-vs-absolute path handling, but the semantic gap is largely closed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete operation and enumerates exactly what it aggregates (dead code, dependencies, security, CVEs, complexity, lint) and the shape of the result (totals by severity, top 5 findings per engine). It does not, however, distinguish itself from the sibling 'check', so an agent cannot tell from the text alone whether this is the aggregate run or the individual check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or routing statement relative to the sibling 'check'. The only usage hint is the path default, which is about invocation, not selection. The agent must infer that this is the umbrella/summary tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkA
Read-only

Findings of one check, sorted by severity and paginated per engine.

Args: name: check to list. path: project folder (defaults to the working directory). severity: minimum severity to include. file: only paths matching this glob or prefix (e.g. "app/Http/**"). offset: findings to skip per engine (pagination). limit: maximum findings per engine. false_positives: include those flagged as likely false positives.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNo
nameYes
pathNo.
limitNo
offsetNo
severityNoinfo
false_positivesNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, non-destructive, open-world safety. The description adds genuinely new behavior: results are sorted by severity and pagination (offset/limit) is applied per engine, not globally, plus the ability to include likely false positives. That per-engine pagination detail is the kind of non-obvious trait annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded summary sentence, then a compact per-argument list with no filler. Every line adds information; the only minor cost is the generic 'Args:' scaffolding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with no output schema, the description covers inputs well but says nothing about the shape of a returned finding or what 'per engine' means in practice. It is adequate to invoke the tool, but an agent cannot anticipate the response structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden, and it does: all seven parameters get meaningful semantics (path defaults to the working directory, file is a glob or prefix like 'app/Http/**', offset/limit are per-engine, false_positives includes flagged items). Only 'name' and 'severity' rely on their enum values rather than prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource precisely ('findings of one check') and adds two scoping traits (sorted by severity, paginated per engine), so an agent knows what comes back. However, the verb is implicit and the sibling tool 'audit' is never mentioned, so it cannot be distinguished from the alternative on the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no statement of when to prefer this over 'audit', and no prerequisites or exclusions. The argument list implies usage but the agent must infer the selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedaudit
    • First observedcheck

TDQS

A3.7/5.0

Scored across 2 tools

Disambiguation4/5

audit and check are distinct: audit gives a summary across all checks, while check drills into a specific check with filters. The descriptions make the boundary clear, though both deal with findings and could momentarily confuse a new agent.

Naming Consistency4/5

Both tools use single, lowercase verbs (audit, check), which is consistent, but they don't follow the common verb_noun pattern. No mixing of conventions, so readability is high.

Tool Count4/5

Two tools is slightly under the typical 3-15 range, but they are well-scoped: audit provides the overview and check handles per-check details via a parameterized interface. Each earns its place, though the surface feels minimal.

Completeness3/5

Core read functionality is covered: summary and detailed findings with pagination/filtering. However, there is no way to list available check names, manage false positives, or run individual checks directly, which are notable gaps for a code audit workflow.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    C
    maintenance
    Enables AI agents to perform comprehensive, zero-infrastructure codebase analysis through 24 MCP tools, covering security, quality, architecture, type safety, git history, and dead code detection with high precision and local privacy.
    45
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables LLMs to analyze any codebase from a GitHub URL, producing deep reports with code snippets, line numbers, and prioritized suggestions. It orchestrates seven MCP tools through LangGraph to perform code review, onboarding, bug investigation, and complexity analysis.
    -