code-audit-mcp
Audits Android platform files in Flutter projects for security issues like usesCleartextTraffic and debuggable in release AndroidManifest.xml.
Audits Bun dependencies for known vulnerabilities via bun audit.
Audits Composer dependencies for unused packages, missing requirements, and known vulnerabilities via composer audit.
Audits Dart projects with dart analyze and checks dependencies and vulnerabilities.
Runs the project's ESLint configuration to report lint findings in TypeScript/JavaScript projects.
Audits FastAPI projects for security issues such as debug=True, CORS misconfigurations, and hardcoded SECRET_KEY.
Handles Laravel Filament hooks and component members as false positives in dead-code detection.
Audits Flask projects for security issues such as debug=True, CORS misconfigurations, and hardcoded SECRET_KEY.
Audits Flutter projects for dead code, dependencies, vulnerabilities, complexity, and lint, including Android/iOS security checks.
Checks git history for secrets and sensitive files that were committed or deleted, using gitleaks.
Audits iOS platform files in Flutter projects for security issues like NSAllowsArbitraryLoads in Info.plist.
Detects dead code and unused dependencies in TypeScript/JavaScript projects via knip.
Audits Laravel/PHP projects for dead code, dependencies, security, CVEs, complexity, and lint, with false-positive handling for Eloquent, Blade, Livewire, and Filament.
Handles Laravel Livewire hooks and component members as false positives in dead-code detection.
Audits Next.js projects for dead code, dependencies, vulnerabilities, complexity, and lint.
Audits npm dependencies for known vulnerabilities via npm audit.
Audits PHP projects for dead code, dependencies, vulnerabilities, complexity, and lint.
Audits pnpm dependencies for known vulnerabilities via pnpm audit.
Handles Pydantic model fields as false positives in dead-code detection.
Audits Python projects for dead code, dependencies, vulnerabilities, complexity, and lint.
Audits React projects for security issues such as dangerouslySetInnerHTML and tokens in localStorage.
Audits Redis usage for dangerous commands like KEYS and FLUSHALL/FLUSHDB in application code.
Runs Ruff to report lint findings in Python projects.
Handles SQLAlchemy model fields as false positives in dead-code detection.
Audits TypeScript projects for dead code, dependencies, vulnerabilities, complexity, and lint.
Checks yarn.lock dependencies against OSV.dev for known vulnerabilities.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@code-audit-mcpAudit this repo and show me the critical security and dependency findings"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
code-audit-mcp
Audit your whole stack with a single question to your agent.
An MCP server (stdio) that gives Claude Code (and any MCP-compatible agent) a complete audit of your project: dead
code, dependencies, security, CVEs, complexity and lint. It covers Laravel/PHP, Python (FastAPI, Flask),
TypeScript/React/Next.js and Flutter/Dart, including monorepos (backend/ + frontend/, Laravel with React,
a Flutter app next to its backend).
"Audit this repo" -> within a minute you get a report prioritized by severity, with false positives already flagged and the next step for every finding.
Why it is different
Zero installation in the project.
vendor/,node_modules, the venv and.dart_toolstay untouched. Tools run throughuvx/npxor from~/.cache/code-audit-mcp, and the first run downloads them on its own.Detects stacks by itself. It walks the repo, finds every project (even when several coexist) and audits each one with its native tools.
Built to be used by an LLM. Findings are normalized (file, line, rule, severity, message and how to fix it), sorted by severity and free of the noise of ten different tools.
Flags false positives and explains why. Routes registered by decorator, Laravel and Livewire hooks, Pydantic fields or the
HEADstrings that bandit mistakes for passwords do not count in the summary. This way the agent never proposes deleting code that is in fact used.Checks git history, not just the current code. A secret committed a year ago and deleted later is still compromised, and it finds it.
Secrets never appear in the output. They are used to group findings and to tell whether they are still in the code, but the value never leaves the server.
Related MCP server: ARGUS
Two tools
Tool | Purpose |
| Overview: runs the 6 checks in parallel and returns a summary per check and per stack, totals by severity and the top findings |
| Detail for one aspect: the findings of a check sorted by severity, filterable and paginated |
Checks: dead_code · dependencies · security · vulnerabilities · complexity · lint
Context savings. Responses are a single compact JSON (no structuredContent copy and no indentation). Messages
that repeat identically across several findings of a rule are emitted only once in rules. check returns 50
findings per engine (limit) and leaves out likely false positives (false_positives=true to see them). Narrow
results with severity (minimum) and file (glob or prefix, e.g. app/Http/**) and paginate with offset; the
response tells you in next how to continue.
What it covers in each stack
Check | Laravel / PHP | Python (FastAPI, Flask) | TypeScript / React / Next.js | Flutter / Dart |
| PHPStan + shipmonk (understands Eloquent, Blade and Livewire) | vulture | knip | — |
| composer-unused + composer-require-checker | deptry | knip | — |
|
| pip-audit (venv, | npm / pnpm / bun audit; |
|
| the project's PHPStan (with its Larastan) | ruff | the project's |
|
| lizard | lizard | lizard | — |
| semgrep | bandit + semgrep | semgrep | semgrep |
In every stack, security adds gitleaks over the whole git history and sensitive files that are tracked, not
ignored, or deleted but still in history.
Flutter/Dart. dart analyze needs .dart_tool/ (flutter pub get); without it, it is skipped, because
generating it would write to the project. The dart from the Flutter SDK is used if available. The Pub ecosystem
has few advisories published in OSV: 0 CVEs does not guarantee there are none.
Security: what generic scanners miss
On top of the public semgrep packs, it ships its own rules (code_audit_mcp/rules/stack.yml) for the problems of
these stacks:
Postgres/SQL: SQL with interpolation or concatenation in
DB::raw/select,whereRaw/orderByRaw…,cursor.execute(f"…"),text(f"…"),pool.query(\…${x}`),$queryRawUnsafe`; connection URLs with a password in the code.Redis:
KEYS(blocks the server; useSCAN) andFLUSHALL/FLUSHDBin application code.Laravel:
env()outsideconfig/(breaks withconfig:cache),$guarded = [],{!! !!}in Blade,eval.FastAPI/Flask:
debug=True, CORS*with credentials,SECRET_KEYin the code.React:
dangerouslySetInnerHTML, tokens inlocalStorage.Flutter/Android/iOS:
badCertificateCallbackaccepting any certificate, tokens and passwords inSharedPreferences,usesCleartextTraffic/debuggablein the releaseAndroidManifest.xmlandNSAllowsArbitraryLoadsinInfo.plist.
Secrets and sensitive files:
Files:
.env*, private keys and certificates (.pem,.key,.p8,id_rsa,.p12…), credentials (.pypirc,.netrc,.pgpass,.git-credentials,.htpasswd, Androidkey.properties;.npmrc/.yarnrc.ymland Composer'sauth.jsononly if they carry tokens), data dumps (.sqlite,.db,*dump*.sql,.sql.gz,.bak) and logs. They are reported when tracked in git, onegit add .away from being committed (not ignored) and deleted but still present in history, with commit, author and date so you know what to rotate.Secrets in history (gitleaks): tokens and keys that were ever committed on branches, tags or remotes, even if they are no longer in the code. One finding per secret, with the oldest commit and whether it is still in the code.
False positives
Dead-code detectors cannot see implicit usages. Every likely false positive of this kind comes out with
likely_false_positive and the reason, and does not count in the severity summary:
Python: functions registered by decorator (
@app.get,@router.post,@bp.route,@command…), Pydantic/SQLAlchemy/dataclass/Enum model fields, Alembic migrations, framework hooks. In dependencies, packages that are used without being imported (asyncpg/psycopgthrough the connection URL,python-multipart,uvicorn,alembic…).Laravel: Laravel/Livewire/Filament hooks, Eloquent accessors and scopes, public members of Livewire/Filament components and names referenced from Blade views.
knip:
laravel-vite-pluginentries (passed to knip as entry) and files of plugins that could not be loaded.bandit:
import subprocess(B404), commands with literal arguments (["git", "status"], B603/B607) and "passwords" that are not (HEAD,refresh_token:, uppercase constants; B105-B107).
Tests count as usage, but their findings are not reported.
Optional configuration: .code-audit.toml at the project root
exclude = ["legacy/**"] # paths to skip
ignore = ["B404", "DEP002:ruff", "unused-method:*Resource::form"] # rule or rule:name (globs)
[dead_code]
ignore_names = ["on_*"]
ignore_decorators = ["@command"] # Python
min_confidence = 60 # Python (vulture)
[complexity]
threshold = 11 # minimum cyclomatic complexity (C = 11-20, D = 21-30…)
[security]
registry = true # public semgrep packs (needs network)Requirements
uv (always).
PHP 8.2+ and Composer for PHP projects; Node.js for JS/TS projects; bun if the project uses
bun.lock; Flutter or Dart SDK for linting Dart projects.gitleaks downloads itself (official GitHub binary, sha256 verified) to
~/.cache/code-audit-mcp/bin, unless it is already in the PATH. CVEs foryarn.lockandpubspec.lockare queried against the OSV.dev API (network).For complete results, the project must have its dependencies installed (
composer install,npm ci, venv). Without them, each affected engine says so inskippedornotes.
The first run downloads the tools (a few minutes). On a large Laravel project, dead code takes 1-2 minutes; the rest of the checks, seconds.
Installation
Globally in Claude Code:
claude mcp add -s user code-audit /path/to/code-audit-mcp/bin/code-audit-mcp-serverPer project (.mcp.json):
{
"mcpServers": {
"code-audit": {
"command": "/path/to/code-audit-mcp/bin/code-audit-mcp-server",
"args": []
}
}
}Environment variables
Variable | Default | Purpose |
|
| maximum seconds per tool |
|
| simultaneous heavy processes |
|
| cache for PHP tools, PHPStan and gitleaks |
| from PATH | alternative binaries |
Development
uv sync
uv run pytest
uvx ruff check .Layout: core.py (processes, cache, finding format), detect.py (stacks), config.py,
engines/{python,php,node,dart,shared}.py, checks.py (which engine resolves each check), server.py.
The own semgrep rules are tested against tests/fixtures/rules, where every line that must be detected carries
BAD:<rule>.
Contributing
Issues and pull requests are welcome; see CONTRIBUTING.md. Security reports go through SECURITY.md. Release notes live in CHANGELOG.md and planned work in ROADMAP.md. Licensed under the MIT License.
code-audit-mcp: the quality and security report your agent can read, prioritize and fix.
Available Tools
2 toolsauditARead-only
Summary of all checks (dead code, dependencies, security, CVEs, complexity, lint) per stack, with totals by severity and the top 5 findings of each engine. Downloads tools on first run.
Args: path: project folder (defaults to the working directory).
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, but the description adds a genuinely useful trait beyond them: 'Downloads tools on first run', which flags network activity and a first-run side effect (consistent with openWorldHint=true). It also discloses the return shape. It stops short of noting runtime cost or caching behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the result shape, then the notable side effect, then the argument in a compact Args block. Two sentences and an arg line with little waste; only minor redundancy in restating the default already present in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully characterizes the return value (per-stack summary, severity totals, top 5 findings per engine) and covers the single parameter and first-run download. Adequate for a one-parameter, read-only tool; missing only guidance on cost/runtime and sibling selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (the lone parameter has only a title and a default), so the description must compensate, and it does: it explains that 'path' is the project folder and that it defaults to the working directory. It does not clarify relative-vs-absolute path handling, but the semantic gap is largely closed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete operation and enumerates exactly what it aggregates (dead code, dependencies, security, CVEs, complexity, lint) and the shape of the result (totals by severity, top 5 findings per engine). It does not, however, distinguish itself from the sibling 'check', so an agent cannot tell from the text alone whether this is the aggregate run or the individual check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or routing statement relative to the sibling 'check'. The only usage hint is the path default, which is about invocation, not selection. The agent must infer that this is the umbrella/summary tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkARead-only
Findings of one check, sorted by severity and paginated per engine.
Args: name: check to list. path: project folder (defaults to the working directory). severity: minimum severity to include. file: only paths matching this glob or prefix (e.g. "app/Http/**"). offset: findings to skip per engine (pagination). limit: maximum findings per engine. false_positives: include those flagged as likely false positives.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | ||
| name | Yes | ||
| path | No | . | |
| limit | No | ||
| offset | No | ||
| severity | No | info | |
| false_positives | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, non-destructive, open-world safety. The description adds genuinely new behavior: results are sorted by severity and pagination (offset/limit) is applied per engine, not globally, plus the ability to include likely false positives. That per-engine pagination detail is the kind of non-obvious trait annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded summary sentence, then a compact per-argument list with no filler. Every line adds information; the only minor cost is the generic 'Args:' scaffolding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema, the description covers inputs well but says nothing about the shape of a returned finding or what 'per engine' means in practice. It is adequate to invoke the tool, but an agent cannot anticipate the response structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden, and it does: all seven parameters get meaningful semantics (path defaults to the working directory, file is a glob or prefix like 'app/Http/**', offset/limit are per-engine, false_positives includes flagged items). Only 'name' and 'severity' rely on their enum values rather than prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource precisely ('findings of one check') and adds two scoping traits (sorted by severity, paginated per engine), so an agent knows what comes back. However, the verb is implicit and the sibling tool 'audit' is never mentioned, so it cannot be distinguished from the alternative on the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of when to prefer this over 'audit', and no prerequisites or exclusions. The argument list implies usage but the agent must infer the selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
audit - First observed
check
TDQS
Scored across 2 tools
audit and check are distinct: audit gives a summary across all checks, while check drills into a specific check with filters. The descriptions make the boundary clear, though both deal with findings and could momentarily confuse a new agent.
Both tools use single, lowercase verbs (audit, check), which is consistent, but they don't follow the common verb_noun pattern. No mixing of conventions, so readability is high.
Two tools is slightly under the typical 3-15 range, but they are well-scoped: audit provides the overview and check handles per-check details via a parameterized interface. Each earns its place, though the surface feels minimal.
Core read functionality is covered: summary and detailed findings with pagination/filtering. However, there is no way to list available check names, manage false positives, or run individual checks directly, which are notable gaps for a code audit workflow.
Maintenance
Related MCP Connectors
Four IaC audits in one call: Compose, Dockerfile, GitHub Actions, Kubernetes. 131 checks.
Repo intel for AI coding agents: overview, PRs, contributors, hot files, CI, deps. Remote MCP.
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Scan any public GitHub MCP-server repo for security issues. 37 MCP-specific L1 rules, 8 languages.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables agents to audit and safeguard repositories by detecting dependency pinning issues, license compliance problems, hardcoded secrets, and dead code through MCP tools.4MIT
- FlicenseBqualityCmaintenanceEnables AI agents to perform comprehensive, zero-infrastructure codebase analysis through 24 MCP tools, covering security, quality, architecture, type safety, git history, and dead code detection with high precision and local privacy.45-
- FlicenseNot gradedqualityCmaintenanceEnables LLMs to analyze any codebase from a GitHub URL, producing deep reports with code snippets, line numbers, and prioritized suggestions. It orchestrates seven MCP tools through LangGraph to perform code review, onboarding, bug investigation, and complexity analysis.-
- AlicenseNot gradedqualityBmaintenanceEnables MCP-compatible clients to audit public GitHub repositories and receive structured launch-readiness reports with evidence-backed findings, reproducible scoring, and ready-to-paste launch copy.569 npm3MIT