mcp-skylos
This MCP server exposes Skylos static-analysis capabilities as tools for AI agents, enabling code scanning, verification, remediation, and triage.
analyze: Run Skylos' default dead-code scan on a path with confidence and exclusion options.
security_scan: Detect security flaws (SQLi, XSS, SSRF, etc.) in a codebase.
quality_check: Assess quality regressions like complexity, duplication, and deep nesting.
secrets_scan: Find hard-coded secrets, API keys, and high-entropy strings.
remediate: Apply AI-assisted fixes for findings (supports model selection, dry-run, severity filter, test command, and max fixes).
verify_dead_code: Use an LLM to verify static dead-code findings (with max verify/challenge counts).
provenance_scan: Scan provenance with an optional diff base for change-aware analysis.
generate_fix: Generate (and optionally apply) code fixes for dead code (delete mode) with a minimum safety threshold.
learn_triage: Record human triage decisions (accept/dismiss) for findings to improve future suggestions.
get_triage_suggestions: Retrieve triage suggestions for a path based on learned patterns.
validate_code_change: Validate a unified diff for security regressions, dangerous patterns, secrets, and AI-defense issues before merge; returns pass/fail.
get_security_context: Obtain security-relevant context for a path (likely to aid agent decision-making).
Scans Docker Compose and other deployment configs for privileged access, host networking, missing sandboxing, and other edge-device security misconfigurations.
Generates GitHub Actions PR gates with annotations and failure thresholds, and scans workflows for CI/CD misconfigurations.
Scans GitLab CI configurations for dangerous triggers, unpinned actions/includes, broad tokens, and other supply-chain risks.
Website | Docs | Repo Map | Quick Start | GitHub Action | VS Code Extension | Real-World Results | Benchmarks | Roadmap | Contributing
English | Deutsch | 简体中文 | Translations
What Is Skylos?
Skylos is an open-source static analysis CLI for Python, TypeScript, JavaScript, Java, Go, Kotlin, PHP, Rust, Dart, C#, C++, Shell, and deployment config. It runs locally by default and can also be used as a CI/CD PR gate.
Use Skylos when you want one command to check a repo or pull request for:
dead code and unused files
security flaws and dangerous data flows
secrets and dependency CVEs
CI/CD and edge-device deployment misconfigurations
quality regressions such as complexity, duplicate branches, and deep nesting
common AI-generated code mistakes, including missing guards, fake helpers, invented package APIs, and impossible dependency versions
LLM app risks such as unsafe tool use and missing output validation
Related MCP server: loctree-mcp
Choose the Right Command
Each command answers a different question. The source scan requires PATH.
Bracketed paths on verify, suite, defend, and clean default to the
current directory. suite and defend require a directory; verify and
clean also accept a file.
Question | Command | Input checked |
What problems are in this source tree? |
| Source files and project configuration; dead code by default, every main source analyzer with |
Should a model review static dead-code findings? |
| LLM by default; optional Jev-only or Jev-plus-LLM review |
Does this code contain AI-code mistakes? |
| AI-defect checks over the selected file or tree, plus a separate Git HEAD behavior comparison for supported Python working changes |
Will this exact local GPU build fit the machines we ship to? |
| A local built file or directory and |
What vulnerabilities are in this container image? |
| A remote registry image scanned by a separately installed Trivy; |
What does the combined repo suite report? |
| Static analysis, technical debt, AI defense, and provenance; local findings are report-only by default |
Does an agent implementation have deployment guardrails? |
| Recognized Python and TypeScript/JavaScript LLM integrations; gating requires a threshold flag or policy |
Which Python dead code can Skylos remove? |
| Python import/function cleanup candidates; |
Run skylos --help for this chooser, skylos <command> --help for one
command, and skylos commands for the command-family map.
The report commands have different gate and network behavior:
image scanrequires Trivy on trustedPATHand uses Trivy's remote image source, so registry network access and any required registry credentials must already be available. Without--fail-on, a completed scan exits0even when it reports vulnerabilities.suiteruns in the local process and does not upload unless--uploadis set, but its dependency scan can query OSV. Findings are report-only: the command exits0regardless of their count. Operational and output failures are nonzero;--uploadcan also fail for an upload error or Cloud quality gate.defendreports guardrail findings by default. It becomes a gate with--fail-on,--min-score, or gate settings in an explicit policy.cleanwithout--dry-runor--applyis interactive and can write after the final confirmation. Its current codemods support Python imports and functions. An apply pass still exits0if an individual edit prints a failure, so review the completion output.
Start In 60 Seconds
pip install skylos
skylos .The default scan focuses on dead code. Run every main source analyzer,
including security, secrets, quality, dependency, and AI-defect checks, with
-a:
skylos . -aRun only evidence-backed AI defect checks with:
skylos . --ai-defectsVerify a repository, file, or range before an agent hands it to review:
skylos verify . --file src/app.py --range 40:75 --project-contextFor a directory such as ., the AI-defect scan covers the selected tree. The
separate behavior result models supported Python working-tree changes against
Git HEAD; it is not the scope selector for the AI-defect scan. Dependency
hallucination checks are enabled for path targets and can query package
registries; use --no-dependency-hallucinations to disable those lookups.
High/critical security findings (category: "security") and hard-coded
secrets (category: "secret", value redacted) in the selected file/range also
fail verification; use --no-security for the AI-code-only verdict.
Interactive terminals get a human report. Redirected stdout and -o produce
the versioned JSON result.
skylos verify schema version 2 returns pass, fail, or incomplete.
incomplete means a requested proof could not be established, such as a
third-party TS/JS import, computed namespace member, unsupported language-local
API check, or parser surface that Skylos could not prove; it exits 2 unless
--no-fail is set. The coverage object lists detected languages, expected
checks, language support, missing checks, completed/skipped checks, checked
references, and deterministic skip reasons.
Deterministic local/workspace API verification currently covers Python, TypeScript/JavaScript, Go, and Java without executing target code. PHP, Rust, Dart, C#, Kotlin, and Shell retain their existing static-analysis coverage, but their local API proof is reported as unsupported and therefore incomplete. See AI Code Verification Coverage.
Create a local AI hallucination contract for repo-specific generated-code
truth. skylos verify auto-discovers .skylos/ai-contract.yml:
skylos contract init
skylos contract inspect
skylos verify .Test a running agent against deterministic response and tool-use scenarios:
skylos agent init
skylos agent test --allow-contract-endpointCreate a project config with thresholds, ignores, template hooks, and vibe dictionary extensions:
skylos initCreate a starter local rule pack:
skylos rules init
skylos rules validate .skylos/rules/local.yml
skylos rules list --json
skylos rules list cross --json
skylos rules list --packs --json
skylos cache statsGenerate a GitHub Actions PR gate:
skylos cicd init
git add .github/workflows/skylos.yml
git commit -m "Add Skylos CI gate"
git pushNeed more commands? Read the CLI Reference.
Check an Exact GPU Release Artifact
skylos preflight checks the built artifact itself against the repository's
declared GPU fleet. This is separate from the SKY-GPU* source scan, which
checks Dockerfiles, CUDA build settings, and TensorRT packaging intent before
the artifact exists.
Declare every machine that receives the same release:
# .skylos/gpu-targets.yml
version: 1
targets:
- name: inference-t4
vendor: nvidia
driver: "535.104.05"
compute_capability: "7.5"
platform: "linux/amd64"The canonical filename is .skylos/gpu-targets.yml; the same schema is also
accepted as .skylos/gpu-targets.yaml.
Then inspect a local build. Skylos uses a trusted system cuobjdump, takes a
private snapshot, and never loads or executes the artifact:
skylos preflight build/appTo make the command argument-free in CI, bind it to a project-relative artifact with a strict release receipt:
{"version": 1, "artifact": "build/app"}Save that file as .skylos/release.json, then run skylos preflight.
For requests that reach report generation, terminal output is concise and
redirected output is schema-versioned JSON. Argument or adapter errors can be
plain text and exit 2.
Status | Exit | Meaning |
|
| Every declared target is compatible within the stated static evidence scope, and the exact artifact identity is verified |
|
| Artifact evidence proves at least one declared target incompatible |
|
| Required evidence is missing, ambiguous, unsupported, or incomplete; it never silently becomes a pass |
cuobjdump can report several code-object groups, identified by the producer
label in its selected executable fatbin. Skylos requires a compatible route
for every reported group. A PTX-only route stays UNKNOWN because static
inspection cannot prove that the deployment driver will JIT it successfully.
The overall result is FAIL if any target or required check fails; otherwise
it is UNKNOWN if any result is unknown, and PASS only when all results pass.
Version 1 proves the local artifact identity, Linux ELF platform, selected
executable-fatbin architecture routes, a static packaged $ORIGIN CUDA
runtime route, and documented CUDA driver-family compatibility. It does not
prove runtime execution, workload correctness, memory demand, performance,
nonselected or relocatable fatbins, or per-kernel symbol parity. Windows PE
runtime import proof is not implemented. Digest-pinned OCI references are
accepted as identities but are never pulled or started, so the CLI always
reports UNKNOWN for them; trusted callers can supply digest-bound inspection
facts through skylos.preflight.run_preflight(...).
Common Workflows
Goal | Command | What You Get | More Detail |
First dead-code scan |
| Finds unused functions, classes, imports, files, and framework entrypoint mistakes | |
Deterministic cleanup preview |
| Shows planned Python import/function removals without writing; add | |
Security and quality audit |
| Adds dangerous flow, secrets, dependency, config, quality, and AI-defect checks | |
Combined repo report |
| Reports static findings, technical debt, AI defense, and provenance; SCA can query OSV, and findings alone exit | |
Optional Python linting |
| Runs Ruff with its native configuration, output, fixes, and exit codes through the Skylos CLI | |
PR gate |
| Generates a GitHub Actions workflow with annotations and failure thresholds | |
GitLab merge request report |
| Exports a native Code Quality report for GitLab CI artifacts | |
Offline dependency SBOM |
| Lists supported recorded dependencies as CycloneDX 1.6 JSON without network requests | |
SPDX SBOM + license policy |
| SPDX 2.3 JSON with declared licenses; | |
Container-image scan |
| Uses separately installed Trivy, registry network/auth, and an explicit severity gate for a pinned remote image | |
Built GPU artifact preflight |
| Verifies the exact local artifact identity, selected CUDA architectures, packaged runtime route, and declared fleet compatibility; returns | |
Container-image report import |
| Converts an existing Trivy image vulnerability report to Skylos JSON/SARIF; optional digest-bound severity check | |
Readable terminal report |
| Groups findings by file with severity badges, snippets, and copyable | |
Single-rule review |
| Enables the matching analyzer family and reports only that exact rule with its full message | |
Selectable terminal triage |
| Opens a keyboard-driven category list, finding list, and detail pane | |
IDE/test-script output |
| Prints untruncated | |
In-loop AI-code verification |
| Reports a narrow set of hallucinated helpers, unfinished code, stale references, disabled controls, and API/dependency hallucinations; JSON is used for redirected or | |
AI hallucination contracts |
| Auto-discovers | |
Changed-lines review |
| Keeps findings focused on active work instead of legacy debt | |
Incumbent scanner comparison |
| Runs Skylos beside the current scanner and produces a revision-aware scorecard: active overlap, raw unique findings by category, and eligible findings inside evidence-backed unused symbols—without replacing the current gate. | |
Kubernetes exposure proof |
| Checks an explicitly external, explicitly plain-HTTP Ingress chain inside one rendered multi-document bundle; route checks compare exact framework wiring with the workload's declared source file and required guards | |
GPU source/build intent gate |
| Checks declared CUDA image, architecture, driver, and TensorRT packaging intent before the artifact is built | |
Runtime-assisted dead-code check |
| Uses runtime traces to reduce dynamic-code false positives | |
Local rule pack |
| Scaffolds YAML rules for project-specific security and quality checks | |
Security agent quick scan |
| One-shot LLM security audit; compatibility alias for | |
Security agent deep scan |
| Three-stage security workflow with threat-model context, static threat traces, discovery/validation, and remediation handoff | |
AI-assisted review |
| Static analysis plus optional LLM review and fix suggestions | |
Agent harness replay |
| Validates and summarizes saved agent verification phases, tool calls, decisions, and budgets | |
Runtime agent behavior test |
| Checks final responses, tool selection, explicit refusals, and source IDs against a versioned contract | |
Verification-backed remediation |
| Scans and fixes supported findings, then re-scans them and records proof-test metadata when available | |
Agent-loop hooks |
| Verifies every agent edit, blocks secret reads and hallucinated package installs, and holds "done" while new issues are open | |
MCP agent verification |
| Lets Claude, Cursor, and other MCP clients verify an edited file/range with the same schema as | |
LLM integration inventory |
| Maps recognized LLM calls, agent tools, prompt sites, and input sources in Python and TypeScript/JavaScript | |
Pre-deployment agent verification |
| Verifies agent guardrails, scores OWASP LLM/Agentic coverage, and emits an attested evidence report | |
Agent verification CI gate |
| Blocks deploys with unguarded LLM integrations; SARIF for code scanning via | |
MCP agent pre-flight |
| Lets coding agents statically verify the agents they build — scores, failed checks, attestation digest. Requires | |
Technical debt triage |
| Ranks hotspots and debt trends |
What Skylos Catches
Category | Examples | Why It Matters |
Dead code | unused functions, classes, imports, package entrypoints, route handlers | reduces maintenance cost without breaking dynamic frameworks |
Security flaws | SQL injection, XSS, SSRF, path traversal, command injection, unsafe deserialization | catches exploitable flows before code reaches main |
Secrets | API keys, tokens, private credentials, high-entropy strings | prevents credentials from leaking through commits and PRs |
CI/CD workflows | GitHub Actions and GitLab CI dangerous triggers, unpinned actions/includes, broad tokens, OIDC misuse, cache poisoning, mutable images | reduces CI/CD supply-chain risk before release jobs run |
Edge deployment config | Docker Compose privileged device access, host networking, systemd root services, broad capabilities, missing sandboxing | catches repo-controlled settings that turn app bugs into device compromise |
Kubernetes deployment exposure | Explicitly external Ingress paths reaching a sensitive FastAPI/Flask route without its deployment-required guard, Flask | reports only when the resources are in one rendered bundle and every deployment edge resolves unambiguously |
GPU release compatibility | source/build contract mismatches plus built-artifact identity, selected CUDA architecture, packaged runtime, and fleet checks | catches declared intent errors early and verifies the resulting local artifact before release |
Quality regressions | complexity, deep nesting, duplicate branches, long functions, inconsistent returns | keeps AI-assisted refactors from adding brittle code |
AI code mistakes | phantom security calls, missing decorators, unfinished stubs, disabled controls, real packages called with invented APIs, impossible npm/Go versions | catches common hallucinated or incomplete code paths before they reach review |
LLM app risks | unsafe tool use, prompt injection exposure, missing output validation, missing rate limits | helps teams ship AI features with guardrails |
See the full Rules Reference.
Dependency licenses
skylos sbom . writes declared dependency licenses into CycloneDX, and
skylos sbom . --format spdx-json writes an SPDX 2.3 SBOM. Licenses come from
lockfiles and installed package metadata, offline. Unknown or ambiguous values
(such as BSD) stay NOASSERTION rather than being guessed. --license-lookup
opts in to a deps.dev query. To block licenses in -a scans, set a policy:
[tool.skylos]
license_deny = ["GPL-*", "AGPL-3.0-only"]
license_severity = "HIGH"Violations are reported as SKY-SCA-LIC001. See
License compliance.
Check Every Agent Edit (Claude Code, Codex, Cursor)
Install local hooks so the agent is checked while it works, not after:
skylos agent install-hooks # Claude Code (.claude/settings.json)
skylos agent install-hooks --codex # Codex (.codex/hooks.json)
skylos agent install-hooks --cursor # Cursor (.cursor/hooks.json)After each edit: verifies only the changed lines (security, secrets, AI-code mistakes) and tells the agent what to fix.
Before a file read: blocks files that contain hard-coded secrets.
Before a package install: blocks hallucinated or typosquatted packages.
At stop: blocks "done" while issues the agent added are still open.
Hooks fail open, log to .skylos/hook.log, and merge with your existing hooks.
--uninstall removes only the Skylos entries. See
Agent-loop hooks for the contract, latency, and
limits.
Verify AI Agents Before They Ship
Runtime guardrails are the WAF; Skylos is the SAST. skylos discover scans
Python and TypeScript/JavaScript for recognized LLM integrations (provider
SDKs, agent frameworks including the OpenAI Agents SDK, Claude Agent SDK, and
Google ADK, MCP servers and their tools, direct HTTP calls to LLM APIs or
OpenAI-compatible gateways, plus agent tools, prompt sites, and input
sources). skylos defend checks the guardrails around those detected
integrations deterministically, in the local process, with no model in the
loop. It emits evidence by default and becomes a CI gate only when a threshold
flag or policy supplies gate criteria.
skylos discover . # inventory LLM integrations and agent tools
skylos defend . # score guardrails (13 weighted checks)
skylos defend . --format md -o evidence.md # auditor evidence report + attestation
skylos defend . --format sarif -o defend.sarif # GitHub code scanning upload
skylos defend . --fail-on critical # CI gate: exit 1 on critical gaps
skylos defend . --owasp-framework agentic # report against OWASP Agentic ASI Top 10These commands do not inspect integrations in other languages. If discovery
finds no supported integration, defend returns an empty inventory; do not
treat the resulting score by itself as proof that an unsupported or
unrecognized agent implementation has guardrails.
Per integration it verifies: dangerous output sinks (eval/exec/subprocess), agent tool scope and typed schemas, prompt-injection exposure (delimiters, untrusted input paths, RAG context isolation), output validation, PII filtering, and model pinning — plus ops checks (logging, cost controls, rate limiting) scored separately so they never inflate the security score.
OWASP mapping: LLM Top 10 (2024/2025) and Agentic ASI Top 10 (2026).
Evidence report (
--format md): integration inventory, per-check results, OWASP coverage, regulatory framework evidence (EU AI Act, NIST AI RMF, ISO/IEC 42001 — "evidence toward" mappings, never compliance claims), and a remediation appendix.Attestation: JSON/md/SARIF reports carry a reproducible SHA-256 digest over file contents, policy, plugin set, integration inventory, scores, and full check evidence — re-run on the same tree with the same flags and Skylos version, and the digest must match.
CI-native:
skylos cicd init --defendgenerates the workflow step, theskylos-defendpre-commit hook gates locally, and$GITHUB_STEP_SUMMARYgets a score summary automatically in Actions.Policy as code:
skylos-defend.yamlpins gate thresholds and severity overrides (--policy).Agent-native: the
verify_agentMCP tool lets coding agents verify the agents they build — deterministic verification, not AI checking AI.
Static pre-deployment verification complements runtime controls (gateways, policy engines, human approval flows); it does not replace them. Full guide: docs/agent-verification.md.
Test Running Agent Behavior
Skylos separates generated-code truth, static agent guardrails, and observed runtime behavior:
Command | Verification question |
| Did the agent generate valid, non-hallucinated code? |
| Does the agent implementation contain the required guardrails? |
| Did the running agent behave according to its contract? |
Create .skylos/agent-test.yml, then test a live OpenAI-compatible endpoint:
skylos agent init
skylos agent test --allow-contract-endpointOr evaluate captured evidence without a network call:
skylos agent test --observations agent-observations.json
skylos agent test --observations agent-observations.json \
--format json --output agent-results.jsonVersion 1 deterministically checks exact response substrings, required,
allowed, and forbidden tool calls, tool arguments and sequence, maximum call
count, explicit refusals, and explicit source IDs. Missing typed evidence is
incomplete, never pass; exit codes are 0 pass, 1 violation, and 2
incomplete/invalid. Tool selection and final-answer source-ID checks are separate
one-turn scenarios: Skylos records local replayable evidence but never executes
tools returned by the target agent. Offline observations are marked as
unverified fixtures rather than runtime proof.
For an authenticated remote endpoint, keep the destination and secret choice in the trusted CLI invocation:
skylos agent test --endpoint https://agent.example.com/v1/chat/completions \
--allow-remote --auth-env MY_AGENT_API_KEYFull guide: docs/agent-behavior-testing.md.
How Skylos Fits
Skylos is not a replacement for every specialized scanner. It is a local-first repo and PR checker that puts several common review checks behind one CLI.
Framework-aware dead code detection: FastAPI, Django, Flask, pytest, SQLAlchemy, Next.js, React, package entrypoints, and common plugin patterns.
PR-focused output: diff scanning, CI thresholds, GitHub annotations, and baselines for existing findings.
Local-first operation: core analysis executes locally and does not upload source or call an LLM. Dependency/CVE checks can query OSV or package registries, and direct container scanning contacts a registry through Trivy.
AI-assisted change review: checks for removed validation, auth, logging, CSRF, rate limiting, timeouts, real-package API hallucinations, and other guardrails in generated or edited code.
Agent-loop verification:
skylos verifyand MCPverify_changeuse a versioned result schema for AI-code trust, security, and secret findings, so coding agents can self-correct before a human sees the change. The CLI renders a human report on a terminal and JSON when redirected or written with-o.Evidence-backed AI defects:
--ai-defectsand full scans put strict AI-code failure checks underai_defects, including phantom references, fake package APIs, nonexistent packages, impossible dependency versions, and weakened test assertions. The category/tag isai_defect; several rules intentionally keep historicalSKY-LorSKY-DIDs for suppression and baseline compatibility, while new AI-defect-only checks useSKY-A.Verification-backed remediation: security fixes are checked by re-running analysis, and supported findings can include targeted regression-test proof metadata.
Project-specific rules: add local YAML rules and extend prompt, credential, sensitive-file, and timeout dictionaries from config.
One command surface: dead code, security, secrets, dependency, quality, technical debt, agent review, and pre-deployment agent verification commands share the same CLI.
Agent Harness Artifacts
skylos agent verify . and skylos agent test record replayable artifacts
under .skylos/runs/<run-id> and print the run directory in table output. JSON
output includes the same harness summary under the harness key.
Use skylos agent replay .skylos/runs/<run-id> to validate and inspect a saved
run without making LLM calls. Add --format json when another agent or CI job
needs machine-readable status. A valid replay exits 0; an invalid or corrupt
artifact set exits 1 with issue codes. Replay output includes
schema_version so CI and agents can detect artifact-contract changes.
Replay checks internal consistency and corruption; artifacts are not signed and
are not proof against an actor that can rewrite the entire run directory.
Each run directory contains:
events.jsonl: chronological run, phase, and tool-call events.state.json: full observable state, including phases, tool calls, decisions, and budget usage.summary.json: compact status, counts, budget, and artifact paths.behavior-results.json: normalized runtime assertions, provenance, coverage, and a digest-bound evidence report forskylos agent testruns.
The current harness state is observable and replay-validated. It is not yet a resume mechanism for continuing interrupted verification runs.
Install Options
# Core static analysis
pip install skylos
# LLM-powered agent workflows
pip install "skylos[llm]"
# Ruff Python linting through `skylos lint`
pip install "skylos[lint]"
# All published optional extras
pip install "skylos[all]"Container image:
docker pull ghcr.io/duriantaco/skylos:latest
docker run --rm -v "$PWD":/work -w /work ghcr.io/duriantaco/skylos:latest . --json --no-provenanceThe unqualified image uses Python 3.14. Runtime-specific tags are also published for Python 3.11 through 3.14, so container scans can match a local or CI parser exactly:
docker run --rm -v "$PWD":/work -w /work ghcr.io/duriantaco/skylos:latest-python3.13 . --json --no-provenanceIf a Python file cannot be parsed, Skylos reports analysis_errors, omits the
grade, and exits with code 2 instead of treating the skipped file as clean.
See Installation for source installs, container usage, and optional dependencies.
Configure Templates And Vibe Checks
Run skylos init to add these sections to pyproject.toml:
[tool.skylos]
exclude = ["node_modules", "dist"]
[tool.skylos.templates]
# security = ".skylos/templates/security.md"
# quality = ".skylos/templates/quality.md"
# security_audit = ".skylos/templates/security_audit.md"
# review = ".skylos/templates/review.md"
[tool.skylos.vibe]
extra_phantom_names = ["verify_enterprise_auth"]
extra_phantom_decorators = ["tenant_admin_required"]
extra_credential_names = ["tenant_signing_secret"]
extra_network_timeout_calls = ["vendor_sdk.fetch"]
[tool.skylos.dead_code]
entrypoints = []
[[tool.skylos.dead_code.entrypoints]]
type = "method"
name = ["create", "pre_hook", "post_hook"]
parent = { name = "Main", base_classes = ["Application"] }
path = "src/**"
reason = "project framework lifecycle hook"
[tool.skylos.contribution]
collect_local_signals = false
contribute_public_corpus = false
structural_signatures_only = true
include_source = falseTemplate files extend Skylos' built-in prompts; they do not replace the
JSON-only output contract or untrusted-code safety rules. Vibe dictionary
extensions let teams teach Skylos about local fake-auth helpers, project
credential names, sensitive files, and network calls that must set timeouts.
Dead-code entrypoints let teams mark proprietary framework classes, lifecycle
methods, and decorator-registered functions as live using precise rules for
type, name, path, decorators, base classes, and parent classes.
Rules must include a symbol selector such as name, decorators,
base_classes, or parent; path and module only narrow the match.
Contribution signals are off by default; when enabled, Skylos records local
structural accept/dismiss/learn events under .skylos/contribution/ without raw
source.
By default Skylos discovers [tool.skylos] in pyproject.toml by walking up
from the scan path. To use a dedicated TOML config, pass --config-file PATH
or set SKYLOS_CONFIG_FILE; standalone files may use either [tool.skylos]
or top-level [skylos]. Synced Skylos Cloud policy keeps its protected
precedence over repository-controlled config. The top-level
[tool.skylos].exclude list applies to the main scan and commands such as
skylos debt and skylos clean; pass --exclude for command-local additions
or --include-folder to override an excluded folder.
Language Support
Language | Dead Code | Security | Quality | Local API Proof ( | Notes |
Python | Yes | Yes | Yes | Supported | strongest coverage; framework-aware static analysis and optional tracing |
TypeScript / JavaScript | Yes | Yes | Yes | Supported | Tree-sitter parsing, package graph reachability, framework conventions |
Java | Yes | Yes | Yes | Supported | Tree-sitter parsing, structured security-flow analysis, conservative static-member proof |
Go | Yes | Partial | Partial | Supported | native engine status remains separate from deterministic workspace API proof |
PHP | Yes | Yes | Partial | Unsupported | PHP parser coverage plus taint-style security sinks and sources |
Rust | Yes | Yes | Partial | Unsupported | Rust parser coverage plus security sink/source checks |
Dart | Yes | Yes | Partial | Unsupported | Dart parser coverage plus selected security sinks and sources |
C# | Partial | Partial | Partial | Partial | C# symbols, direct-block unreachable code, selected security sinks, and direct NuGet inventory |
C++ | Partial | No | No | Unsupported | conservative unused file-local functions in |
Kotlin | Yes | Partial | Partial | Unsupported | Kotlin symbol extraction with conservative static-analysis coverage |
Shell | No | Yes | Partial | Unsupported | shell-script security checks for command injection, SSRF, and path traversal |
C# dead-code findings are conservative: in a complete executable or web
application scan, unreferenced public types and methods are low-confidence
candidates; library APIs, protected members, and known framework or configured
entry points remain externally reachable. C# source-file reachability is not
implemented. Quality scanning detects statements after unconditional
return, throw, break, or continue in a direct block; it is not full
control-flow analysis. Security coverage includes selected tainted-input
sinks and generic secret scanning of .cs files when --secrets or -a is
enabled. Interpolated raw-string expressions are not yet followed by C# taint
analysis. SCA inventories direct NuGet PackageReference entries in .csproj
files. Only unconditional exact Version pins ([version]) without local
Update/Remove mutations, not VersionOverride or centrally managed
versions, can be checked against
advisories. Ordinary NuGet version values are minimum bounds, so without a
resolved lockfile their installed versions and vulnerability status remain
unknown. Transitive packages are not inventoried.
Java security analysis follows directly implemented request-data helpers in same-package files or source files identified by exact imports or fully qualified names under a verified local source root. Helper reads are bounded and reject symlinks; no Java code or build scripts are executed. This is not full classpath or recursive helper analysis. Unknown helpers are not assumed to be request sources or sanitizers.
C++ analysis covers .cpp, .cc, .cxx, .hpp, .hh, and .hxx. The first
release reports only apparently unused file-local free functions. Without a
build configuration, it cannot fully resolve templates, overloads, macros, or
external usage; these findings are a conservative heuristic, not a proof of
C++ deadness. Ambiguous .h files and C files are not analyzed as C++.
Java weak-hash checks also follow local algorithm variables and values loaded
through java.util.Properties from literal classloader resources. Resource
lookup stays within the matching src/main/resources or src/test/resources
directory. Missing resources, unsupported loaders/layouts, conflicting branch
values, and unresolved mutations remain unknown. Properties are used only as
crypto evidence, never to prove a security guard or choose a safe branch. This
does not resolve arbitrary runtime classpaths, JAR resources, or environment
overrides.
TypeScript and JavaScript dead code analysis recognizes package.json entry
fields, including bin. For targets under dist/ or out/, it checks the
matching src/ location first, then the package root, before the declared
output. This also covers dist/bin/palee.js mapping to bin/palee.ts and
dist/src/index.js mapping to src/index.ts. If both source locations exist,
the src/ mapping keeps priority; unrelated files are not treated as entries.
For ESM build scripts invoked by package scripts, Skylos also follows top-level esbuild calls using unchanged constants, spreads, templates, Node path helpers, and simple literal-array maps. Build scripts are never executed. Nested build calls and filesystem-generated entry lists remain unsupported and may still produce unused-file findings.
VitePress configs at .vitepress/config.* and .vitepress/config/index.*
are recognised as development entrypoints for .js, .ts, .mjs and .mts.
Other files in .vitepress still need a reference or another entrypoint rule.
Existing directory conventions such as scripts/ work with native Windows
separators too; this does not add general discovery of commands in CI workflows.
Vue single file components (.vue) are skipped by source analysis, including
when passed explicitly. Skylos does not yet parse their <script> or
<script setup> blocks; separate JavaScript, TypeScript and backend source
files are still analyzed. Existing browser script and event references in
templates are unaffected.
Go dead-code and security checks require the native skylos-go engine. If
Skylos discovers Go files but cannot run that engine, the report is marked
incomplete, no grade or clean result is produced, and the CLI exits with status
2. Run skylos doctor to verify engine availability and configure
SKYLOS_GO_BIN when using a separately built engine. The official GitHub
Action builds the matching native engine automatically.
See Rules Reference for rule families and scanner scope.
Config And Deployment Support
Surface | Files | Security Scope |
GitHub Actions |
| dangerous triggers, token permissions, unpinned actions, template injection, secrets, OIDC, cache, and artifact policy |
GitLab CI |
| mutable images, unpinned includes, literal secrets, untrusted eval, Docker-in-Docker, OIDC, cache, timeout, and runner-tag policy |
Dockerfile |
| dangerous |
Edge Docker Compose |
| privileged containers, broad host device/control mounts, GPU/device runtime, and host networking |
Edge systemd |
| root edge services, mutable |
Rendered Kubernetes exposure | one multi-document | opt-in proof from an Ingress annotated |
GPU release contract |
| static target-fleet checks for driver, compute-architecture, and serialized-engine compatibility; no hardware probing |
Benchmark Snapshot
Skylos has checked-in regression benchmarks for dead code, security, quality, and agent review. These are strict regression gates, not broad proof that any tool is universally state of the art.
Suite | Current Skylos Result | Baseline |
Dead code regression | 16 cases, TP=36 FP=0 FN=0 TN=59, score 100.0 | Ruff score 62.67; Vulture not installed in latest local rerun |
Security regression | 56 cases, TP=35 FP=0 FN=0 TN=23, score 100.0 | Bandit score 47.14 on Python-applicable cases |
Quality regression | 13 cases, score 100.0 | regression gate only |
Agent review | 25 cases, score 100.0 | regression gate only |
AI-code defect regression | curated verifier cases for hallucinated references, package APIs, and dependency versions | run |
Frozen golden-v0.2 highlights:
Frozen Suite | Skylos Result | Caveat |
Dead code seeded dev | overall score 96.28; TS/JS/Go/Java score 100.0; Python score 93.33 | Python residuals are label-review items |
Security seeded dev | overall score 96.52; full recall with one Python | label should be reviewed |
OWASP Java security dev | TP=120 FP=0 FN=0 TN=120, score 100.0 | 240-case development subset, not general Java coverage; direct static analysis plus focused CLI checks |
Quality seeded dev | TP=1 FP=0 FN=0 TN=1, score 100.0 | one seeded case only |
For methodology, commands, competitor rows, and caveats, see BENCHMARK.md.
An experimental Jev runner can blindly score the pinned jev-1.13.0 typed
decision model against the checked-in 124 dead-code labels or the frozen
skylos-benchmarks corpus. It withholds labels and review reasons from the
request, retains golden label IDs locally for exact classifier comparison, and
repeats the test with answer-signaling identifiers neutralized when that can be
done without changing fixture semantics. An offline comparator now reports
label-by-label corrections, regressions, abstentions, and unsafe removals
against a frozen Skylos result. Live mode sends
fixture source to TypeSafe, requires an explicit flag and a separate
TYPESAFE_API_KEY; it supports bounded, resumable runs and records request
versus local contract-validation latency. Normal Skylos scans need no Jev key.
We ran a paid six-request Jev-only fresh holdout on 33 frozen labels: 26/33
decisions at confidence 0.8, 24/26 correct among those decisions, and no
known-used symbol classified as unused. A paired paid cascade run on the same
holdout kept F1 at 0.72 while reducing broad-verifier calls from 8 to 4; it
did not improve classification. See the
dead-code benchmark guide.
The 59-label same-run comparison reports total accuracy
for pure Skylos (57.6%), LLM-only (69.5%), the original Jev precheck router
(67.8%), and the initial Jev judge mode (88.1%). After two safety fixes, a
separate paid judge-only run on the final code scored 51/59 (86.4%); it has no
same-run LLM-only arm. The judge threshold was chosen on this synthetic suite,
so these are development results, not an independent holdout or a guarantee
for real repositories.
An expanded 125-label repository-style suite
adds cross-file workflow, installable-package, and async event-bus traps.
Its pure Skylos baseline is 71/125 (56.8%); no paid Jev/LLM result has been
run on that expanded suite yet.
On a harder 19-label dynamic-dispatch challenge, Jev correctly routed all 11 used symbols to the LLM and skipped it for six unused symbols. The paired final score still tied at F1=0.727: both arms retained six false positives from a JSON-configured router. This is a measured limitation, not a claimed accuracy improvement.
Dead-code review is opt-in. A normal skylos . scan stays local and uses no
model. skylos agent verify . uses the LLM verifier by default; choose Jev
explicitly when you want it:
skylos agent verify . --format json # LLM only
skylos agent verify . --dead-code-review jev --format json # Jev only
skylos agent verify . --dead-code-review jev-llm --format json # Jev, then LLM if uncertainJev-only needs TYPESAFE_API_KEY, not an LLM key. Confident Jev "unused"
decisions retain a static finding and confident "used" decisions suppress it;
uncertain or unavailable decisions leave the finding visible as unverified.
Jev-plus-LLM also needs a configured LLM provider and sends uncertain
decisions to the LLM. Jev alone never authorizes --fix. If you select either
Jev mode without TYPESAFE_API_KEY, Skylos stops with setup instructions
instead of silently switching review modes. Get a key by signing in at the official
TypeSafe console (access may require an
invitation), then set TYPESAFE_API_KEY in your environment; do not put it in
source control or a command-line argument.
Jev sends a bounded project source snapshot to TypeSafe (currently at most
64 KB). Check it for embedded secrets and confirm sharing is authorized; the
local file guard is not a secret scanner. Oversized or unsafe snapshots cannot
be judged by Jev. The older --jev-judge and --jev-precheck flags remain
available for compatibility; use --dead-code-review for new workflows.
See the dead-code review guide for key setup,
fallback behavior, and limitations.
Real-project regression testing
liveness_primer, created and maintained by Matthew Digman, is Skylos's official real-project regression testing tool. On every PR, it compares the base and proposed merge result against the same pinned Python projects and reports which findings were added, removed, or changed.
Read the Analyzer Blast Radius check for the comparison and downloadable reports. These results complement the labeled benchmarks above; a change in finding counts alone does not establish accuracy. See the liveness_primer guide for scope, review steps, and reproduction commands.
Project Evidence
Skylos-assisted dead-code cleanup PRs have been merged in Black, NetworkX, Optuna, mitmproxy, pypdf, beets, and Flagsmith. These are accepted cleanup PRs, not project endorsements. See Real-World Results.
A local Astronomer scan on April 26, 2026 computed 420 stargazers and returned overall trust: A. StarGuard also reported low fake-star risk.
Integrations
Integration | Link | Purpose |
GitHub Action | Repository PR gates or optional digest-pinned container-image gates | |
GitLab Code Quality | merge request report artifacts; no comment-posting bot or API token | |
Bitbucket Pipelines / Azure Pipelines | detects pull request context for Skylos Cloud checks; | |
VS Code extension | in-editor findings and AI-assisted fixes | |
Claude Code / Codex / Cursor hooks | check each agent edit, file read, and package install locally | |
MCP server | expose Skylos scans to AI agents and coding assistants | |
Ruff | optional Python linting through | |
Docker image | run Skylos without a local Python install | |
Skylos Cloud | optional upload and dashboard workflows |
Generate a GitHub Actions workflow from the CLI:
skylos cicd init --upload
skylos cicd init --upload --scan-path apps/apiThe generated workflow reviews changed lines on pull requests and uploads full
scans on pushes using GitHub OIDC. It supports monorepo subprojects through
--scan-path.
To scan a built image with the composite Action, install a pinned Trivy version
in the caller's job and set image to a trusted repository@sha256:<digest>
build output plus image-platform. This runs an image-only scan; mode: gate
uses image-fail-on (default high), while mode: scan only reports findings.
See container-image scanning.
Documentation Map
Need | Read This |
Install options, source install, and Docker | |
First scan and core workflows | |
CLI commands, flags, and examples | |
CLI output modes, pretty reports, and TUI controls | |
Optional Ruff linting through the Skylos CLI | |
CI setup, PR gates, annotations, and branch protection | |
GitLab merge request reports and CI example | |
Bitbucket Pipelines and Azure Pipelines pull request checks | |
Dead-code behavior and framework awareness | |
Security scanning and taint analysis | |
Dependency CVEs, uv/npm/pnpm/Poetry/Yarn lockfiles, offline SBOM, and SCA in CI | |
Dependency licenses, SPDX 2.3 SBOM, and license deny/allow policy | |
Digest-pinned container-image scanning and GitHub Action setup | |
Source GPU contracts and built-artifact preflight | |
Rule ID prefixes and product terminology | |
Agent scan, verification, remediation, and model setup | |
AI defense checks and LLM guardrails | |
Claude Code, Codex, and Cursor hooks | |
MCP server setup | |
Real-world merged cleanup PRs | |
Baselines, filtering, suppressions, and whitelists | |
Smart tracing | |
Rule families and language support | |
Cloud uploads and dashboard flow | |
VS Code extension | |
Benchmarks and methodology | |
Security policy | |
Release process | |
Contribution priorities | |
Contributing |
Common Questions
Does Skylos replace Bandit, Semgrep, CodeQL, or Vulture?
No. Skylos can run alongside them. It focuses on framework-aware dead-code signal, PR gating, AI-era regression checks, and a combined workflow across dead code, security, secrets, quality, and AI-defect checks.
Does Skylos require an LLM?
No. Core static analysis runs locally without API keys. LLM features are
optional through skylos[llm] and agent commands.
Does the MCP server need an account?
Without SKYLOS_API_KEY the MCP server only exposes analyze (dead code,
5 calls/day). Every other MCP tool (verify_change, verify_agent,
security_scan, secrets_scan, and so on) returns an authentication error
until the server is started with SKYLOS_API_KEY set; create a key in the
Skylos Cloud dashboard settings. The CLI equivalents (skylos verify,
skylos defend, skylos . -a) and skylos agent install-hooks run fully
locally with no account or key.
Does Skylos replace Ruff?
No. skylos lint is an optional convenience entry point that delegates to
Ruff. Install it with pip install "skylos[lint]"; normal Skylos scans do not
run Ruff or merge Ruff violations into SKY-* findings.
Can I use it only on changed code?
Yes. Use skylos . -a --diff origin/main locally or configure CI gates to focus
on new findings.
How should I handle intentional dynamic code?
Use baselines, whitelists, inline suppressions, or runtime tracing. See the configuration docs and smart tracing docs.
Contributing And Support
Report security issues through SECURITY.md.
Open bugs and false-positive reports with minimal repros.
Check ROADMAP.md for useful contribution areas.
Read CONTRIBUTING.md before sending a pull request.
See QUALITY.md for project quality and gate expectations.
Join the Discord for community support.
License
Skylos is licensed under the Apache License 2.0.
Available Tools
12 toolsanalyzeD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| confidence | No | ||
| exclude_folders | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_fixD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| mode | No | delete | |
| min_safety | No | ||
| apply | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_security_contextD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_triage_suggestionsD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
learn_triageD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| action_id | Yes | ||
| action | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provenance_scanD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| diff_base | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quality_checkD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| confidence | No | ||
| exclude_folders | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remediateD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| max_fixes | No | ||
| dry_run | No | ||
| model | No | gpt-4.1 | |
| test_cmd | No | ||
| severity | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secrets_scanD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| confidence | No | ||
| exclude_folders | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_scanD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| confidence | No | ||
| exclude_folders | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_code_changeA
Validate a code diff for security regressions and issues before it lands.
Takes a unified diff and checks for:
Security control regressions (auth, CSRF, TLS, rate limiting removal)
New dangerous patterns (eval, exec, SQL injection, etc.)
Secrets in added code
AI defense issues in added code
Returns pass/fail with findings.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | Yes | ||
| path | No | . | |
| policy | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It states the tool returns pass/fail with findings and lists checks, but does not disclose whether it is read-only, auth requirements, rate limits, or side effects. For a validation/analysis tool, the assumed read-only behavior is not explicitly confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct and well-structured, using a one-line summary followed by bullet points for specific checks. Every sentence adds value without redundancy. No unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main purpose and checks, and output schema handles return values. However, it omits explanations for the 'path' and 'policy' parameters, which are only present in the schema. Given the complexity of security validation, more contextual detail would help an agent, especially missing guidance on how 'policy' affects validation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning for parameters. It clarifies the 'diff' param expects a unified diff, but does not explain 'path' (optional, default '.') or 'policy' (optional, default null). This leaves two of three parameters undocumented beyond their names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool validates a code diff for security regressions and issues before landing. It lists specific security checks (control regressions, dangerous patterns, secrets, AI defense), distinguishing it from sibling tools like security_scan or remediate, which likely target different stages or scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use before code lands ('before it lands'), but does not explicitly state when to use this tool versus alternatives like security_scan or quality_check. No exclusions or when-not-to-use guidance is provided, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_dead_codeD
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| confidence | No | ||
| model | No | gpt-4.1 | |
| max_verify | No | ||
| max_challenge | No | ||
| exclude_folders | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v4.10.0- Added
analyze - Added
generate_fix - Added
get_security_context - Added
get_triage_suggestions - Added
learn_triage - Added
provenance_scan - Added
quality_check - Added
remediate - Added
secrets_scan - Added
security_scan - Added
validate_code_change - Added
verify_dead_code
TDQS
Scored across 12 tools
Multiple tools appear to overlap in purpose, such as security_scan, secrets_scan, provenance_scan, and quality_check, all of which could be interpreted as scanning or checking code. Without detailed descriptions, an agent could easily misselect among these. Other tools like analyze and get_security_context also have ambiguous boundaries.
All names use snake_case consistently, but the pattern is mixed: some start with a verb (remediate, verify, get, generate, validate, learn), while others start with a noun (security_scan, secrets_scan, provenance_scan, quality_check). This inconsistency could confuse agents expecting a uniform verb_noun structure.
With 12 tools, the server is well within the typical 3-15 range for a security/quality scanning and remediation service. The count feels appropriate for the apparent scope, covering scanning, triage, fixing, and validation without being overwhelming.
The tool set covers a reasonable lifecycle for security scanning and remediation: scanning (security_scan, secrets_scan, provenance_scan, quality_check), analysis (analyze, get_security_context), triage (get_triage_suggestions, learn_triage), fixing (generate_fix, remediate), and validation (validate_code_change). Minor gaps like explicit reporting or configuration are not critical to the core workflow.
Maintenance
Related MCP Connectors
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Zero-install security baseline for AI coding agents — OWASP/CWE-cited rules over MCP.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceA high-performance MCP server providing lightning-fast hybrid code search using TF-IDF and vector embeddings for AI assistants. It enables real-time codebase indexing and semantic retrieval with sub-50ms latency and offline support.12MIT

loctree-mcpofficial
FlicenseNot gradedqualityAmaintenanceStructural code intelligence for AI agents. Scan once, query everything — dead exports, circular imports, dependency graphs, and more. CLI + MCP server.6 npm9-- AlicenseAqualityAmaintenanceAI-powered codebase health analysis — detects dead code, circular dependencies, coupling issues, and architectural drift. 6 MCP tools for Claude Desktop, Cursor, Windsurf, and Slack.647 npmMIT
- FlicenseNot gradedqualityDmaintenanceGive your AI coding agents superpowers — a local MCP server for fast, token-efficient code navigation, search & analysis.-