guardvibe
GuardVibe is a local, deterministic security-scanning MCP server for AI-written code — covering prompts, code, dependencies, secrets, config, auth, compliance, and fixes.
Scan code — inline snippets (
check_code), single files (scan_file), directories (scan_directory), multiple files (check_project), staged files (scan_staged), and git-changed files (scan_changed_files)Prompt-level security —
secure_promptanalyzes a coding prompt before code is written and injects severity-ranked requirements (NO_MOD/LIGHT_MOD/HEAVY_MOD)Dependencies & supply chain — CVE checks via OSV (
check_dependencies,scan_dependencies), package health/typosquat checks (check_package_health), and AI-hallucinated/slopsquatted package detection (scan_hallucinated_packages)Secrets — scan files/directories and full git history for leaked keys, tokens, and credentials (
scan_secrets,scan_secrets_history)Vulnerability analysis — single-file and cross-file taint tracking (
analyze_dataflow,analyze_cross_file_dataflow), auth coverage mapping across routes (auth_coverage), and LLM-powered deep scan for IDOR/business logic/race conditions (deep_scan)Fixes — suggestion patches (
fix_code), verified auto-fix that rolls back regressions (secure_this), and fix verification (verify_fix)Config & host security — app config audits (
audit_config), config-downgrade detection (scan_config_change), MCP config audits (audit_mcp_config), host hardening (guardvibe_doctor,scan_host_config), shell-command risk verdicts (check_command), and policy generation (generate_policy)Compliance & reporting — map findings to SOC2/PCI-DSS/HIPAA/GDPR/ISO27001/EU AI Act (
compliance_report), enforce.guardvibercpolicies (policy_check), export SARIF (export_sarif), and review PRs (review_pr)End-to-end workflows — full audit with PASS/FAIL verdict and hash (
full_audit), mandatory remediation plans (remediation_plan), remediation verification (verify_remediation), repo posture assessment (repo_security_posture), workflow guidance (security_workflow), and cumulative stats (security_stats)Guidance — security best practices (
get_security_docs) and per-rule exploit/fix explanations (explain_remediation)Output formats — markdown, JSON, SARIF, and an agent-actionable contract with exact edits, confidence, and verify steps
Provides security analysis for Clerk authentication, including auth coverage and secure configuration detection.
Provides security analysis for Drizzle ORM, including SQL injection detection and CVE version detection.
Provides security analysis for Express.js applications, including vulnerability scanning and best practices.
Provides security analysis for FastAPI applications, including vulnerability detection and secure coding patterns.
Integrates with GitHub Actions for CI/CD security scanning, including SARIF report generation and PR review annotations.
Provides security analysis for Hono applications, including CVE version detection and secure configuration.
Provides security analysis and remediation for Next.js applications, including auth coverage, config audit, and vulnerability scanning for App Router and Pages Router.
Integrates with pre-commit hooks to block insecure code before it reaches the repository.
Provides security analysis for Prisma ORM, including SQL injection detection and secure query patterns.
Provides security analysis for React applications, including XSS prevention and secure coding patterns.
Provides security analysis for Stripe integrations, including webhook signature verification, secrets detection, and secure coding patterns.
Provides security analysis for Supabase projects, including RLS policy generation, auth detection, and secret scanning.
Provides security analysis for tRPC applications, including vulnerability detection and CVE version intelligence.
GuardVibe
Security infrastructure your AI can't be. No matter how good your coding agent gets, it can't know the CVE published after its training cutoff, it can't deterministically guarantee the same check every run, it can't hold your whole repo in context, and it can't objectively review its own code. GuardVibe does all four — the deterministic, post-cutoff-current, whole-repo, author-independent verification layer for AI-written code.
🗓️ Knows what your AI doesn't. CVE rules refreshed daily from GHSA / OSV.dev / CISA KEV — GuardVibe flags vulnerable dependencies published after your model's training cutoff. (188 CVE rules,
npm run inteldaily triage.)🎯 Deterministic, not probabilistic. Same code = same result, every run (content-hashed). Your AI guesses; GuardVibe doesn't.
🗺️ Sees the whole repo. Cross-file taint + auth-coverage across every route — catches the unprotected endpoint your agent's narrow context missed.
🔍 An independent second pair of eyes. The thing that wrote the code can't review itself. GuardVibe is the outside checker on AI-written code — in the loop while your AI codes (real-time edit hook), not after.
⬅️ NEW: Starts before the first line of code. Every scanner on earth — including your agent reviewing itself — acts after the code exists.
secure_promptacts before: it analyzes the coding prompt itself, detects the stack and attack surfaces it implies, and embeds severity-ranked GuardVibe requirements into the prompt your AI executes. The vulnerability is prevented, not caught. Deterministic, zero LLM calls — and if the prompt is already secure, it passes through untouched.
The security MCP built for vibe coding. 563 security rules, 39 tools covering the entire AI-generated code journey — from the prompt itself to production deployment.
Works with Claude Code, Cursor, Gemini CLI, Codex, VS Code (Copilot), Windsurf, and any MCP-compatible coding agent.
Why a tool, when your AI is so good?
"More rules" was never the moat — a strong model already knows most security rules by heart. What it can't do is be deterministic, know the CVE published after its training cutoff, hold your whole repo in context, or objectively review the code it just wrote. Those four gaps are structural; they don't close as models improve. GuardVibe is the layer that fills them — running while your AI codes, not in a separate audit later. And since v3.19, it runs before your AI codes too: secure_prompt rewrites the task itself so the security requirements are in the prompt, not in the post-mortem.
Related MCP server: supership-scan
Why GuardVibe
Most security tools are built for enterprise security teams. GuardVibe is built for you — the developer using AI to build and ship web apps fast.
563 security rules, 39 tools purpose-built for the stacks AI agents generate
Zero setup friction —
npx guardvibeand you're scanningNo account required — runs 100% locally, no API keys, no cloud
Understands your stack — not generic SAST, but rules that know Next.js, Supabase, Stripe, Clerk, and the tools you actually use
CVE version intelligence — detects 188 known vulnerable package versions in package.json, refreshed every day from GHSA / OSV.dev / CISA KEV
AI agent & MCP security — detects MCP server vulnerabilities, tool-description prompt injection (OWASP MCP Top 10), model-controlled sandbox-disable flags, excessive AI permissions, indirect prompt injection
Auto-fix suggestions —
fix_codetool returns concrete patches and structured edits the AI agent can apply mechanically. Coverage: hardcoded credentials → env-var migration; public-prefix LLM keys (NEXT_PUBLIC_/VITE_/EXPO_PUBLIC_/REACT_APP_) → prefix removal; CORS wildcards → env allowlist;dangerouslyAllowBrowserflags → drop; sandbox bypass flags (unsafe/noSandbox/allowEval) → drop; agent loops → addmaxSteps; raw-HTML React props →<ReactMarkdown>; missing auth checks → insert auth guard; SQL injection → parameterized queries; missing rate limiters / CSRF / security headers → snippet templates.Pre-commit hook — block insecure code before it reaches your repo
CI/CD ready — GitHub Actions workflow with SARIF upload to Security tab
Agent-friendly output — JSON format for AI agents, Markdown for humans, SARIF for CI/CD
Plugin system — extend with community or premium rule packs
New in v3.53.x
CISA KEV Strapi, Capacitor, compression, ProseMirror and Vue SSR gaps closed — v3.53.0 adds
VG1208Strapi unauthenticated private-field filtering that leaks admin password-reset tokens (CVE-2023-22894, in CISA KEV),VG1209Capacitor Android/iOS remote content loaded at the app origin through the internal HTTP proxy path (CVE-2026-103922),VG1210compression memory-leak DoS on premature response close (CVE-2026-87776),VG1211prosemirror-view XSS through pasted HTML (CVE-2026-104847) andVG1212@vue/server-renderer SSR XSS through a carriage return in dynamic attribute names (GHSA-g2v6-rqmx-r4w6). 188 CVE version-pin rules.
New in v3.52.x
Widely installed dependency gaps closed — v3.52.0 adds
VG1203source-map-js indexed source-map offset event-loop DoS (CVE-2026-93749),VG1204proxy-addr IP spoofing through an IPv4-mapped IPv6 trust subnet that matches every client (CVE-2026-90711),VG1205simple-git unsafe-operation guard bypasses — trailer command config, config includes, abbreviated long options and theVISUALeditor (CVE-2026-102826/102827/102828/102829),VG1206SerovalfromJSONthenable assimilation and unbounded TypedArray allocation (CVE-2026-104846/104845) andVG1207Tinypool prototype-pollution gadgets to RCE (CVE-2026-104848/104849). 183 CVE version-pin rules.
New in v3.51.x
figlet, Mockoon, libp2p and dev-tool gaps closed — v3.51.0 adds
VG1198figletwhitespaceBreakinfinite-loop DoS (CVE-2026-96780),VG1199@yeger/turbo-graph unauthenticated network-exposed task execution via/api/run(CVE-2026-59160),VG1200Payload alt-text plugin authorization bypass through an omittedoverrideAccess(CVE-2026-59965),VG1201Mockoon unauthenticated admin API with wildcard CORS (CVE-2026-59148) andVG1202libp2p PeerStore accepting attacker-signed PeerRecords for a victim peer ID (CVE-2026-86039). 183 CVE version-pin rules.
New in v3.50.x
Coverage audit of every CVE rule — each version rule is now compared with the advisories it cites on every published npm release. v3.50.0 closes the three windows that audit found uncovered:
VG1195Next.js 10.x–11.x for the AVIF/libheif image-optimizer RCE (GHSA-2xp9-vwfh-vxw4, critical),VG1196end-of-life @angular/router lines (19.x and older) for the SSR matrix-parameter DoS (CVE-2026-101896), andVG1197the original mcp-from-openapi / FrontMCP$refSSRF window (CVE-2026-39885).
New in v3.49.x
basic-ftp, A2UI, Trigger.dev, @fastify/busboy and probe-image-size gaps closed — v3.49.0 adds
VG1190basic-ftp quadratic directory-listing parser DoS (CVE-2026-102990),VG1191@a2ui/web_coreopenUrljavascript: URI execution from agent-supplied buttons (CVE-2026-10032, critical),VG1192Trigger.dev default secrets / cross-tenant SQL injection / replay IDOR / webhook SSRF cluster (6 advisories, critical),VG1193@fastify/busboy prototype-named header and oversized boundary DoS (CVE-2026-19481, CVE-2026-19484) andVG1194probe-image-size quadratic SVG parser DoS (CVE-2026-104861).
New in v3.48.x
NestJS, Astro, vm2, Piscina and devalue gaps closed — v3.48.0 adds
VG1185@nestjs/platform-fastify absolute-form middleware bypass (GHSA-9c5c-9qcx-q35q),VG1186@astrojs/node malformed Host port crash (CVE-2026-102984),VG1187vm2 sandbox escape cluster residual window (12 advisories, critical),VG1188Piscina ThreadPool options prototype-pollution RCE (CVE-2026-102992, critical) andVG1189devalue shared-memory / uneval expansion / stringifyAsync rejection cluster (3 advisories).
New in v3.47.x
Axios, Fastify, gRPC, NestJS and Angular SSR gaps closed — v3.47.0 adds
VG1180NestJS microservices nested message pattern crash (CVE-2026-102281),VG1181axios prototype-pollution gadget / fetch maxRedirects / HTTP/2 / ReDoS cluster residual window (7 advisories),VG1182@grpc/grpc-js getAuthContext unauthorized certificate (CVE-2026-101916),VG1183Fastify not-found auth bypass and validation bypass cluster (4 advisories) andVG1184Angular router SSR numeric matrix parameter DoS (CVE-2026-101896).
New in v3.46.x
Next.js og-image RCE and widely installed dependency gaps closed — v3.46.0 adds
VG1175Next.jsnext/ogImageResponse remote code execution residual window (GHSA-vcvr-r3jv-pc5j, critical),VG1176Nodemailer addressparser quadratic backtracking DoS past 9.1.0 (2 advisories),VG1177engine.io protocol revision mismatch crash (CVE-2026-102599),VG1178webpack-dev-middleware publicPath path traversal (CVE-2026-76844) andVG1179Electron sandbox inheritance / webview worker / protocol CORS / preload cache cluster (5 advisories).
New in v3.45.x
Widely installed dependency gaps closed — v3.45.0 adds
VG1170Cline Hub dashboard cross-origin WebSocket hijacking (CVE-2026-59723),VG1171@xhmikosr/decompress symlink-chain path traversal (CVE-2026-101894, critical),VG1172brace-expansion stack-exhaustion DoS on nested and comma-chained braces (2 advisories),VG1173undici WebSocket crash / BalancedPool TLS check drop / cache poisoning (3 advisories) andVG1174joiisoDate()quadratic regex DoS (GHSA-6h2x-m376-mqjq).
New in v3.44.x
SSR, package-manager and dev-tool gaps closed — v3.44.0 adds
VG1165Angular SSR infinite-loop DoS on a malformed DOCTYPE (CVE-2026-101895, the releases that fixed last month's SSR advisories),VG1166pnpm 12 pre-release lockfile symlink escape (GHSA-2rx9-3g3h-c2jv),VG1167claude-code-templates Studio server unauthenticated command injection (CVE-2026-73222),VG1168OpenClaw Feishu per-account disablement bypass (2 advisories) andVG1169DOCX editor font-name CSS injection / print XSS (GHSA-x7m8-jrm8-hpvx).
New in v3.43.x
Parser, workflow and MCP gaps closed — v3.43.0 adds
VG1160smol-toml six-byte infinite-loop DoS (CVE-2026-85730),VG1161n8n expression sandbox escape / SSRF / ReDoS / OAuth registration cluster (5 advisories),VG1162FrontMCP / mcp-from-openapi$refSSRF fix bypass (CVE-2026-59973),VG1163node-opcua client socket leak (CVE-2026-68904) andVG1164Plate DOCX export SSRF (CVE-2026-65842).Widest-reach gaps of the month closed — v3.42.0 adds
VG1153nanoid size-overflow that freezes every later ID to one constant string (CVE-2026-73086),VG1154mysql2 plaintext-password auth downgrade,VG1159Nodemailer addressparser quadratic DoS,VG1157toml prototype pollution + recursion crash,VG1158SVGOremoveScriptsXSS bypass, and residual windows for the September pnpm path-traversal / proxy-secret cluster (VG1155) and Orval$refSSRF (VG1156).30-day backlog cleared — every uncovered high/critical advisory of the last month on a package with real reach now has a rule (v3.40.0–v3.41.0): js-yaml, browserslist, multer, engine.io, @xmldom/xmldom, faker, adm-zip, @tiptap/core, @angular/platform-server SSR, @sap/cds-mtxs, mariadb and more. Older hand-written xmldom, RSC and MikroORM rules were regenerated so ranges that resolve past the fix are no longer flagged.
Backlog worked by reach — the daily run now ranks uncovered advisories by how many projects actually install the package. v3.40.0 closed the widest-reach gaps of the last month:
VG1141js-yaml merge-key CPU exhaustion,VG1136browserslist custom-stats prototype write,VG1135engine.io WebTransport crash,VG1137fakerhelpers.fake()code execution,VG1134mariadb password-before-TLS-validation, plus link-preview-js, TOON and LiquidJS.Advisory-accurate version rules — every CVE version-pin rule is checked against the GitHub Advisory Database on every published npm release: ranges that resolve past the fix are never flagged (v3.37.2), releases outside a cited advisory are no longer flagged and missed ones are added, and three rules with no advisory behind them were removed (v3.38.0).
Daily threat-intel pipeline — rule set tracks GHSA / OSV.dev / CISA KEV every day. Latest shipments (v3.39.0) added
VG1129Orval generated-client code-injection RCE cluster (11 advisories, CVE-2026-62681 and more),VG1130Astro AVIF/libheif image-optimization RCE (GHSA-26w7-cxv4-gfx2, CVSS 9.8),VG1131Vendure external-auth account takeover (CVE-2026-63472),VG1132MapLibre GLDOM.sanitize()zero-click XSS (CVE-2026-85061, CVSS 10.0), andVG1133yayson prototype pollution (CVE-2026-61534). v3.37.0 addedVG1128ws memory-exhaustion DoS + uninitialized memory disclosure (CVE-2026-48779 / CVE-2026-45736) and corrected VG917's caret/tilde matching. v3.36.0 addedVG1124@bytebase/dbhubMCP-server DNS-rebinding SQL execution + read-only bypass (CVE-2026-61742/-61788),VG1125Unleash missing-await permission bypass + cross-project IDOR (CVE-2026-77426),VG1126Elysia multipart quadratic-CPU DoS (CVE-2026-56669), andVG1127request-filtering-agent process crash (CVE-2026-62985). v3.35.0 addedVG1119SunEditor sanitizer-bypass stored XSS (CVE-2026-59167, CVSS 10.0),VG1120sharp bundled-libheif RCE residual window (GHSA-rgj7-g3m4-5g8c, 0.33.3–0.35.3),VG1121@roomi-fields/notebooklm-mcpMCP-tool path traversal (CVE-2026-61647),VG11229router LLM-router auth-bypass cluster (CVE-2026-56681/-56675/-56676/-56679), andVG1123deepstream PATCH_MULTI permission-bypass residual (CVE-2026-63116). v3.34.0 addedVG1115Next.js AVIF/libheif RCE + Windows-hosted RCE (GHSA-2xp9-vwfh-vxw4 / CVE-2026-75604, CVSS 9.5),VG1116@clerk/clerk-react5.x org/billing/reverification authorization bypass (CVE-2026-42349),VG1117@zereight/mcp-gitlabunauthenticated file read / SSRF / DNS rebinding (CVE-2026-61560/-61559/-61568), andVG1118PostCSSsourceMappingURLresidual window (GHSA-fxqj-rqcc-2cmp / GHSA-r28c-9q8g-f849). v3.32-3.33 addedVG1109-VG1114: crypto-jsWordArray.random()insufficient entropy, jsii-diff command injection, the keyv/cacheable "ChainDrop" npm supply-chain worm, React2Shellreact-server-dom-*RCE,@trigger.dev/coreprototype pollution, and an axios Basic-auth injection window. Earlier: Auth.js v5 beta fail-open, Next.js/PostCSS July residual windows,@asyncapi/*supply-chain IOC, Clerk 5.x middleware bypass, jscrambler/@injectivelabsIOCs, n8n-mcp cross-tenant isolation, Next.js May 2026 13-advisory cluster, Drizzle ORM SQL identifier injection (CVE-2026-39356),@tanstack/*Mini Shai-Hulud supply-chain attack, Kysely JSON-path traversal, and moreOWASP MCP Top 10 alignment —
VG1068flags MCP / AI tool definitions whosedescription,instructions, orsystemPromptfields carry prompt-injection markers (ignore previous instructions,you are now,jailbreak mode,system prompt:,override safety, …); pair withVG1063which catchesdangerouslyDisableSandbox: truein agent runtimesInline suppress —
// guardvibe-ignore VG001silences individual findings per-lineCLI-first approach —
npx guardvibe audit,npx guardvibe scan,npx guardvibe doctorall work standalone without MCPEmbedded remediation plan —
remediation_plangenerates a section-by-section fix checklist after every auditScore reflects all sections — security score now factors code, dependencies, config, secrets, auth coverage, and taint analysis
Gitignored secrets excluded — files matched by
.gitignoreare automatically skipped during secret scanningTaint sanitizer recognition — dataflow analysis recognizes common sanitizers (DOMPurify, escape functions, parameterized queries) and stops propagation
How GuardVibe Compares
GuardVibe is purpose-built for the AI coding workflow. Traditional tools are excellent for enterprise CI/CD pipelines — GuardVibe fills a different gap.
Capability | GuardVibe | Traditional SAST | Dependency Scanners |
Runs inside AI agents (MCP) | Native | Not supported | Not supported |
Zero config setup |
| Account + config required | Built-in (limited) |
Vibecoding stack rules (Next.js, Supabase, Clerk, tRPC, Hono) | 100+ dedicated | Generic patterns | Not applicable |
AI/LLM security (prompt injection, MCP, tool abuse) | 68 rules | Experimental/None | None |
AI host security (CVE-2025-59536, CVE-2026-21852) |
| Not supported | Not supported |
Auto-fix suggestions for AI agents |
| CLI autofix | Not supported |
CVE version detection | 188 packages, refreshed daily | Extensive | Extensive |
Compliance mapping (SOC2, PCI-DSS, HIPAA) | Built-in | Paid tier | None |
SARIF CI/CD export | Yes | Yes | Limited |
Rule count | 563 (focused, 68 AI-native) | 5000+ (broad) | N/A |
When to use GuardVibe: You're building with AI agents and want security scanning integrated into your coding workflow — no dashboard, no account, no CI setup.
When to use traditional tools: You need deep AST analysis, enterprise dashboards, org-wide policy enforcement, or coverage across hundreds of languages.
Quick Start
Claude Code
npx guardvibe init claudeCreates .mcp.json MCP config (pinned to current version), .claude/settings.json auto-scan hooks, and CLAUDE.md security rules. Restart Claude Code after setup.
Cursor
npx guardvibe init cursorCreates .cursor/mcp.json and .cursorrules with security rules. Restart Cursor after setup.
Gemini CLI
npx guardvibe init geminiCreates ~/.gemini/settings.json MCP config and GEMINI.md security rules.
Codex (OpenAI)
codex mcp add guardvibe -- npx -y guardvibeVS Code (GitHub Copilot)
Create .vscode/mcp.json in your project:
{
"servers": {
"guardvibe": {
"command": "npx",
"args": ["-y", "guardvibe"]
}
}
}Note: VS Code uses
"servers", not"mcpServers".
Windsurf
Add to ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"guardvibe": {
"command": "npx",
"args": ["-y", "guardvibe"]
}
}
}All platforms at once
npx guardvibe init all # Claude + Cursor + GeminiPre-commit hook
npx guardvibe hook install # Blocks commits with critical/high findings
npx guardvibe hook uninstall # Remove hookCI/CD (GitHub Actions)
npx guardvibe ci github # Generates .github/workflows/guardvibe.yml (SARIF scan)
npx guardvibe ci github --pr # + a diff-aware PR review workflow that posts inline commentsWhat GuardVibe Scans
Application Code
Next.js App Router, Server Actions, Server Components, React, Express, Hono, tRPC, GraphQL, FastAPI, Go
Authentication & Authorization
Clerk, Auth.js (NextAuth), Supabase Auth, OAuth/OIDC (state parameter, PKCE) — middleware checks, secret exposure, session handling, SSR cookie auth, admin method protection
Database & ORM
Supabase (RLS, anon vs service role), Prisma (raw query injection, CVEs), Drizzle (SQL injection — including CVE-2026-39356 identifier-injection), MikroORM (CVE-2026-44680 runtime-identifier injection), Kysely (CVE-2026-44635 JSON-path traversal), Turso/LibSQL (client exposure, SQL injection), Convex (auth bypass, internal function exposure)
Payments
Stripe (webhook signatures, replay protection, secret keys), Polar.sh, LemonSqueezy
Third-Party Services
Resend (email HTML injection), Upstash Redis, Pinecone, PostHog, Google Analytics (PII tracking), Uploadthing (auth, file type/size)
AI / LLM Security
Prompt injection detection, LLM output sinks, system prompt leaks, MCP server SSRF/path traversal/command injection, MCP tool description prompt-injection markers (OWASP MCP Top 10 alignment, VG1068), model-controlled sandbox-disable flags (dangerouslyDisableSandbox, VG1063), AI agent unrestricted shell/database access, dangerouslyAllowBrowser, missing maxTokens, agent loop without maxSteps, AI API key client exposure, indirect prompt injection via external data, RAG/vector poisoning, public-prefix LLM key leaks (NEXT_PUBLIC_*, VITE_*, EXPO_PUBLIC_*)
AI Host Security
guardvibe doctor — unified host hardening scanner detecting CVE-2025-59536 (hook injection via .claude/settings.json), CVE-2026-21852 (API key exfiltration via ANTHROPIC_BASE_URL override), MCP config audit, environment scanner, permission analysis. Supports Claude, Cursor, VS Code, Gemini, Windsurf. Host-specific remediation with platform-tailored fix steps.
OWASP API Security
BOLA/IDOR (Broken Object Level Authorization), mass assignment (spread request body, Object.assign), missing pagination, rate limiting, admin endpoint authorization, verbose error leaks
Modern Stack
Zod .passthrough() mass assignment, z.any() bypass, file upload validation, server-only import guard, webhook replay protection, CSP headers, unsafe-inline/unsafe-eval detection, cron endpoint auth
Mobile
React Native, Expo — AsyncStorage secrets, deep link token exposure, hardcoded API URLs, ATS configuration
Firebase
Firestore security rules, Firebase Admin SDK exposure, storage rules, custom token validation
CVE Version Intelligence (188 CVEs, refreshed daily)
Frameworks: Next.js (CVE-2024-34351, CVE-2024-46982, CVE-2025-29927, CVE-2026-23869, CVE-2026-44573 / 44574 / 44575 / 44578 / 44579 / 45109 May 2026 cluster), React + react-server-dom-* (CVE-2025-55182, CVE-2026-23870), Express, Hono pre-4.12.18 cluster, @vitejs/plugin-rsc, Strapi content-type-builder (CVE-2026-22599)
Auth: Clerk middleware bypass (GHSA-vqx2), Clerk has() org/billing/reverification bypass (GHSA-w24r), Clerk clerkFrontendApiProxy SSRF (CVE-2026-34076), NextAuth.js (2 CVEs), jsonwebtoken
ORMs / SQL: Drizzle SQL identifier injection (CVE-2026-39356) + Drizzle sql.raw interpolation (VG1073), MikroORM SQL injection (CVE-2026-44680), Prisma raw-query call-form, Kysely JSON-path traversal (CVE-2026-44635)
AI ecosystem: @anthropic-ai/sdk (CVE-2026-34451 + memory tool path escape), Vercel AI SDK file-type bypass (CVE-2025-48985), LangSmith untrusted prompt manifest (CVE-2026-45134), OpenClaude sandbox bypass (CVE-2026-42074), @nyariv/sandboxjs Function.caller escape (CVE-2026-43898)
HTTP / parsing: Axios pre-1.15.2 cluster (SSRF + prototype-pollution + DoS + CRLF) + axios proxy-auth redirect leak (VG1071), Hono setCookie attribute injection (VG1072, override pinned ^4.12.21), fast-uri path traversal + host confusion (CVE-2026-6321 / 6322), fast-xml-parser CDATA injection, xmldom CDATA, protobuf.js multi-CVE cluster, undici (2 CVEs), ws
Tools / supply chain: node-ipc protestware (VG1069), Miasma @redhat-cloud-services namespace compromise IOC (VG1074), Session messenger exfil endpoint IOC (VG1075), @tanstack/* Mini Shai-Hulud (84 malicious versions, May 2026), @wdio/browserstack-service command injection (CVE-2026-25244), @babel/plugin-transform-modules-systemjs arbitrary code (CVE-2026-44728), @opentelemetry exporter-prometheus DoS (CVE-2026-44902), systeminformation Linux cmd injection (CVE-2026-44724), velocityjs prototype pollution, defu, sharp, lodash, node-fetch, tar, xml2js, crypto-js, angular-expressions RCE, i18next-http-backend, vm2 sandbox breakouts
Deployment & Config
Vercel (vercel.json, cron secrets, headers), Next.js config, Docker, Docker Compose, Fly.io, Render, Netlify, Cloudflare
Infrastructure
Dockerfile security, GitHub Actions CI/CD, Terraform (S3, IAM, RDS, security groups)
Secrets & Environment
API keys (AWS, GitHub, Stripe, OpenAI, Resend, Turso), .env management, .gitignore coverage, high-entropy detection, NEXT_PUBLIC exposure
Compliance Control Mapping
Maps security findings to SOC2, PCI-DSS, HIPAA, GDPR, ISO27001, and EU AI Act (EUAIACT) controls. Identifies which code-level vulnerabilities are relevant to specific compliance requirements. Not a substitute for professional compliance audits.
Supply Chain
Malicious postinstall scripts, unpinned GitHub Actions, CI npm provenance / --ignore-scripts hardening (VG1070), typosquat detection, node-ipc protestware versions (VG1069), Miasma @redhat-cloud-services namespace compromise IOC (VG1074, RHSB-2026-006), Session messenger exfil endpoint IOC (VG1075, filev2.getsession.org), @tanstack/* Mini Shai-Hulud mass-malware versions (May 2026), @wdio/browserstack-service command injection via git branch names (CVE-2026-25244), lockfile poisoning patterns
Prompt-Level Security (Shift Left)
Most vulnerabilities in AI-generated code are born in the prompt: "add login to my app" says nothing about password hashing, session handling, or rate limiting — so the model picks defaults, and the defaults are where the CVEs live. secure_prompt moves the security gate to before code generation: it analyzes the raw prompt, detects the stack and attack surfaces it implies, matches them against GuardVibe's rule set, and returns a directive the host LLM uses to rewrite the prompt with security requirements embedded.
This is not a prompt beautifier. It is deterministic (no LLM, no network), it never restructures intent, and its first job is do no harm: a prompt that is already specific and security-aware gets verdict NO_MOD and passes through untouched.
NO_MOD— prompt is already specific and security-aware → proceed with the original prompt unchangedLIGHT_MOD— intent is clear but security constraints are missing → inject requirements onlyHEAVY_MOD— prompt is vague and security-relevant → inject requirements + surface clarifying questions (never invent the answers)
Before (what the user typed):
add login to my appAfter (what the host LLM executes, having applied the secure_prompt directive):
Add login to my app, with these security requirements:
- [VG001] Use environment variables or a secrets manager — never hardcode credentials.
- [VG1008] Always verify the caller has admin privileges before allowing role elevation.
- [VG105] Always specify allowed algorithms explicitly in jwt.verify().
Before implementing, confirm: which framework/stack is this for, and which auth
provider should be used (e.g. Clerk, Auth.js/NextAuth, Supabase Auth, custom JWT)?Same user intent — but the model now generates auth code with the guardrails stated up front, instead of GuardVibe catching the missing pieces after the fact.
Tools (39 MCP tools)
Tool | What it does |
| Analyze a code snippet for security issues |
| Scan multiple files with security scoring (A-F) |
| Scan a project directory from disk |
| Pre-commit scan of git-staged files — diff-aware (blocks only newly-staged lines; |
| Check all dependencies for known CVEs (OSV) — annotates each vulnerable package with reachability (is it actually imported in your source?) |
| Detect leaked secrets, API keys, tokens |
| Check individual packages against OSV |
| Typosquat detection, maintenance status, adoption metrics |
| Map security findings to compliance controls (SOC2, PCI-DSS, HIPAA, GDPR, ISO27001, EU AI Act) |
| SARIF v2.1.0 export for CI/CD integration |
| Security best practices and guides |
| Auto-fix suggestions with concrete patches for AI agents |
| Close the loop — scan, apply only the fixes that verifiably land (each re-scanned, rolled back on regression), return the verified code + a definition-of-done gate |
| Audit project configuration files for cross-file security misconfigurations |
| Detect project stack and generate tailored security policies (CSP, CORS, RLS) |
| Review PR diff for security issues with severity gating |
| Scan git history for leaked secrets (active and removed) |
| Check project against compliance policies defined in .guardviberc |
| Track tainted data flows from user input to dangerous sinks |
| Cross-file taint analysis — track tainted data across module boundaries |
| Analyze shell commands for security risks before execution |
| Compare config file versions to detect security downgrades |
| Assess overall repository security posture and map sensitive areas |
| Get detailed remediation guidance with exploit scenarios and fix strategies |
| Real-time single-file scan — designed for post-edit hooks |
| Scan only git-changed files — for PRs and incremental CI; diff-aware (only newly-added lines; |
| Cumulative security dashboard — scans, fixes, grade trend over time |
| Host security audit — CVE-2025-59536, CVE-2026-21852, MCP config, env scanner |
| Audit MCP server configurations for hook injection, file:// abuse, sensitive paths |
| Scan shell profiles, .env files for base URL hijack and credential sniffing |
| Verify a security fix was applied correctly — returns fixed/still_vulnerable/new_issues |
| Get recommended tool workflow for your current task (writing, pre-commit, PR review, etc.) |
| Auth coverage map — enumerate routes, parse middleware matchers, detect auth guards, report coverage % |
| LLM-powered deep analysis — IDOR, business logic, race conditions, auth bypass. Defaults to Claude Haiku 4.5 (~cents/scan). Pass |
| Single source of truth — runs ALL checks in one call, returns PASS/FAIL/WARN verdict + score + coverage % + deterministic result hash |
| Remediation plan — generates section-by-section fix checklist after audit |
| Remediation verification — compares before/after audit, flags skipped sections |
| Prompt-level security (shift left) — analyze a coding prompt BEFORE code is written; deterministic triage (NO_MOD/LIGHT_MOD/HEAVY_MOD), stack + attack-surface detection, severity-ranked GuardVibe requirements embedded via a rewrite directive |
| Slopsquat / AI-hallucination detector — flags phantom imports (imported but in no manifest) and typosquats fully offline + deterministic; opt-in online tier adds npm-registry truth (404 = nonexistent, brand-new low-download = slopsquat pattern). CLI: |
All scanning tools support format: "json" for machine-readable output.
Slopsquat / hallucinated-package detection
AI assistants invent package names — ~20% of AI-generated code references packages that don't exist, and attackers register those hallucinated names ("slopsquatting"). Commodity SCA scans known, published packages against vuln databases; it can't see a name that doesn't exist yet, was never installed, or was published yesterday. scan_hallucinated_packages / slopscan targets exactly that seam, at code-gen/PR time:
Offline (deterministic, air-gapped):
phantom_import(a package imported in source but absent from everypackage.json— a classic LLM tell) and typosquats of popular packages. Statement-anchored + comment/template-aware, so example imports in docs/strings are never miscounted.Online (opt-in, graceful degrade): npm-registry truth —
nonexistent(404), brand-new + low-download (slopsquat-registration pattern), deprecated/unmaintained.
The offline tier is also a full_audit section (online never runs inside the audit, keeping the result hash deterministic). Allowlist intentional unpublished/workspace names via .guardviberc:
{ "slopscan": { "online": true, "allow": ["@myorg/internal-pkg"] } }Security Rules (563 rules across 25 modules)
Category | Rules | Coverage |
Core OWASP | 39 | SQL injection, XSS, CSRF, command injection, CORS, SSRF, hardcoded secrets |
Next.js App Router | 18 | Server Actions, secret exposure, auth bypass, CSP, redirects |
Auth (Clerk / Auth.js / Supabase Auth) | 17 | Middleware, secret keys, session storage, role checks, SSR cookies |
Database (Supabase / Prisma / Drizzle) | 13 | Raw queries, client exposure, service role leaks, NoSQL injection, Drizzle identifier injection (CVE-2026-39356) |
OWASP API Security | 11 | BOLA/IDOR, mass assignment, pagination, rate limiting, error leaks |
Modern Stack | 47 | Zod, tRPC, Hono, GraphQL, Uploadthing, Turso, Convex, OAuth, CSP, webhooks, AI SDK, React Server Action validation (React2Shell) |
Deployment Config | 21 | Vercel, Next.js config, Docker Compose, Fly, Render, Netlify, Cloudflare, K8s secrets |
Payments (Stripe / Polar / Lemon) | 9 | Webhook signatures, key exposure, price manipulation |
Services (Resend / Upstash / Pinecone / PostHog) | 11 | API key leaks, PII tracking, email injection |
Web Security | 20 | Webhooks, CSP, .env safety, AI key exposure, cookie handling |
React Native / Expo | 10 | AsyncStorage secrets, deep links, ATS, hardcoded URLs |
Firebase | 7 | Firestore rules, admin SDK, storage, custom tokens |
AI / LLM Security | 33 | Prompt injection, MCP SSRF, excessive agency, indirect injection |
AI Host Security | 14 | CVE-2025-59536 hook injection, CVE-2026-21852 base URL hijack, MCP config audit |
AI Tool Runtime | 14 | MCP tool output sanitization, obfuscated descriptions, safety bypass |
CVE Version Intelligence | 188 | Known vulnerable versions in package.json — incl. Vite dev-server cmd injection (CVE-2024-52011), React Router 7 cluster (CVE-2026-33245/42211/42342), DOMPurify XSS (CVE-2026-47423), Better Auth bypass (CVE-2026-45337), Axios supply-chain backdoor |
Shell / Bash | 5 | Pipe to bash, chmod 777, rm -rf, sudo password |
SQL | 4 | DROP/DELETE without WHERE, stacked queries, GRANT ALL |
Supply Chain | 19 | Malicious install scripts, lockfile integrity, dependency confusion, typosquat detection |
Go | 6 | SQL injection, command injection, template escaping |
Dockerfile | 7 | Root user, secrets in ENV, untagged images, non-root user |
CI/CD (GitHub Actions) | 8 | Secrets interpolation, unpinned actions, write-all permissions |
Terraform | 6 | Public S3, open security groups, IAM wildcards |
Advanced Security | 31 | ReDoS, CRLF injection, race conditions, XXE, brute force, audit logging |
Other Services | 5 | AWS, GCP, MongoDB, Convex, Sentry, Twilio |
CLI Commands
# Scanning
npx guardvibe scan [path] # Scan a directory for security issues
npx guardvibe scan . --format json # JSON output for automation
npx guardvibe check <file> # Scan a single file
npx guardvibe diff [base] # Scan changed files — reports only newly-introduced issues
npx guardvibe diff [base] --all-lines # Include pre-existing findings in changed files too
# Close the loop — scan, apply verified fixes, re-verify
npx guardvibe secure-this <file> # Dry run: show the fixes that would land + remaining manual work
npx guardvibe secure-this <file> --write # Apply only the fixes that re-verify clean (rolled back on regression)
npx guardvibe secure-this <file> --format json
# Full security audit
npx guardvibe audit [path] # Full audit with PASS/FAIL verdict + hash
npx guardvibe audit . --format json # JSON output for CI pipelines
npx guardvibe audit --skip-deps # Skip dependency CVE check
npx guardvibe audit --full # Disable MCP-output truncation (full finding set)
# Host security audit
npx guardvibe doctor # Host hardening audit (project scope)
npx guardvibe doctor --scope host # + shell profiles, global MCP configs
npx guardvibe doctor --scope full # + home dir configs
npx guardvibe doctor --format json # JSON output
# LLM-powered deep scan (IDOR, business logic, race conditions, auth bypass)
npx guardvibe deep-scan <file> # Default: Haiku 4.5, all focus areas
npx guardvibe deep-scan <file> --focus idor # Narrow to IDOR
npx guardvibe deep-scan <file> --model sonnet # Deeper analysis (more expensive)
npx guardvibe deep-scan <file> --max-bytes 5000 # Truncate input for cost control
# Requires ANTHROPIC_API_KEY or OPENAI_API_KEY env var
# Setup
npx guardvibe init <platform> # Setup MCP server (claude, cursor, gemini, all)
npx guardvibe hook install # Install pre-commit hook
npx guardvibe hook uninstall # Remove pre-commit hook
npx guardvibe ci github # Generate GitHub Actions workflow
# Pre-commit / CI
npx guardvibe-scan # Scan staged files (for pre-commit)
npx guardvibe-scan --format sarif --output results.sarif # CI mode
# Options (scan commands)
# --format <type> scan / diff: markdown|json|sarif
# check: markdown|json|sarif|buddy|agent
# agent = guardvibe.agent.v1 — per finding: { id, severity, confidence, exactEdit, manualFix, verify }
# so an AI agent can apply the exact edit and run the verify step to prove the fix
# (an unsupported format errors rather than silently falling back to markdown)
# --output <file> Write results to file
# --fail-on <level> critical|high|medium|low|none — exit 1 when a finding at/above this level exists
# check, audit, and the pre-commit gate (guardvibe-scan / scan --staged) gate on
# critical by DEFAULT; scan and diff are reports (exit 0) unless --fail-on is passed
# --full Bypass response-size caps (50 JSON / 30 markdown / 200-file taint)Plugin System
Extend GuardVibe with custom or community rule packs.
npm install guardvibe-rules-awesomePlugins matching guardvibe-rules-*, @guardvibe/rules-*, or @guardvibe-pro/rules-* are discovered automatically.
Writing a Plugin
A plugin is an npm package that exports a GuardVibePlugin object:
// index.ts
import type { GuardVibePlugin } from "guardvibe/plugins";
const plugin: GuardVibePlugin = {
name: "my-rules",
version: "1.0.0",
description: "My custom security rules",
rules: [
{
id: "CUSTOM001",
name: "My Custom Rule",
severity: "high", // "critical" | "high" | "medium" | "low" | "info"
owasp: "A01:2025 Broken Access Control",
description: "What this rule detects and why it's dangerous",
pattern: /vulnerable_pattern_here/g, // RegExp with global flag
languages: ["javascript", "typescript"], // which file types to scan
fix: "How to fix the vulnerability",
fixCode: "// Copy-paste secure code example",
compliance: ["SOC2:CC6.1"], // optional compliance mapping
},
],
};
export default plugin;Plugin Rule Schema
Field | Type | Required | Description |
| string | Yes | Unique rule ID (e.g., "CUSTOM001") |
| string | Yes | Human-readable rule name |
| string | Yes |
|
| string | Yes | OWASP category mapping |
| string | Yes | What the rule detects |
| RegExp | Yes | Regex pattern to match vulnerable code (use |
| string[] | Yes | File types to scan |
| string | Yes | How to fix the issue |
| string | No | Copy-paste secure code example |
| string[] | No | SOC2/PCI-DSS/HIPAA control IDs |
Loading Plugins
Plugins are loaded from three sources:
Auto-discovery: Any installed npm package matching
guardvibe-rules-*or@guardvibe/rules-*Config-specified: Packages listed in
.guardvibercpluginsarrayLocal paths: Relative paths in
.guardvibercpluginsarray
// .guardviberc
{
"plugins": [
"guardvibe-rules-awesome",
"./my-local-rules"
]
}Configuration
Create a .guardviberc file in your project root:
{
"rules": {
"disable": ["VG030"],
"severity": {
"VG002": "medium"
}
},
"scan": {
"exclude": ["fixtures/", "coverage/"],
"maxFileSize": 1048576
},
"plugins": ["guardvibe-rules-awesome"]
}Inline Suppression
const key = process.env.API_KEY; // guardvibe-ignore VG001
// guardvibe-ignore-next-line VG002
app.get("/api/health", (req, res) => res.json({ ok: true }));Supports //, #, and <!-- --> comment styles.
GuardVibe Scans Itself
We run GuardVibe on its own codebase as a pre-commit hook. Every commit is scanned before it reaches the repository — the same workflow GuardVibe enables for your projects.
How It Works
You write code with AI
|
AI agent calls GuardVibe MCP tools
|
GuardVibe scans locally (no cloud, no API)
|
Returns findings with severity, OWASP mapping, and fix suggestions
|
AI agent fixes issues before they reach productionPerformance
Tested on real AI-built projects (837 files, Next.js + Supabase + Clerk):
Scan time: ~1.2s (837 files)
False positive rate: near zero — context-aware detection (React Native, Supabase client/server, static innerHTML, git-aware secrets)
Detection rate: 100% on known vulnerability patterns
Security score: A (99/100) on production projects
Troubleshooting
MCP connection issues
If your AI agent cannot connect to GuardVibe:
Restart your IDE/agent. MCP servers are started by the host application. After running
npx guardvibe init, restart Claude Code, Cursor, or Gemini CLI for the config to take effect.Check the config path. Run
npx guardvibe init claudeagain and verify the output shows the correct config file location (.mcp.jsonin your project root for Claude Code,.cursor/mcp.jsonfor Cursor).Re-run
initto upgrade. When upgrading GuardVibe, re-runnpx guardvibe init claude—.mcp.jsonis pinned to a specific version (e.g.guardvibe@3.1.36) at init time for fast deterministic startup. As of v3.1.2 the re-run also rewrites stale pins automatically (Upgraded GuardVibe pin (3.1.27 → 3.1.28)); since v3.1.27 the PostToolUse hook command is pinned to the same version (was@latest) and re-run upgrades a stale hook too. The same applies tonpx guardvibe hook installandnpx guardvibe ci github(since v3.1.3) — both are version-pinned at install/generate time and re-run to upgrade.Pre-3.1.1 users won't see the auto-update banner. GuardVibe started writing a once-per-day "newer version available" notice to stderr in v3.1.1. If your install predates that, you'll never see it — run
npx -y guardvibe@latest init <host>once to bake in the latest pin and start receiving banners on subsequent sessions.Verify Node.js version. GuardVibe requires Node.js >= 18.0.0. Check with
node --version.Check npx cache. If you upgraded GuardVibe and the old version is cached, run
npx -y guardvibe@latestto force the latest version.
Node.js version requirements
GuardVibe requires Node.js >= 18.0.0. Earlier versions will fail with syntax errors or missing APIs. Node.js 22 LTS is recommended.
False positives
If a rule triggers on safe code:
Inline suppression: Add
// guardvibe-ignore VG001on the same line, or// guardvibe-ignore-next-line VG001on the line above. Supports//,#, and<!-- -->comment styles.Config exclusion: Add the rule ID to
rules.disablein.guardviberc:{ "rules": { "disable": ["VG030"] } }Path exclusion: Add directories to
scan.excludein.guardviberc:{ "scan": { "exclude": ["fixtures/", "test-data/"] } }
Pre-commit hook issues
Hook not running: Verify the hook file exists at
.git/hooks/pre-commitand is executable (chmod +x .git/hooks/pre-commit).Hook blocking valid commits: Use
git commit --no-verifyto skip the hook temporarily, then investigate the findings.Removing the hook: Run
npx guardvibe hook uninstall.
Security Model
GuardVibe is designed for use on sensitive and proprietary codebases:
100% local execution. All scanning happens on your machine. No code, findings, or metadata are sent to any server.
No accounts, no API keys, no telemetry. There is no signup, no cloud dashboard, and no usage tracking of any kind.
One optional network call. The
scan_dependenciesandcheck_dependenciestools query the OSV API to check for known CVEs. This is opt-in -- you only call it when you explicitly use those tools. No other tool makes network requests.Safe for air-gapped environments. All code analysis rules run entirely offline. Only dependency vulnerability checks require network access.
Configuration (.guardviberc)
Create a .guardviberc JSON file in your project root to customize GuardVibe behavior.
Full example
{
"rules": {
"disable": ["VG030", "VG045"],
"severity": {
"VG002": "medium",
"VG010": "low"
}
},
"scan": {
"exclude": ["fixtures/", "coverage/", "dist/", "vendor/"],
"maxFileSize": 1048576
},
"plugins": [
"guardvibe-rules-awesome",
"./my-local-rules"
],
"compliance": {
"frameworks": ["SOC2", "HIPAA"],
"failOn": "high",
"exceptions": [
{
"ruleId": "VG030",
"reason": "Accepted risk per security review 2026-03",
"approvedBy": "security-team",
"expiresAt": "2026-12-31",
"files": ["src/legacy/**"]
}
],
"requiredControls": ["SOC2:CC6.1"]
},
"scoring": {
"densityModel": "exponential"
}
}Configuration fields
Field | Type | Default | Description |
|
|
| Rule IDs to skip during scanning |
|
|
| Override severity for specific rules |
|
|
| Glob patterns for directories/files to skip |
|
|
| Maximum file size in bytes (files larger than this are skipped) |
|
|
| npm package names or local paths to load as plugins |
|
| -- | Compliance frameworks to map against ( |
|
|
| Minimum severity that causes compliance failure |
|
|
| Approved exceptions with expiration dates |
|
| -- | Controls that must pass regardless of exceptions |
|
|
| Score decay curve. |
Security
GuardVibe takes supply chain security seriously:
npm provenance — every published version is cryptographically signed via Sigstore, linking the package to this exact GitHub repo and commit. Verify with
npm audit signatures2FA enabled — npm account protected with two-factor authentication
Branch protection — force push disabled on main, admin enforcement enabled
Tag protection — version tags (
v*) cannot be deleted or force-pushedMinimal CI permissions — GitHub Actions workflows use
permissions: contents: readonlyMinimal, fully-audited runtime dependencies — only three direct dependencies: the MCP SDK, Zod, and the TypeScript compiler (used for AST-based dataflow analysis). Zod and TypeScript are zero-sub-dependency, pure-JS packages. The MCP SDK pulls a small set of widely-used, audited transitive packages (e.g.
express,cors,ajv) for its optional HTTP transport — GuardVibe itself runs over stdio. No native bindings anywhere in the tree, and all code analysis runs 100% locally and offline
To report a vulnerability, please email info@goklab.com or open a GitHub issue.
License
Apache 2.0 — open source, patent-safe, enterprise-ready. Built by GokLab.
Available Tools
39 toolsanalyze_cross_file_dataflowA
Track user input flowing across module boundaries — detects injection vulnerabilities spanning multiple files. Pass files array with file contents. For single-file analysis, use analyze_dataflow instead. Example: analyze_cross_file_dataflow({files: [{path: 'src/api.ts', content: '...'}, {path: 'src/db.ts', content: '...'}]})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project directory path. When provided, auto-discovers all JS/TS files — no need to pass file contents manually. | |
| files | No | List of files to analyze (ignored when path is provided) | |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It does not explicitly state read-only nature or other behavioral traits, though the purpose implies analysis without side effects. Lacks disclosure of authentication needs or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences plus an example, front-loaded with purpose and usage guidance. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool, it covers purpose, usage, and alternative. Lacks details about output format behavior, but the format parameter is documented in schema. Could mention return format briefly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and includes detailed descriptions. The tool description adds minimal value beyond the schema, except reiterating the files array. Example is helpful but not additive to semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it tracks user input flow across multiple files to detect injection vulnerabilities, and explicitly distinguishes from sibling tool analyze_dataflow for single-file analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear guidance: use for multi-file analysis, with a direct alternative (analyze_dataflow) for single-file, plus an example invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_dataflowB
Track user input (request body, URL params, form data) flowing into dangerous sinks (SQL queries, eval, file operations, redirects). Detects injection vulnerabilities that regex rules miss by following variable assignments through code.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Code to analyze for tainted data flows | |
| format | No | Output format | markdown |
| language | Yes | Language (JS/TS only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It states the tool follows variable assignments through code, but does not mention whether it modifies code, requires authentication, has rate limits, or any side effects. As a static analysis tool, it likely has no side effects, but this is not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action. Every sentence adds value: the first defines the operation, the second explains its benefit over regex rules. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose and methodology well, but lacks details about output format or return value structure. Since no output schema exists, the description could briefly state what the tool returns (e.g., list of vulnerabilities). This is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage, with each parameter described. The description adds no additional parameter meaning beyond the schema, but it does set context for 'code' as the analysis target. Baseline 3 applies due to full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: tracking user input into dangerous sinks to detect injection vulnerabilities. It specifies the types of sinks (SQL queries, eval, file operations, redirects) and distinguishes itself from regex-based approaches, but does not explicitly differentiate from sibling tool 'analyze_cross_file_dataflow' which likely handles multi-file analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for detecting injection vulnerabilities that regex rules miss, but provides no explicit guidance on when to use this tool versus siblings or alternatives. It lacks context on prerequisites, such as requiring complete code or specific language support.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_configA
Audit application config files (next.config, middleware, .env, vercel.json) for cross-file security gaps: missing headers, unprotected routes, exposed secrets. NOT the same as guardvibe_doctor which checks AI host security (MCP configs, hooks). Example: audit_config({path: '.'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Project root directory to audit | |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description mentions security gap checks but does not disclose if the tool is read-only or has side effects. Some transparency but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-structured sentences plus an example. Front-loaded with purpose, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given schema richness (100% covered params) and no output schema, description explains what files and checks are performed. Adequate but could mention output format implications.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3. Description adds example usage but does not provide extra meaning beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it audits application config files (next.config, middleware, .env, vercel.json) for cross-file security gaps, and distinguishes from guardvibe_doctor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'NOT the same as guardvibe_doctor' and provides an example invocation, helping the agent decide when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_mcp_configA
Scan MCP configuration files (.claude/settings.json, .cursor/mcp.json, .vscode/mcp.json) for security issues: malicious hooks (CVE-2025-59536), suspicious MCP servers, overly permissive tool access, and shell injection patterns. Use this to verify MCP configurations are safe before use.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory to scan | . |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description implies read-only scanning but does not explicitly confirm non-destructiveness, side effects, or security considerations. Adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences covering purpose, target files, issues, and usage guidance with no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lacks details on return values or behavior when issues are found; no output schema provided. For a scanning tool with no annotations, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are fully described in the input schema; the description adds no additional meaning beyond what the schema provides. Baseline score applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly defines scanning MCP config files for specific security issues (malicious hooks, suspicious servers, etc.), distinguishing from sibling tools like 'audit_config' that may target different configs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use ('verify MCP configurations are safe before use'), but lacks explicit when-not-to-use or comparison with alternative tools like 'audit_config'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_coverageA
Analyze authentication coverage across Next.js App Router routes. Detects auth guards (Clerk, NextAuth, Supabase, custom) and reports protected vs unprotected routes. Pass files array with route file contents and middleware content. Example: auth_coverage({files: [{path: 'app/api/users/route.ts', content: '...'}], middleware: '...'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project directory path. When provided, auto-discovers all route, page, layout, and middleware files — no need to pass file contents manually. | |
| files | No | Route and page files from app/ directory (ignored when path is provided) | |
| format | No | Output format | markdown |
| middleware | No | Content of middleware.ts file (ignored when path is provided) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It explains the tool detects auth guards (Clerk, NextAuth, Supabase, custom) and reports protected vs unprotected routes, but does not disclose potential side effects (none expected), permissions required, or whether it modifies files. The description is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus a code example, front-loading the purpose and key inputs. Every sentence earns its place: first sentence states purpose, second describes input, example shows usage. No superfluous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain return values. It mentions 'reports protected vs unprotected routes' and an output format parameter, but does not describe the structure of the report (e.g., list of routes, statuses). Given the complexity and 4 parameters, the description is missing details about output shape, making it only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all parameters. The description adds an example call and clarifies that 'files' and 'middleware' are ignored when 'path' is provided. This adds meaning beyond the schema, but the description does not elaborate on each parameter beyond the example. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes authentication coverage across Next.js App Router routes, detecting auth guards and reporting protected vs unprotected routes. It specifies inputs (files array with route contents and middleware), distinguishing it from sibling tools that focus on general code scanning or security audits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an example call and mentions that when path is provided, file contents are not needed, implying usage context. However, it lacks explicit guidance on when to use this tool versus alternatives, such as when a full file-by-file scan is needed or when the project structure is complex.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_codeA
Analyze inline code for security vulnerabilities (OWASP Top 10, XSS, SQL injection, insecure patterns). Pass code as a string parameter. For scanning files on disk, use scan_file instead. Example: check_code({code: 'app.get(...)', language: 'javascript'})
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The code snippet to analyze | |
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| language | Yes | Programming language of the code | |
| framework | No | Framework context (e.g. express, nextjs, fastapi, react, django) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It states the tool analyzes code for vulnerabilities, implying a read-only operation, but does not explicitly mention whether it modifies anything, requires network access, or has other side effects. While not misleading, it lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences: main purpose, usage guideline, and example. Front-loaded with the core function, no extraneous words. Efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (analyze code), the description covers the essential: what it does, how to use it, and distinguishes from a sibling. No output schema is provided, but the example implies return of a security analysis. It could mention the output format more explicitly, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds value by explaining the purpose of the 'format' parameter (markdown vs json) and providing an example that contextualizes the parameters. This goes beyond the schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it analyzes inline code for security vulnerabilities, listing specific categories (OWASP Top 10, XSS, SQL injection). It distinguishes itself from the sibling tool 'scan_file' which scans files on disk, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit instructions: pass code as a string parameter, use 'scan_file' for files on disk. Includes an example call with parameters, providing clear guidance on when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_commandA
Analyze a shell command for security risks before execution. Returns allow/ask/deny verdict with blast radius, safer alternatives, and context-aware risk assessment. Detects: destructive ops, git history rewrites, secret exposure, data exfiltration, deploy triggers, privilege escalation, database drops.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Current working directory | . |
| branch | No | Current git branch (for branch-specific risk) | |
| format | No | Output format | json |
| command | Yes | Shell command to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
In the absence of annotations, the description fully discloses the tool's behavior: it returns a verdict with blast radius, safer alternatives, and risk assessment. It is a read-only analysis tool with no destructive side effects, as implied by the analysis nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at three sentences, front-loading the core purpose in the first sentence and listing outputs and risk types efficiently without extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description thoroughly explains return values (allow/ask/deny verdict, blast radius, safer alternatives, context-aware risk assessment) and enumerates detected risks, providing complete context for an analysis tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already described. The description adds no further detail on the optional parameters (cwd, branch, format) beyond what the schema provides, meeting the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes a shell command for security risks before execution, listing specific risk types. It distinguishes itself from sibling tools focused on files, dependencies, or secrets by being command-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use before running a command ('before execution') but does not explicitly state when not to use it or provide direct alternatives among the many sibling security tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_dependenciesA
Check npm, PyPI, or Go packages for known security vulnerabilities (CVEs) using the OSV database. Use this before adding new dependencies or to audit existing ones.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| packages | Yes | List of packages to check: [{name, version, ecosystem}] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not mention that the tool is read-only, calls an external API, or other behavioral traits. It is adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and purpose, followed by usage context. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (two parameters, no output schema), the description covers purpose and use case well, though a note on return format would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description only echoes schema info (ecosystem and format hints). It does not add new meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks npm, PyPI, or Go packages for CVEs using the OSV database, distinguishing it from sibling tools focused on scanning files or secrets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises using the tool before adding new dependencies or auditing existing ones, but lacks comparison to siblings like scan_dependencies or check_package_health.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_package_healthA
Check npm packages for typosquat risk, maintenance status, adoption metrics, and deprecation. Use this before adding new dependencies to catch suspicious or risky packages.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| packages | Yes | List of package names to check (e.g. ['lodash', 'expres', 'react-qeury']) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that the tool checks for typosquat risk, maintenance status, adoption metrics, and deprecation, implying a read-only operation. It does not mention any destructive behavior or limitations beyond the checks listed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first states the tool's function, the second provides usage guidance. No unnecessary words, front-loaded with purpose, and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is provided, but the description explains the types of metrics returned and the output format options. For a check tool, it covers the necessary context about what is checked and the possible outputs. The presence of similar sibling tools doesn't detract from completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with good parameter descriptions. The description adds value by specifying the types of checks performed (typosquat risk, maintenance, adoption, deprecation) and the output formats (markdown for humans, json for agents), which augments the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks npm packages for specific risks (typosquat, maintenance, adoption, deprecation). It uses a specific verb 'Check' and resource 'npm packages', and differs from sibling tools like check_dependencies or scan_dependencies which are broader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly advises using this tool before adding new dependencies to catch suspicious packages, providing clear context. It does not mention when not to use or alternatives, but the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_projectA
Scan multiple files for security vulnerabilities and generate a project-wide security report with a security score. Use this for comprehensive security audits.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | List of files to scan: [{path, content}] | |
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes core behavior (scan, generate report) but no disclosure of side effects, auth, or rate limits. No annotations to contradict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two-sentence description, no fluff, front-loaded with action and outcome.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lacks output schema, but description mentions report and score. Could elaborate on report contents, but sufficient given parameter coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters with descriptions; description adds no new parameter-level detail. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states scanning multiple files and generating a project-wide security report with a score. Differentiates from single-file scanners like scan_file, but similar to full_audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Says 'Use this for comprehensive security audits' which gives context but lacks explicit when-not-to-use or comparisons to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compliance_reportA
Map security findings to compliance controls (SOC2, PCI-DSS, HIPAA, GDPR, ISO27001, EUAIACT). Scans a directory and groups issues by control. Output includes a summary section at the top; for large projects, findings are truncated to top 50. Use mode=executive for C-level summary. Example: compliance_report({path: '.', framework: 'SOC2'})
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Report mode: full (detailed) or executive (C-level summary) | full |
| path | Yes | Directory to scan | |
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| framework | Yes | Compliance framework |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses truncation to top 50 for large projects, a summary section at top, and different modes. This is adequate behavioral context for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences plus an example, no extraneous information. Front-loaded with key purpose and frameworks, then details. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, behavior (truncation), output structure (summary), modes, and provides an example. No output schema, but return format is implicitly described. Complete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds context for mode ('C-level summary') and an example, but does not significantly expand on schema descriptions. Appropriate score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verb 'map' and clearly identifies the resource: security findings to compliance controls. It lists supported frameworks (SOC2, PCI-DSS, etc.) and distinguishes from sibling scanning tools by focusing on compliance mapping rather than general scanning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States when to use: mapping security findings to compliance frameworks. Provides mode choices (executive for C-level) and an example. Does not explicitly state when not to use or contrast with siblings, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deep_scanA
LLM-powered deep security analysis for vulnerabilities that pattern-matching cannot detect: IDOR, business logic flaws, race conditions, stale auth, mass assignment, privilege escalation. Defaults to Claude Haiku 4.5 (~cents per scan); pass model: 'sonnet' for deeper analysis at higher cost. Requires ANTHROPIC_API_KEY or OPENAI_API_KEY env var.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Code to analyze | |
| focus | No | Focus area — narrows the prompt to a specific vulnerability class | all |
| model | No | LLM model. haiku = fast & cheap (default), sonnet = deeper analysis | haiku |
| format | No | Output format | markdown |
| context | No | Additional context (e.g., 'This is a payment endpoint') | |
| language | Yes | Programming language | |
| maxBytes | No | Max prompt size in bytes — caps cost. Code over this limit is truncated. | |
| existingFindings | No | Already-detected findings to avoid duplicating |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the use of LLM (Claude Haiku 4.5 default), model options with cost implications, required environment variables, and truncation behavior. It lacks details on expected output structure but covers key operational traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief (three sentences) but packed with essential information. It front-loads the purpose and then efficiently adds behavioral details, model options, and prerequisites. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (8 parameters, enums, defaults, no output schema), the description covers the main aspects: purpose, model choice, cost, env vars, and truncation. It does not describe the return value structure, but since no output schema is provided, the parameter 'format' gives some indication. Slightly more detail on expected output would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 8 parameters are described in the schema (100% coverage). The description adds value by providing context beyond the schema, such as the example of passing 'model: sonnet' and explaining that 'maxBytes caps cost' and 'focus narrows the prompt'. This additional guidance enhances parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a clear verb-resource combination: 'LLM-powered deep security analysis' for vulnerabilities that pattern-matching cannot detect. It enumerates specific vulnerability types (IDOR, business logic flaws, etc.), distinguishing it from sibling tools that likely rely on pattern matching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool (for vulnerabilities undetectable by pattern matching) and mentions default model and required API keys. However, it does not explicitly state when not to use it or compare with sibling tools for alternative use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_remediationA
Pass a GuardVibe rule ID (e.g. VG154) to get a detailed explanation: risk assessment, exploit scenario, minimum fix, secure alternative, and test strategy. Optionally pass the affected code snippet for context-aware guidance. Example: explain_remediation({rule_id: 'VG402'})
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | Affected code snippet for context | |
| format | No | Output format | markdown |
| rule_id | Yes | GuardVibe rule ID (e.g. VG001, VG402) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the output content (risk assessment, exploit scenario, etc.) and mentions optional context-aware guidance. No annotations exist, so the description carries the full burden. It is transparent about what the tool returns, though it does not explicitly state that it is read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus an example, efficiently conveying the tool's purpose and optional parameters. It is front-loaded and contains no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately covers the return structure. It lists the components of the explanation and covers input parameters. However, it does not address error handling or behavior for invalid rule IDs, which would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% description coverage, so the baseline is 3. The description adds value by explaining the purpose of rule_id and code in context, and provides an example, enhancing understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides a detailed explanation for a GuardVibe rule ID, listing specific components like risk assessment and exploit scenario. It differentiates from siblings by focusing on explanation rather than planning or fixing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates when to use the tool (for rule ID explanations) but lacks explicit guidance on when not to use it or alternatives. It does not contrast with sibling tools like remediation_plan or fix_code.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_sarifA
Scan a directory and export results in SARIF v2.1.0 format for CI/CD integration (GitHub, GitLab, Azure DevOps). Returns JSON string.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Directory to scan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavioral traits. It mentions the return format (JSON string) but does not state whether the operation is read-only, what side effects occur, or any rate limits. Given the tool's simplicity, the description is adequate but lacks details beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the tool's purpose, output format, and context. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a single input parameter and a simple output (JSON string), the description covers purpose, input, and output. No output schema exists, but the return format is specified. The description is complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the only parameter 'path' is described in the schema as 'Directory to scan'). The description does not add additional semantic meaning beyond what the schema provides. Baseline 3 is appropriate since the schema already covers the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'scan and export', the resource 'directory', and the output format 'SARIF v2.1.0'. It also specifies the context (CI/CD integration for GitHub, GitLab, Azure DevOps), which distinguishes it from sibling tools like scan_directory that likely do not produce SARIF format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates the tool is intended for CI/CD integration, providing context for when to use it. However, it does not explicitly exclude alternative tools or mention when not to use it. For a simple tool, this is clear enough but could be more explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fix_codeA
Pass vulnerable code as a string and get fix suggestions with before/after patches. Returns structured edit instructions (line numbers, severity, confidence). Use verify_fix afterwards to confirm the fix resolved the issue. Example: fix_code({code: '...', language: 'typescript'})
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The code snippet to analyze and fix | |
| format | No | Output format: json (for agent auto-fix) or markdown (human review) | json |
| language | Yes | Programming language of the code | |
| framework | No | Framework context (e.g. express, nextjs, fastapi, react, django) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given no annotations, the description discloses that it returns structured edit instructions with line numbers, severity, and confidence. It also mentions before/after patches. This covers the essential output behavior. However, it does not discuss any side effects, permissions, or limitations (e.g., whether it modifies the input code).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: two sentences and an example. It front-loads the core action and output, and every sentence adds value. No redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool without output schema or annotations, the description covers purpose, usage hint, output structure, and follow-up tool. It is sufficient for an agent to understand how to use it. Minor gap: no mention of asynchronous behavior or failure modes, but not critical for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 4 parameters are described in the input schema (100% coverage). The description adds little beyond restating the schema; it mentions 'vulnerable code' and provides an example but does not clarify nuances like the purpose of 'framework' or 'format' beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it takes vulnerable code and returns fix suggestions with before/after patches. The verb 'fix' and resource 'code' are specific. However, it does not explicitly differentiate from sibling tools like check_code or verify_fix, leaving some ambiguity about when to use this tool over others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear use case ('vulnerable code') and suggests to use verify_fix afterwards. It also gives an example. However, it does not specify when not to use this tool or describe prerequisites beyond the required parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
full_auditA
Single command that runs ALL checks: code scan (429 rules), secret detection, dependency CVEs, config audit, taint analysis, and auth coverage. Returns PASS/FAIL/WARN verdict with deterministic hash. IMPORTANT: If verdict is FAIL or WARN, you MUST call remediation_plan next to get a section-by-section fix checklist — do NOT skip any section. After fixing, call verify_remediation to confirm ALL sections are addressed. Example: full_audit({path: '.'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory | . |
| format | No | Output format | markdown |
| skipDeps | No | Skip dependency vulnerability check | |
| skipSecrets | No | Skip secret scanning |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool runs multiple checks and returns a deterministic hash, but it does not mention potential side effects, performance implications, or required permissions. For a read-heavy tool, this is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph with front-loaded key information ('Single command that runs ALL checks'), followed by a list, verdict description, and imperative workflow steps. Every sentence is essential, and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (running many checks) and the presence of a format parameter, the description explains what the tool does and the post-processing workflow. However, with no output schema, it could better describe the output structure beyond 'PASS/FAIL/WARN verdict with deterministic hash'. Overall, it is quite complete for a high-level tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 4 parameters. The description does not add meaningful detail beyond the schema, except implicitly referencing 'path' in the example. With baseline 3, the description provides no extra parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it runs ALL checks, listing specific types (code scan, secret detection, dependency CVEs, etc.), and returns a verdict with hash. This distinguishes it from the many sibling tools that focus on individual checks, making its comprehensive purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent to call 'remediation_plan' and 'verify_remediation' based on the verdict, providing a clear workflow. It also gives an example usage. However, it does not explicitly mention when NOT to use this tool (e.g., for smaller, targeted checks), but the purpose implies it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_policyA
Auto-detect project stack (Next.js, Supabase, Stripe, Clerk, Prisma, etc.) and generate tailored security policies. Outputs ready-to-use CSP headers, CORS configuration, Supabase RLS policies, rate limiting rules, and security headers based on detected frameworks.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Project root directory to scan | |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does not disclose whether the tool is read-only, modifies files, or has side effects. For a generation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action, and includes key output details. No extraneous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that generates multiple policy types, the description lists key outputs. However, it omits whether it writes files or returns content, and lacks error handling details. Still reasonably complete given the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for both parameters. The description adds context about outputs but does not enhance parameter meaning beyond the schema. Baseline score applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool auto-detects the project stack and generates tailored security policies, listing specific outputs. It differentiates from siblings like 'policy_check' and 'full_audit' by focusing on generation versus checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when generating policies based on detected stack, but does not explicitly state when not to use it or suggest alternatives. Clear enough but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_security_docsA
Get security best practices and remediation guidance for a specific topic, framework, or vulnerability type. Covers OWASP Top 10, framework-specific hardening (Next.js, Supabase, Stripe), and secure coding patterns. Returns actionable guidance with code examples.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | Security topic to look up (e.g. "express authentication", "sql injection prevention", "nextjs csrf", "react xss", "owasp top 10") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It states the tool returns actionable guidance with code examples, implying a read-only operation. However, it does not disclose any behavioral details like authentication needs or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action, and contains no extraneous information. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieval tool with one parameter and no output schema, the description covers purpose, scope, and return type. It could be more explicit about output format limitations, but remains mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds slight context about covered topics but does not significantly extend the parameter meaning beyond the schema's examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves security best practices and remediation guidance for a specific topic, covering OWASP, framework hardening, and secure coding patterns. It distinguishes from siblings like scan_file or check_code which perform active scanning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies using the tool when needing guidance on a security topic, but does not explicitly contrast with siblings like explain_remediation, nor provide when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
guardvibe_doctorA
Check AI host security: MCP configurations, hooks, base URL hijacking, environment variable exposure. NOT the same as audit_config which checks application config files (next.config, .env, headers). Use scope=project (default) for project-only, scope=host to include shell profiles and global AI configs. Example: guardvibe_doctor({scope: 'project'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory | . |
| scope | No | Scan scope: project (default, .claude.json + .cursor/ + .vscode/ + .env), host (+ shell profiles + global MCP configs), full (+ home dir configs) | project |
| format | No | Output format: markdown (human) or json (machine-readable) | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the tool's inspection scope (MCP configs, hooks, etc.) and behavior across scopes. It does not explicitly state it is read-only, but 'Check' implies no side effects. Minor omission: no mention of permissions or network calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences: purpose, sibling differentiation, and usage guidance with example. No redundant words, front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema and no annotations, the description effectively explains what it does, when to use which scope, and how it differs from a sibling. It does not describe return values beyond format options, but the example implies a report. Slightly incomplete for a security check tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed descriptions for all three parameters (path, scope, format). The description adds an example call but does not provide significant new semantics beyond what the schema already offers for path and format. The scope explanation in the schema is equally detailed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Check AI host security: MCP configurations, hooks, base URL hijacking, environment variable exposure' – a specific verb and resource. It clearly distinguishes from sibling 'audit_config' by saying 'NOT the same as audit_config which checks application config files'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use this vs sibling: 'NOT the same as audit_config'. Provides guidance on scope parameter: 'Use scope=project (default) for project-only, scope=host to include shell profiles and global AI configs.' Includes example call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
policy_checkA
Check project against compliance policies defined in .guardviberc. Use this in CI/CD pipelines to enforce security gates, or before releases to verify compliance requirements are met. Validates custom framework requirements, severity thresholds, required controls, and risk exceptions. Returns pass/fail status with detailed findings per control.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Project root directory | |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description implies a non-destructive check by stating it 'Returns pass/fail status' and 'Validates...' but does not explicitly declare it as read-only. Without annotations, this is an adequate baseline, though the agent may need to infer it does not modify state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no fluff: purpose, usage context, and a brief summary of what it validates and returns. Every sentence carries weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple input schema (2 params, no nested objects) and no output schema, the description adequately covers purpose, usage, and output nature (pass/fail with findings). It lacks precise output format detail, but the context suggests it returns a structured result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for both parameters ('Project root directory', 'Output format' with enum). The description adds marginal value beyond the schema, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check project against compliance policies defined in .guardviberc.' It specifies the verb 'check' and resource 'project against compliance policies,' and differentiates from sibling tools by referencing a specific configuration file and use cases like CI/CD gates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly recommends usage in CI/CD pipelines and before releases, and lists what it validates: custom framework requirements, severity thresholds, etc. It does not mention alternatives among siblings, but the guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remediation_planA
Generate a mandatory section-by-section remediation plan from full_audit results. MUST be called after full_audit when verdict is FAIL or WARN. Returns ordered steps for ALL 6 sections (secrets, code, dependencies, config, taint, auth-coverage) with specific tool calls and actions. AI assistants MUST complete every section — skipping sections is not allowed. Example: remediation_plan({path: '.'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory | . |
| format | No | Output format: json for agents (recommended), markdown for humans | json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It reveals the tool returns ordered steps for six specific sections with tool calls and actions, and emphasizes mandatory completion. It does not explicitly state side effects (read-only), but the planning nature implies no destructive actions. This adds context beyond what annotations would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus an example, all front-loaded with the core purpose and key usage rule. Every sentence adds value with no fluff. Highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately describes the output as ordered steps for all six sections with specific tool calls. It also sets context by linking to full_audit verdicts. Could be more detailed on failure modes or edge cases, but sufficient for an AI to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description only provides an example usage with path. It adds no new meaning beyond the schema's descriptions for path and format. Baseline score of 3 is appropriate as the description neither enhances nor detracts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates a mandatory section-by-section remediation plan from full_audit results, specifically after a FAIL or WARN verdict. It lists the six sections and emphasizes completeness, distinguishing it from siblings like full_audit or verify_remediation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: MUST be called after full_audit when verdict is FAIL or WARN. Also provides a constraint: AI assistants MUST complete every section. This gives clear guidance and exclusions, meeting all criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_security_postureA
Analyze a repository's overall security posture. Maps sensitive areas (auth, payments, PII, admin, API, infrastructure), identifies high-risk workflows, recommends guard mode, and lists priority fixes.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Repository root path | |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, but description explains the tool analyzes and recommends (no mutation implied). However, it does not disclose if the tool modifies the repository, requires specific permissions, or has rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences. First sentence states the core purpose, second lists key outputs. No wasted words, front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two simple parameters and no output schema, the description adequately states what it maps, identifies, and recommends. However, it does not describe the return format beyond the format parameter, nor explain 'guard mode' or how priority fixes are presented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% with clear descriptions for both parameters. Description adds context about what the tool examines (sensitive areas list) but does not provide additional details about the parameters beyond what schema includes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a clear verb+resource ('Analyze a repository's overall security posture') and lists specific outputs (maps sensitive areas, identifies high-risk workflows, recommends guard mode, priority fixes). It distinguishes from siblings like 'deep_scan' and 'compliance_report' which have different scopes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. Does not mention prerequisites, when-not-to-use, or contrast with sibling tools like 'deep_scan' or 'compliance_report'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_prA
Review a pull request for security issues. Scans only changed lines (diff-only mode) and produces output for GitHub Check Runs, PR comments, or inline annotations. Supports severity gating to block PRs.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | Base branch to diff against | main |
| path | No | Repository root path | . |
| format | No | Output: markdown (PR comment), json (structured), annotations (GitHub Check Runs) | markdown |
| fail_on | No | Block PR if findings at this severity or above exist | high |
| diff_only | No | Only report findings in changed lines (true) or all findings in changed files (false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full burden of behavioral disclosure. It discloses that it uses diff-only mode, supports multiple output formats, and enables severity gating. However, it does not mention required permissions (e.g., write access to post comments) or whether it modifies any state. The description is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: two sentences that cover the essential aspects without unnecessary details. Every sentence adds value, and the most critical information (purpose and key features) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (5 parameters, no output schema), the description covers core functionality: diff-only scanning, output formats, and severity gating. It omits the structure of the returned findings, which would be helpful for an agent, but the description still provides sufficient context for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% parameter description coverage, so the schema already documents parameter meanings. The description reiterates high-level concepts like 'diff-only mode' and output formats but does not add new meaning beyond the schema's parameter descriptions. Thus, it provides marginal additional value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: reviewing a pull request for security issues. It specifies the diff-only mode, output formats, and severity gating, which are distinguishing features. This differentiates it from siblings like 'scan_changed_files' which may not have the same output integration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Review a pull request for security issues,' which provides clear usage context. However, it does not mention when not to use this tool or suggest alternative tools for different scenarios, such as full scans of the repository. The context is clear but lacks exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_changed_filesA
Scan only files that have changed since a given git ref (branch, commit, or HEAD~N). Ideal for PR checks, pre-push hooks, and incremental CI. Diff-aware by default: returns only findings on newly-added lines (set diff_aware:false for whole changed files).
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | Git ref to diff against (e.g. 'main', 'HEAD~3', commit SHA) | HEAD~1 |
| path | No | Repository root path | . |
| format | No | Output format | markdown |
| diff_aware | No | Report only newly-introduced findings on added lines (true, default) vs. all findings in changed files (false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations present, so description carries full burden. It describes the diff-aware behavior (default returns only new findings on added lines) and the toggle to disable it. Lacks mention of side effects or auth needs, but scan implies read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no waste. Front-loaded with purpose and ideal uses. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but description implies return type (findings on changed lines). Lacks detail on output structure, but given the tool's simplicity and no required params, it is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage 100% so baseline 3. Description adds meaning: explains 'diff_aware' toggle and default 'base' value. Clarifies 'base' accepts branches, commits, or HEAD~N. Adds value beyond basic schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it scans files changed since a git ref, with specific verb 'scan' and resource 'changed files'. Distinguished from sibling tools like scan_file (single file) and scan_staged (staged changes) by focusing on a diff base.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly suggests use cases: PR checks, pre-push hooks, incremental CI. Does not explicitly state when not to use or list alternatives, but the context is clear enough for an agent to infer proper usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_config_changeA
Compare before/after versions of a config file to detect security downgrades: CORS relaxation, CSP weakening, HSTS removal, debug mode, cookie flag changes, TLS disabling, new hardcoded secrets, removed security headers.
| Name | Required | Description | Default |
|---|---|---|---|
| after | Yes | New config file content | |
| before | Yes | Previous config file content | |
| format | No | Output format | json |
| file_path | No | Config file path for context | config |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations present, so description carries full burden. It discloses detection types but omits side effects, permissions, or output behavior. Safe read operation is inferred, not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence but packs many specific checks. No waste, though a bit lengthy. Front-loads verb and purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but description covers the tool's core function and detectable downgrades. Lacks detail on output structure (e.g., format options described in schema). Adequate for a diff tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline 3. Description adds no extra parameter details beyond schema; parameters are self-explanatory from names and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it compares before/after config files to detect security downgrades, listing specific checks like CORS, CSP, HSTS. Distinct from siblings like scan_file or scan_directory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use vs alternatives like scan_file or scan_directory. Implies config diff use case but doesn't exclude other scanning needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_dependenciesA
Parse a lockfile or manifest (package.json, package-lock.json, requirements.txt, go.mod) and check all dependencies for known CVEs via the OSV database. Reads the file directly. Use this after installing dependencies, during CI, or when auditing existing projects for vulnerable packages.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| manifest_path | Yes | Path to manifest file (e.g. 'package.json', 'requirements.txt', 'go.mod') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully bears the burden of transparency. It discloses that the tool 'Reads the file directly' and checks against the OSV database, making its operation clear. There is no contradiction with annotations (none exist).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three sentences: the first states the core functionality, the second adds a key behavioral detail, and the third provides usage guidance. Every sentence earns its place with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description does not fully explain the return value structure beyond output format options. For example, it doesn't specify whether the tool returns a list of CVEs, severity levels, per-dependency results, or a summary. This gap makes the description less complete for an agent to understand what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage for both parameters, meaning the schema already documents them. The description adds value by explaining that the tool reads the file directly (context for manifest_path) and by elaborating on the output format options (human vs machine-readable). This exceeds the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: parse a lockfile/manifest and check dependencies for known CVEs via the OSV database. The verb 'Parse' and resources 'lockfile or manifest' specify the action and target, and the description distinguishes it from sibling tools like 'check_dependencies' by mentioning direct file reading and the OSV database.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides usage context: 'Use this after installing dependencies, during CI, or when auditing existing projects for vulnerable packages.' This covers when to use the tool, though it does not mention when not to use it or suggest alternatives, which would improve the score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_directoryA
Scan all files in a directory on disk for security vulnerabilities. Pass a directory path — reads files from filesystem. Returns security score (A-F) and findings. Results may be truncated for large projects — check fileRanking in JSON output for top files. Example: scan_directory({path: './src'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Directory path to scan (e.g. './src', '.') | |
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| exclude | No | Additional directories to exclude | |
| baseline | No | Path to a previous scan JSON output file for baseline comparison (new/fixed/unchanged findings) | |
| recursive | No | Scan subdirectories |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states that the tool reads files from the filesystem (access behavior), returns security scores and findings, and mentions truncation. However, it does not disclose whether the tool is entirely non-destructive, any authentication or permissions required, or other side effects like network access or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences including an example. It is front-loaded with the main purpose, and every sentence serves a purpose: stating action, explaining input and behavior, noting truncation, and giving an example. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema, the description covers the key output features (security score A-F, findings, fileRanking) but does not enumerate all possible result fields. It provides enough context for basic usage but could be more complete regarding error handling or structure of findings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value with an example usage and a note about fileRanking for truncated results, but does not provide additional semantics beyond what the schema already offers for parameters like recursive, exclude, format, and baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Scan all files in a directory on disk for security vulnerabilities', which is a specific verb+resource action. This distinguishes it from sibling tools like 'scan_file' (single file) and 'scan_dependencies' (dependencies).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context: 'Pass a directory path — reads files from filesystem' and notes that results may be truncated for large projects, suggesting when to check fileRanking. However, it does not explicitly guide when to use this tool vs alternatives like scanning individual files or dependency scans.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_fileA
Scan a single file on disk by path for security vulnerabilities. Pass a file path — the tool reads the file itself. For inline code snippets, use check_code instead. The 'agent' format returns the structured guardvibe.agent.v1 contract (finding + exact edit + confidence + verify step). Example: scan_file({file_path: 'src/api/route.ts', format: 'agent'})
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format. 'agent' = machine-actionable guardvibe.agent.v1 (exact edits + confidence + verify) | json |
| file_path | Yes | Absolute or relative path to the file to scan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries the burden. It discloses that the tool reads the file itself and describes output formats, but lacks details on side effects, permissions, or error handling. Some context is added but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with front-loaded purpose. No wasted words; the example at the end aids understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema, the description covers purpose, usage, param semantics, and behavioral context. Minor omission of error handling but overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, baseline 3. The description adds value by explaining the 'agent' format in detail (structured contract) and providing an example call, which goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans a single file for security vulnerabilities and distinguishes itself from sibling tool 'check_code' for inline snippets. The verb 'scan' and resource 'file on disk' are specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to use this tool (scan file paths) and when not to (use 'check_code' for inline code snippets). It provides a clear alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_hallucinated_packagesA
Detect AI-hallucinated and slopsquatted packages in a repo — the supply-chain seam commodity SCA misses. OFFLINE (deterministic): flags phantom imports (a package imported in source but absent from every package.json — a classic LLM hallucination tell) and typosquats of popular packages. ONLINE (opt-in, default on; gracefully degrades offline): adds npm-registry truth — packages that return 404 (definitive hallucination) and brand-new low-download packages (slopsquat-registration pattern). Run on AI-generated code at PR time, before npm install. Pass online:false for a fully deterministic, air-gapped scan.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Repository root to scan (default current directory) | . |
| format | No | Output format: markdown (human) or json (guardvibe.slopscan.v1 for agents) | markdown |
| online | No | Query the npm registry for existence/age/downloads. false = deterministic offline-only (phantom imports + typosquats). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It details both offline and online behaviors, including graceful degradation when offline. However, it does not explicitly state the tool is read-only or mention any permissions or side effects, but the nature of scanning is implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat lengthy but well-structured with clear sections for offline and online modes. It front-loads the core purpose and provides necessary detail without excessive verbiage. Could be slightly more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the three parameters are fully described in the schema and there is no output schema, the description adequately covers the tool's behavior, modes, and recommended usage. It provides sufficient context for an agent to correctly invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description repeats the offline parameter's effect but adds no new meaning beyond the schema's parameter descriptions. No additional value is provided for the other parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool detects AI-hallucinated and slopsquatted packages, specifying both deterministic offline detection and opt-in online registry checks. It distinguishes itself from sibling SCA tools by calling out what 'SCA misses' and targeting AI-generated code.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises running on AI-generated code at PR time before 'npm install', and provides guidance on when to use offline mode ('air-gapped') or online mode. This clearly tells the agent when to invoke this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_host_configA
Scan host environment for AI security issues: API base URL hijacking (CVE-2026-21852), credential exposure in shell profiles, .env file leaks, and environment variable sniffing. Checks .env files at project scope; add scope=host to also check shell profiles and global AI configs.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory | . |
| scope | No | Scan scope: project (.env files only), host (+ shell profiles, global configs), full (+ home dir) | project |
| format | No | Output format | markdown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description details what is scanned (specific CVEs, credential exposure) and how scope affects coverage. However, without annotations, it fails to disclose whether the tool is read-only, modifies anything, or requires authentication. The safety profile is left assumed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences. The first front-loads purpose and specific threats; the second adds actionable usage guidance. No wasted words or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers parameter usage and scanning behavior adequately, but does not explain the output structure or return values. With no output schema, the agent lacks information about what the scan results look like, leaving a gap in completeness for a 3-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, providing baseline of 3. The description adds value beyond the schema by explaining the behavioral difference between scope values, e.g., 'Checks .env files at project scope; add scope=host to also check shell profiles and global AI configs.' This clarifies parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans host environment for AI security issues, listing specific CVEs and threat types. It distinguishes itself from sibling scanning tools like scan_secrets or scan_directory by focusing on AI-specific host configuration issues.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that default scope checks .env files, and adding scope=host extends to shell profiles and global configs. It provides clear context for different scopes but does not explicitly state when to use this tool over alternatives or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_secretsA
Scan files and directories for leaked secrets, API keys, tokens, and credentials. Detects high-entropy strings, known API key patterns (AWS, Stripe, OpenAI, GitHub, Supabase), exposed .env files, and missing .gitignore coverage. Returns findings with exact line numbers and remediation steps.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | File or directory path to scan | |
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| recursive | No | Scan subdirectories |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description fully carries the burden. It discloses that the tool is read-only, detects specific patterns, and returns findings with line numbers and remediation steps. Could mention performance or permissions but sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, then specifics. No unnecessary words, and structure is clear and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description explains the result includes findings with line numbers and remediation steps. The tool has 3 well-documented parameters and no nested objects, so coverage is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions, so baseline 3. The description does not add meaning beyond the schema for parameters, only reinforces the detection scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Scan files and directories for leaked secrets...' with specific patterns like AWS, Stripe, OpenAI, etc., clearly distinguishing it from general scanning siblings like scan_file or scan_directory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly indicates when to use (detecting secrets, API keys), but does not explicitly state when not to use or provide alternatives, though the sibling list implies specialization.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_secrets_historyA
Scan git history for leaked secrets. Finds secrets that were committed in the past — even if they were later removed. Marks each finding as 'active' (still in code) or 'removed' (in git history only, needs rotation).
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Repository root path | |
| format | No | Output format | markdown |
| max_commits | No | Maximum number of commits to scan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description discloses marking findings as 'active' or 'removed' and need for rotation, but lacks details on permissions, performance impact, or side effects. Adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences are concise and front-loaded with the core action. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Description explains the output classification (active/removed) but does not detail the output format beyond schema enum values. With no output schema, a bit more detail on return structure would improve completeness, but overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%; description adds no extra meaning to parameters beyond schema definitions. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool scans git history for leaked secrets, including those later removed. Distinguishes from sibling tools like 'scan_secrets' by specifying historical scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description implies this is for historical scanning vs current scanning but does not explicitly state when to use or alternatives. No exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_stagedA
Scan git-staged files for security vulnerabilities before committing. Run this before every commit to catch issues early. No input needed — automatically reads staged files. Diff-aware by default: reports only issues on newly-staged lines (set diff_aware:false for whole staged files).
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: markdown (human) or json (machine-readable for agents) | markdown |
| diff_aware | No | Report only findings on newly-staged lines (true, default) vs. all lines in staged files (false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Discloses automated staged file reading, diff-aware default behavior, and output format options. Does not mention that it is read-only or requires a git repository, but these are reasonable assumptions for a scanning tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero waste: purpose, usage guidance, and behavioral note. Front-loaded with the essential verb and resource, efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers all essential aspects given the tool's simplicity: what, when, how to customize. Lacks mention of prerequisite git repository or potential error states, but these are minor omissions for a focused scan tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, baseline 3. Description adds value by explaining format options as 'human-readable' vs 'machine-readable' and clarifying the diff_aware trade-off. This goes beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans git-staged files for security vulnerabilities, with a specific verb and resource. It distinguishes from sibling tools like scan_file or scan_directory by targeting only staged files, and explicitly ties its use to the pre-commit workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly recommends running before every commit, giving clear context. Does not explicitly exclude alternative tools, but the staged file focus is a natural differentiator. Could mention when not to use, but the guidance is sufficient for most scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secure_promptA
Shift-left security at the prompt level: analyze a raw coding prompt BEFORE any code is written and return a structured enhancement directive that embeds GuardVibe security requirements (auth checks, input validation, webhook signature verification, SQL injection prevention, secrets handling) into the prompt you are about to execute. Deterministic — no LLM, no network: triage verdict NO_MOD (prompt already specific and security-aware → proceed with the ORIGINAL prompt unchanged), LIGHT_MOD (inject missing security constraints only), or HEAVY_MOD (also surface clarifying questions — never invent answers to them). Detects stack (Next.js, Supabase, Clerk, Stripe, Prisma, Express, Hono...) and attack surfaces (auth, payments, file upload, user input, SQL, secrets, redirects) from the prompt text, matches them against GuardVibe's rule set, and returns verdict + intent summary + numbered [rule-id] requirements + rewrite directive. Call this with the user's prompt before generating code; prevents vulnerabilities before code generation instead of scanning after. Example: secure_prompt({raw_prompt: 'add login to my app'})
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | Known stack/framework context if the client has it (e.g. 'Next.js app router, Supabase, Stripe') | |
| raw_prompt | Yes | The user's original coding prompt, verbatim |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully details behavioral traits: deterministic, no LLM, no network, triage verdicts, detection of stack and attack surfaces, matching against GuardVibe rule set, and output contents. It also clarifies that it never invents answers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and structured with specific details. It is moderately concise; each sentence adds value, though it could be slightly shorter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description thoroughly explains the return structure (verdict, intent, numbered requirements, rewrite directive) and behavior. It is complete for an agent to understand invocation and expected results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds context (e.g., 'verbatim' for raw_prompt, 'if the client has it' for context) and an example, but this does not significantly extend beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action (analyze a raw coding prompt), the resource (prompt), and the outcome (structured enhancement directive). It distinguishes itself from sibling tools by focusing on pre-code generation security, with specific verdicts NO_MOD, LIGHT_MOD, HEAVY_MOD.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to 'Call this with the user's prompt before generating code' and contrasts with post-generation scanning. It provides an example but lacks explicit statements about when not to use it, though the context of sibling tools implies alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secure_thisA
Close the loop on vulnerabilities in code: scan, apply only the fixes that VERIFIABLY land (each candidate edit is re-scanned and rolled back if it fails to resolve the issue or introduces a new one), and return the verified code plus a definition-of-done gate. Prefer this over fix_code+verify_fix when you want a guarantee the fix landed — not just a suggestion. Returns { status: clean|secured|partial|no_autofix, fixedCode, applied[], remaining[], definitionOfDone:{passed,message}, proofTest }. Write fixedCode to disk, then require definitionOfDone.passed before claiming the task complete; anything in remaining[] needs a manual fix. When fixes were applied, proofTest is a runnable regression test (GuardVibe-as-oracle) you can drop into the project to guard against regressions. Example: secure_this({code: '...', language: 'typescript'})
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The code to scan and secure | |
| filePath | No | File path for context-aware analysis (the file is NOT written; apply fixedCode yourself) | |
| language | Yes | Programming language of the code | |
| framework | No | Framework context (e.g. express, nextjs, react) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description comprehensively discloses behavior: scanning, applying verifiable fixes, re-scanning, rollback on failure, and the return structure. It also specifies post-invocation actions like writing fixedCode and checking definitionOfDone.passed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is fairly long but well-organized, with the main idea front-loaded, followed by details, and an example. Each sentence adds value, though some trimming could improve conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (4 parameters, no output schema), the description is highly complete: it explains the algorithm, return format, and post-invocation steps. It leaves no major gaps for an AI agent to act correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds minimal extra meaning beyond the schema; it mentions 'code' and 'language' in the example but does not elaborate on 'framework' or 'filePath' beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Close the loop on vulnerabilities in code: scan, apply only the fixes that verifiably land...' and explicitly distinguishes it from siblings like fix_code+verify_fix, making it specific and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool versus alternatives: 'Prefer this over fix_code+verify_fix when you want a guarantee the fix landed — not just a suggestion.' It also implies context for use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_statsB
Show cumulative security statistics, grade trend, and vulnerability fix progress for this project. Use this to demonstrate the value of GuardVibe security scanning over time. Data is stored locally in .guardvibe/stats.json.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root path | . |
| format | No | Output format | markdown |
| period | No | Time period for stats | month |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behaviors. It mentions data is stored locally in .guardvibe/stats.json, implying it reads local data. However, it does not state if the tool is read-only, modifies anything, or requires authentication. The lack of side-effect clarity is a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two sentences with no wasted words. The purpose is front-loaded, and the storage detail is relevant. It earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 3 parameters documented in schema and no output schema, the description provides purpose and storage location but lacks detail on return format or specifics of the statistics. It is adequate for a simple tool but incomplete for understanding output behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters (path, period, format) have clear descriptions in the input schema (100% coverage). The tool description adds no additional meaning beyond the schema, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it shows cumulative security statistics, grade trend, and vulnerability fix progress. The verb 'show' is specific, but it does not explicitly distinguish this tool from siblings like 'scan_directory' or 'remediation_plan'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a use case ('demonstrate the value of GuardVibe security scanning over time') but gives no guidance on when not to use it or alternatives among the 35 sibling tools. No prerequisites or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
security_workflowA
Get the recommended GuardVibe tool sequence for your current task. Returns which tools to call, in what order, and with what parameters. Use this when unsure which tool to use. Example: security_workflow({task: 'pre_commit'})
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Current task: writing_code (after edits), pre_commit (before commit), pr_review (reviewing PR), new_project (initial setup), fix_vulnerabilities (fixing known issues), compliance_mapping (audit against framework), dependency_check (check deps), merge_to_main (pre-merge gate), publish_package (pre-publish checks), security_audit (comprehensive audit), incident_response (post-breach investigation), full_remediation (fix ALL security issues across all 6 sections — secrets, code, deps, config, taint, auth) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return type (tools, order, parameters) and gives an example, but does not mention side effects, permissions, or limitations. It is adequate but lacks deeper behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus an example, no unnecessary words. Front-loaded with the core purpose and usage. Highly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is an orchestrator without an output schema, the description adequately explains what it returns and when to use it. It could optionally describe the output format, but the current level is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers the single parameter 'task' with 100% coverage, including detailed enum descriptions. The description adds a usage example but no additional semantic meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns a recommended tool sequence (tools, order, parameters) for the given task. This distinguishes it from sibling tools like analyze_dataflow or check_code, which are individual tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to 'Use this when unsure which tool to use,' providing a clear usage context. It does not list alternatives or exclusions, but the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_fixA
Verify that a specific security fix was applied correctly. Re-scans the updated code and checks if the target vulnerability (by rule ID) is resolved. Returns 'fixed', 'still_vulnerable', or 'new_issues' status with details.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | Updated code after applying the fix | |
| ruleId | Yes | Rule ID to verify (e.g. VG402) | |
| filePath | No | File path for context-aware analysis | |
| language | Yes | Programming language |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description bears full behavioral disclosure. It states the return statuses ('fixed', 'still_vulnerable', 'new_issues') but does not mention side effects, permissions, idempotency, or whether the tool is read-only. This adequate but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no unnecessary words. First sentence states the main purpose; second explains behavior and output. Every word adds value; extremely concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 parameters, no output schema, and no annotations, the description covers input and output well. It explains the verification process and return statuses. However, it omits behavioral details like whether it modifies state or requires previous fix application, and lacks usage context like when to choose this over sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (all parameters described), providing a baseline of 3. The description adds beyond the schema by clarifying that 'ruleId' targets a specific vulnerability and that 'filePath' is for context-aware analysis, which is not detailed in the schema. This adds meaningful context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool verifies a specific security fix by re-scanning code against a rule ID. This distinguishes it from sibling tools like 'fix_code' (applies fix) or 'remediation_plan' (plans). The verb 'verify' and noun 'fix' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage after applying a fix ('after applying the fix') but does not explicitly state when to use versus alternatives like 'explain_remediation' or 'check_code'. No exclusion criteria or alternative tool suggestions are provided, leaving context implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_remediationA
Compare before/after audit results to verify ALL sections were addressed. MUST be called after completing remediation to confirm success. Runs a fresh audit and compares against the before snapshot. Explicitly flags skipped sections and refuses to return 'complete' status unless every section is addressed. Pass the before audit hash or let it re-run. Example: verify_remediation({path: '.', before_hash: 'abc123'})
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Project root directory | . |
| format | No | Output format | json |
| before_hash | No | Result hash from the initial full_audit (for tracking) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses key behaviors: runs a fresh audit, compares snapshots, flags skipped sections, and refuses to return 'complete' unless all are addressed. This is good transparency, though side effects like permissions are missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is concise: three sentences plus an example. Every sentence serves a purpose, no redundancy, and key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description hints at return values ('flags skipped sections', 'complete' status). For a relatively simple verification tool, this provides adequate context, though explicit output structure is not detailed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, baseline 3. Description adds value by explaining that 'before_hash' is for tracking and can be omitted to re-run, and provides an example usage. This clarifies optionality and usage beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Compare before/after audit results to verify ALL sections were addressed.' It uses specific verbs and resources, and distinguishes itself from sibling tools like 'full_audit' by emphasizing post-remediation verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: 'MUST be called after completing remediation to confirm success.' Provides details like passing a before hash or letting it re-run. However, no explicit when-not-to-use or alternative tools are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v3.35.1- Added
analyze_cross_file_dataflow - Added
analyze_dataflow - Added
auth_coverage - Added
check_dependencies - Added
check_project - Added
deep_scan - Added
explain_remediation - Added
full_audit - Added
generate_policy - Added
get_security_docs - Added
guardvibe_doctor - Added
policy_check - Added
scan_changed_files - Added
scan_config_change - Added
scan_dependencies - Added
scan_host_config - Added
scan_secrets - Added
security_stats
18 tool updates
v3.30.0- Removed
analyze_cross_file_dataflow - Removed
analyze_dataflow - Removed
auth_coverage - Removed
check_dependencies - Removed
check_project - Removed
deep_scan - Removed
explain_remediation - Removed
full_audit - Removed
generate_policy - Removed
get_security_docs - Removed
guardvibe_doctor - Removed
policy_check - Removed
scan_changed_files - Removed
scan_config_change - Removed
scan_dependencies - Removed
scan_host_config - Removed
scan_secrets - Removed
security_stats
1 tool update
v3.27.0- Changed
check_dependencies1 field changed- added
Input schema / properties / formatAdded value: +{ + "default": "markdown", + "description": "Output format: markdown (human) or json (machine-readable for agents)", + "enum": [ + "markdown", + "json" + ], + "type": "string" +}
1 tool update
v3.22.0- Added
scan_hallucinated_packages
5 tool updates
v3.21.0- Changed
scan_changed_files1 field changed- added
Input schema / properties / diff_awareAdded value: +{ + "default": true, + "description": "Report only newly-introduced findings on added lines (true, default) vs. all findings in changed files (false)", + "type": "boolean" +}
- Changed
scan_file2 fields changed- changed
Input schema / properties / format / descriptionPrevious value: -"Output format"New value: +"Output format. 'agent' = machine-actionable guardvibe.agent.v1 (exact edits + confidence + verify)" - changed
Input schema / properties / format / enumPrevious value: -[ - "markdown", - "json" -]New value: +[ + "markdown", + "json", + "agent" +]
- Changed
scan_staged1 field changed- added
Input schema / properties / diff_awareAdded value: +{ + "default": true, + "description": "Report only findings on newly-staged lines (true, default) vs. all lines in staged files (false)", + "type": "boolean" +}
- Added
secure_prompt - Added
secure_this
TDQS
Scored across 39 tools
Multiple tools overlap significantly in scope: scan_host_config, guardvibe_doctor, and audit_mcp_config all target AI host/MCP security; scan_file, check_code, scan_directory, and check_project all scan code for vulnerabilities; check_dependencies and scan_dependencies both check for CVEs. The descriptions try to differentiate them, but an agent choosing between three tools for the same host-config task will likely misselect.
Naming mixes verb-first conventions (scan_file, analyze_dataflow, verify_fix) with noun-first conventions (repo_security_posture, security_stats, auth_coverage, remediation_plan), and there is no consistent verb for the same action (scan/check/analyze/audit). Some names like guardvibe_doctor, secure_this, and full_audit break the pattern entirely, making the surface harder to predict.
With 39 tools, the surface far exceeds the 3-15 tool range for a well-scoped server. While the security domain is broad, many tools are near-duplicates (e.g., three config/host scanners, four code scanners), so the count feels inflated rather than meaningfully comprehensive.
The server covers the full security lifecycle remarkably well: scanning (code, deps, secrets, config, host, history), analysis (dataflow, auth coverage, deep LLM analysis), remediation (fix_code, secure_this, remediation_plan, verify_remediation), reporting (compliance, SARIF, stats), and prevention (secure_prompt, generate_policy). The full_audit → remediation_plan → verify_remediation workflow has no dead ends.
Maintenance
Related MCP Connectors
Zero-config MCP security scanner for AI-generated apps. 25K+ vulnerability patterns.
Security scanner for MCP servers. Detect vulnerabilities, prompt injection, and tool poisoning.
Zero-install security baseline for AI coding agents — OWASP/CWE-cited rules over MCP.
Security, SEO and AI-visibility scanner for web apps · free scans and focused checks via MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceSecurity scanner for MCP servers and AI-generated code. Detects leaked API keys, PII, prompt injection, and MCP misconfigs with A-F security grades.MIT
- AlicenseNot gradedqualityFmaintenancePredeploy security scanner for AI-generated code. 80+ vulnerability patterns across secrets, auth, injection, config, Supabase, and logging. Runs locally, code never leaves your machine. Optional x402 witnessed attestation.67 npmApache 2.0
- AlicenseNot gradedqualityCmaintenanceOpen-source AI code review MCP server for local git diff auditing with deterministic security rules and AI-powered analysis using any OpenAI-compatible model.4MIT
- AlicenseNot gradedqualityCmaintenanceSecurity scanner for MCP servers — vet an MCP before you wire it into an agent. Detects prompt-injection, credential exfiltration (via taint analysis), RCE, and supply-chain risks, and catches cross-server exfil chains no single server reveals. Zero-dependency local CLI, SARIF output, CI-gateable, no account.43 npmMIT