Cybersec Toolkit
This MCP server lets an AI agent discover, recommend, check, and safely execute cybersecurity tools from a 670+ tool registry, plus manage remote hosts and run guided security workflows.
Discover tools:
list_tools,get_tool_info,get_module_info,get_profile_tools,list_profiles— filter by module, method, or install status.Check installation:
check_installedverifies tools locally or on a configured remote host.Get recommendations:
suggest_for_ctf,suggest_for_bounty,recommend_install,get_cve_infomap tasks, challenge categories, target types, and CVEs to appropriate tools/modules/skills.Execute tools safely:
run_toolruns registry tools or system utilities with sanitized arguments, timeouts, output caps, and network policy enforcement (external targets off by default).Chain tools:
run_pipelinepipes multiple tools together without a shell (e.g.,strings | grep).Run scripts (opt-in):
run_scriptexecutes Python/Bash, optionally in named Python venvs (e.g., pwntools, crypto) — disabled unless explicitly enabled.Drive security workflows:
guided_assessmentplans and guides assessments (companion mode) or can autonomously solve tasks using the toolchain, with authorization gating and step limits.Manage remote execution:
manage_remote_hostsadds/lists/tests/removes SSH hosts, lettingrun_toolandcheck_installedrun on remote Kali/Linux boxes.
Enables GitHub Copilot in CLI and VS Code to discover, recommend, and execute installed cybersecurity tools through the MCP server.
Allows Ollama-backed MCP-capable clients to interact with the cybersecurity toolkit through an MCP host, including tool discovery and execution.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Cybersec Toolkitrecommend tools for a CTF web challenge"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
/\ /\ ______ __ _____
(o ) ( o) / ____/_ __/ /_ ___ _____/ ___/___ _____
\ \_/ / / / / / / / __ \/ _ \/ ___/\__ \/ _ \/ ___/
<==\ /==> / /___/ /_/ / /_/ / __/ / ___/ / __/ /__
\ V / \____/\__, /_.___/\___/_/ /____/\___/\___/
/_ _\ /____/ by 26zl
|_| ToolkitA security toolkit that AI agents can drive, under rules you set. One command installs 670+ security tools on Linux or Termux. An MCP (Model Context Protocol) server lets Claude Code, Codex, Gemini CLI, OpenCode, and other MCP clients discover, recommend, and run those tools through a governed execution path, inside a disposable Kata Containers VM by default. 872 Agent Skills supply the methodology for CTF, pentest, bug bounty, DFIR, and blue-team work.
Component | What you get |
670+ tools in 18 modules, 14 profiles, and 12 install methods for Debian/Ubuntu/Kali/Parrot, Fedora/RHEL, Arch, openSUSE, and Termux | |
15 tools for discovery, advice, and governed execution; external targets and script execution are off by default | |
One Kata Containers VM per session by default; | |
872 skills (31 project-authored, 841 curated), also installable as a Claude Code plugin |
What makes it different: most toolkits stop at installing tools. Here an AI can also drive them: infer the problem type, pick tools from every module and profile, and work through the problem with you step by step. When you explicitly authorize it, the same toolchain runs an autonomous solver loop. Companion by default; autonomous only when you ask. It complements Kali, Parrot, or BlackArch rather than replacing them: it runs on the machine you already have, Termux included.
Quick start · How it works · Trust & safety · Installer · MCP server · Agent Skills · Development · License
Quick start
1. Install the tools on a supported Linux distro or on Termux:
git clone --depth 1 --branch v1.3.0 https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
./install.sh --doctor # read-only preflight: distro, prerequisites, MCP server, sandbox
sudo ./install.sh --profile ctf # one profile; with no flags, all 18 modules2. Connect an AI client. The tracked configs for Claude Code (.mcp.json), Codex, Gemini CLI, and OpenCode start the server through scripts/mcp-launch.sh, with external targets and script execution disabled. Choose where the tools run:
Mode | Tools run | Host needs | One-time setup |
Sandbox (default) | In a disposable Kata VM, from the sandbox image | Linux with KVM, Docker 23+ with a Kata runtime, Node.js 22+ |
|
Host ( | As your user, with the tools | Add |
Kata needs KVM: on macOS, on Windows, and in VMs without nested virtualization, use host mode.
3. Restart the client. The 15 tools appear (/mcp in Claude Code). Then ask for what you need, for example "triage this binary" or "map the attack surface of my lab at 10.10.0.0/24".
Related MCP server: toolgovern
How it works
Two entry points share one tool registry. An operator runs the bash installer to put tools on disk: on the host, or into the sandbox image at build time. An AI agent talks to the MCP server, which runs inside a Kata VM by default, to discover, recommend, and execute those tools through one governed path. tools_config.json is the single source of truth that the modules define and the MCP advisors read; CI validators keep the Python and bash sides in sync.

Solid arrows are runtime or installation actions; dashed arrows are validation and context relationships. Client configurations enter through the root-aware launcher, which boots the VM through a Sandcastle Kata provider (sandbox/kata.mjs) — or, with --local, starts the server on the host. security.py governs run_tool and run_pipeline through the allowlist, argument checks, and network policy without invoking a shell. run_script is a separate, disabled-by-default capability that runs arbitrary code. Agent Skills stay outside the execution path: .claude/skills/ is canonical, and scripts/sync-skills.sh generates .agents/skills/ for clients that read the portable mirror. Mermaid source: assets/how-it-works.mmd.
Trust & safety
What runs, and what is gated:
Default-safe MCP.
CYBERSEC_MCP_ALLOW_EXTERNAL=0rejects network targets that do not resolve to private or loopback ranges, andCYBERSEC_MCP_ALLOW_SCRIPTS=0disablesrun_script. External scopes and scripting are explicit opt-ins.One gate for governed execution (
mcp_server/security.py): registry allowlist, no shell (create_subprocess_exec, nevershell=True), argument sanitization, a per-tool blocked-flag denylist (e.g.sqlmap --os-shell,nmap -iL, file-list and target-injection flags), target and network policy, rate limiting, output caps, and timeouts. The policy knows enough CLI grammar to tell a target from a header, wordlist, output path, or target-list flag. Tool output reaches the model without terminal escape sequences or LLM control markers, and lines addressed to an AI reader are flagged.Tools run in a disposable VM by default. The launcher boots a Kata Containers VM with its own kernel, no host filesystem beyond an optional
CYBERSEC_SANDBOX_WORKSPACEmount, no Docker socket, a non-root user with every capability dropped, and memory, CPU, and process limits. It is destroyed when the client disconnects. Startup fails closed instead of falling back to the host;--localis the explicit opt-out.Know the limits. The VM is not a network boundary: it reaches whatever its Docker network reaches, and
CYBERSEC_MCP_ALLOW_EXTERNALis a preflight check on resolved addresses, not a firewall (setCYBERSEC_SANDBOX_NETWORK=noneor use a filtered network). Inside the VM, and on your host in--localmode, allowed tools run with the server user's permissions, and some of them spawn child processes or load plugins.Audit trail without leaks. Actions are logged as JSON lines to an owner-only (
0600), rotating log under the user's state directory (~/.local/state/cybersec-tools-mcp/audit.logby default). Script bodies are never stored, only their SHA256 and length, and credential-shaped strings are redacted from tool arguments. A sandboxed server mirrors its records to the same host log, and records are hash-chained, so an edited or deleted record inside a chain shows up inmake audit-verify(limits).Least privilege in the installer. It runs as root but drops to the invoking user (
$SUDO_USER) for cloned-repo builds andpip/cargo/geminstalls; binary releases are SHA256-verified when checksums are published.Dual-use tooling is gated. C2 and phishing frameworks (Sliver, Caldera, gophish, evilginx, …) are off by default and install only with
--include-c2(theredteamandfullprofiles set it); the MCP layer reflects this and never auto-runs them.Authorized use only. See
SECURITY.md, the supply chain model, and the disclaimer.
Installer
Requirements
A supported Linux distro (Debian/Ubuntu/Kali/Parrot, Fedora/RHEL, Arch, openSUSE) or Termux. Runtimes (Python, Go, Ruby, Java, Rust, Node.js), dev libraries, pipx, and build tools are installed automatically. The installer does not run on macOS or native Windows; use WSL or Docker.
Docker is needed only for
--enable-docker(C2 frameworks, MobSF, BeEF, BloodHound, TheHive, Cortex). Install it yourself: Docker Engine docs.GitHub authentication is recommended. The installer downloads ~50 release binaries and makes ~50+ GitHub API calls; unauthenticated requests are limited to 60 per hour, authenticated ones to 5,000. It uses your
gh auth loginsession automatically, also undersudo. Alternatively, pass a personal access token (no scopes needed) throughsudo, which otherwise drops it:sudo --preserve-env=GITHUB_TOKEN ./install.sh.
Install
From the latest release (pinned; recommended):
git clone --depth 1 --branch v1.3.0 https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit && sudo ./install.shFrom main (newest tools and fixes, including unreleased work):
git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit && sudo ./install.shWith no flags, the installer covers the standard tools of all 18 modules. C2/phishing tools and Docker images stay opt-in; --profile full --enable-docker installs everything the current platform supports. To install a subset:
sudo ./install.sh --profile ctf # CTF tools only
sudo ./install.sh --profile redteam --enable-docker # Red team + Docker C2
sudo ./install.sh --module web --module recon # Specific modules
sudo ./install.sh --tool sqlmap --tool nmap # Individual tools
sudo ./install.sh --dry-run --profile ctf # Preview without installing
./install.sh --doctor # Read-only preflight; no root neededsudo ./install.sh --help # Full help
sudo ./install.sh --list-profiles # Show profiles
sudo ./install.sh --list-modules # Show modules
sudo ./install.sh --skip-heavy # Skip large/slow packages
sudo ./install.sh --skip-pipx # Skip all pipx (Python) installs
sudo ./install.sh --skip-go # Skip all Go tool installs
sudo ./install.sh --skip-cargo # Skip all Cargo (Rust) installs
sudo ./install.sh --skip-gems # Skip all Ruby gem installs
sudo ./install.sh --skip-git # Skip all git clone installs
sudo ./install.sh --skip-binary # Skip all binary release downloads
sudo ./install.sh --skip-source # Skip build-from-source, snap, npm, and curl-pipe installs
sudo ./install.sh --fast # Skip checksum verification (see Supply chain model)
sudo ./install.sh --require-checksums # Fail if a binary release has no checksum file
sudo ./install.sh --production # Strict checksum preset for release downloads
sudo ./install.sh --upgrade-system # Upgrade system packages before installing
sudo ./install.sh --list-sessions # List install sessions and exit
sudo ./install.sh --rollback <id|last> # Roll back tools installed in a session
sudo ./install.sh --version # Show installer version and exit
sudo ./install.sh --enable-docker # Pull Docker images
sudo ./install.sh --include-c2 # Include C2/phishing frameworks (Empire also needs --enable-docker)
sudo ./install.sh -j 8 # 8 parallel install jobs (default: 4)
sudo ./install.sh -v # Verbose / debug output--tool installs only the named tool, without the full dependency setup. Dry-run time estimates count install entries across methods, so they can exceed the de-duplicated 670+ tool registry.
The time goes to I/O-bound work that no scripting language can speed up:
What takes time | Why |
System packages (apt/dnf) | Downloading and unpacking |
Go tools | Downloading modules and compiling each binary |
pipx (Python) | Creating one isolated venv per tool and downloading wheels |
Cargo (Rust) crates | Compiling from source |
Git clones | Cloning each repository |
Binary releases | Downloading pre-built binaries from GitHub |
./install.sh --dry-run prints the live per-method breakdown. pipx, Go, git, and binary installs run in parallel (-j 4 by default); the system package manager and Cargo run sequentially. To go faster: install a profile or --module instead of everything, --skip-cargo to avoid Rust compilation, -j 8 for more parallel jobs, and an apt-cacher-ng proxy for repeated installs.
Docker and Podman
The prebuilt image is the installer on Ubuntu, not a sandbox. Its default command is a dry run of the full profile:
docker run --rm ghcr.io/26zl/cybersec-toolkit # preview only
docker run -it --name cybersec --entrypoint bash ghcr.io/26zl/cybersec-toolkit # keep a container
sudo ./install.sh --profile ctf # inside it; reattach with: docker start -ai cybersecBuild it yourself with docker build -t cybersec-toolkit ., or run a throwaway install through the bundled Compose file with docker compose run --rm installer --profile ctf. On Apple silicon, add --platform linux/amd64 to docker build and docker run. Podman works as a drop-in replacement (podman compose needs a compose provider). The image grants the toolkit user passwordless sudo so the installer can manage packages; treat code inside it as root-capable.
Termux (Android)
pkg install git
git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
./install.sh --profile lightweightProfiles
Profile | Modules | Description |
| All 18 | Complete security toolkit |
| misc, crypto, pwn, reversing, stego, forensics, cracking, web, mobile, blockchain | CTF competitions |
| misc, networking, recon, web, enterprise, pwn, mobile, cracking, cloud, wireless, reversing, crypto | Offensive security |
| misc, networking, recon, web, llm | Web application testing |
| misc, recon | OSINT gathering |
| misc, forensics, blueteam, reversing, stego, cracking | Digital forensics and incident response |
| misc, pwn, reversing, crypto | Binary exploitation and reverse engineering |
| misc, mobile, web, reversing | Mobile application security testing |
| misc, cloud, containers, networking, recon | Cloud and container security auditing |
| misc, blockchain, web, crypto | Smart contract auditing and blockchain security |
| misc, wireless, networking | WiFi, Bluetooth, and SDR security |
| misc, networking, recon, web, cracking | Hobby ethical hacking essentials (HTB, THM, bug bounty) |
| misc, cracking, crypto | Hash cracking |
| misc, blueteam, forensics, reversing, mobile, containers, networking, cloud, recon | Defensive security, IR, malware analysis |
Modules
Module | Tools | Description |
| 40 | Post-exploitation, social engineering, wordlists, resources, C2 (Docker + Loki) |
| 57 | Port scanning, packet capture, tunneling, MITM, protocol tools |
| 82 | Subdomain enumeration, OSINT, DNS, automated recon frameworks |
| 60 | Vulnerability scanning, fuzzing, SQLi, XSS, CMS scanners, API testing |
| 15 | RSA attacks, cipher analysis, hash attacks, constraint solving |
| 36 | Exploit frameworks, binary exploitation, fuzzing, payload generation |
| 33 | Disassemblers, debuggers, emulation, Java/Python reversing |
| 57 | Disk/memory forensics, file carving, timeline analysis, log analysis, hardware/serial |
| 80 | Active Directory, Kerberos, Azure AD, credential harvesting, lateral movement |
| 41 | WiFi cracking, Bluetooth, SDR, rogue AP |
| 34 | Hash cracking (john, hashcat), brute force, wordlist generation |
| 15 | Image/audio steganography, detection, StegCracker |
| 22 | AWS/Azure/GCP security auditing, Checkov |
| 15 | Docker/Kubernetes security (Grype, Syft, Kubescape, kubeaudit) |
| 36 | IDS/IPS, SIEM, incident response, threat intelligence, hardening, malware analysis (YARA, ClamAV, FLOSS, Capa, Loki) |
| 18 | Android/iOS app testing, APK analysis, MobSF (Docker) |
| 15 | Smart contract auditing (Slither, Mythril, Foundry, Aderyn), blockchain forensics, Echidna (Docker) |
| 14 | LLM red teaming, prompt injection, jailbreak testing, AI vulnerability scanning |
Method | Count | Examples |
Git clone | 193 | GitHub repos with auto-setup, resources, wordlists |
System packages (apt/dnf/pacman/zypper) | 166 | nmap, wireshark, john, hashcat |
pipx | 137 | sqlmap, impacket, bloodhound, volatility3 |
Go install | 62 | nuclei, subfinder, ffuf, httpx |
Binary release | 51 | gitleaks, chainsaw, findomain, FLOSS, Capa, Loki, Syft, Kubescape |
Build from source | 23 | massdns, duplicut, AFLplusplus, honggfuzz |
Docker | 13 | Empire, MobSF, BeEF, BloodHound, TheHive, Cortex, PentAGI |
Cargo (Rust) | 8 | feroxbuster, RustScan, pwninit, yara-x-cli |
Ruby gem | 6 | wpscan, evil-winrm, brakeman |
npm | 5 | promptfoo, apk-mitm, surya, solgraph |
Special | 5 | Metasploit, Foundry, Steampipe, patator (curl-pipe installers), crypto venv bootstrap |
Snap | 1 | zaproxy |
Post-install scripts
All scripts support --help and need root on Linux (sudo); Termux needs no root.
Script | Purpose | Example |
| Check which tools are installed |
|
| Update all installed tools |
|
| Remove tools by module |
|
| Purge caches and build artifacts |
|
| Back up and restore tool configs |
|
--deep-clean removes the Go module/build cache, the Cargo registry, pip/pipx/npm/gem caches, orphaned pipx venvs, stale symlinks, and log files. Add --remove-deps to also purge Rustup toolchains.
Non-system tools land in /usr/local/bin/ on Linux and $PREFIX/bin on Termux; system packages use their default location (/usr/bin/).
Method | Binary location (Linux) | Binary location (Termux) | Data location |
pipx |
|
|
|
Go |
|
|
|
Cargo |
|
|
|
Git repos |
|
|
|
Binary releases |
| Skipped (glibc incompatible with Bionic) | -- |
The .versions file records what was installed, how, and when.
If --enable-docker is set and Docker is missing, the installer stops and asks you to install Docker first.
Image | Module | Flag | Description |
| misc |
| Empire C2 |
| misc |
| SpiderFoot OSINT |
| web |
| BeEF browser exploitation |
| mobile |
| MobSF |
| enterprise |
| BloodHound CE |
| blockchain |
| Echidna smart contract fuzzer |
| cloud |
| KICS infrastructure-as-code scanner |
| crypto |
| SageMath (Coppersmith, Groebner bases, curve arithmetic) |
| blueteam |
| TheHive IR platform |
| blueteam |
| Cortex analysis |
| blueteam |
| Zeek network analysis |
| blueteam |
| Zircolite EVTX detection |
| llm |
| PentAGI autonomous pentesting |
Distro support
Debian/Ubuntu/Kali is the primary target and has the strongest test coverage: Kali and Parrot get the full apt set, while plain Debian/Ubuntu skip a handful of Kali-only packages. Fedora, Arch, and openSUSE skip the packages their repositories do not carry (the - entries in lib/distro_compat.tsv, several dozen per distro) and run in the integration workflow.
Platform | Status |
WSL | Supported for installs and MCP use; the wireless module and kernel-level packages are skipped. No dedicated CI job, so validate release-critical changes in a local WSL distro. See Windows Defender false positives if the repo lives on a Windows-mounted path. |
ARM (aarch64/armv7) | Supported, with automatic skips for x86-only binary releases and build-from-source tools. No dedicated CI job. |
Termux (Android) | Supported without sudo. Docker, snap, binary releases, and build-from-source are skipped (Bionic incompatible). No dedicated CI job. |
Windows (native) | Not supported. Use WSL. |
macOS | The installer is not supported; use the Docker image. The MCP server runs on macOS in host mode ( |
Other Linux distros | Distros without |
Supply chain model
The installer downloads and runs code from the internet. On Linux it runs as root (sudo); on Termux it runs in the app's user sandbox.
System packages: signed by your distro's repositories (apt, dnf, pacman, zypper, pkg).
pipx/Go/Cargo/Gem/npm: fetched from their registries without signature verification; pipx tools are isolated in venvs.
Binary releases: SHA256-verified when the release publishes a checksum file, with a hard failure on mismatch. About half of the upstream releases publish none;
--require-checksumsor the--productionpreset fails those tools instead of installing them unverified.--fastdisables all checksum verification, including for releases that publish checksums, and cannot be combined with the strict flags; keep it out of CI and production.Runtime bootstraps: the rustup, uv, and cargo-binstall installers and NodeSource's setup script (run as root) are fetched over HTTPS and sanity-checked before they run, but not signature-verified. cargo-binstall then downloads prebuilt Rust binaries where available;
--skip-sourceskips its installer on a fresh host, so Rust tools compile from crates.io.Go SDK: SHA256-verified against go.dev when the API is reachable; strict mode fails if it is not.
Git repos: cloned at HEAD; dependencies go into isolated venvs, and
setup.pyis not executed.Build from source: runs
make, as root on Linux. Review what you build.MCP Python dependencies: resolved by
uvwith a 3-dayexclude-newerwindow, and Dependabot waits the same 3 days. This does not apply to the security tools themselves, which follow their upstream release channels.Toolkit Docker images: Ubuntu and uv are digest-pinned; apt packages resolve from the current signed Ubuntu repositories, so rebuilds are not bit-for-bit reproducible.
Optional tool images: pulled by mutable tags, several of them
latest.--productiondoes not pin or verify them.
--production does not pin Git clones, language-package registries, or build-from-source tools; those track their upstream release channels.
Windows Defender false positives
On a Windows-mounted path (e.g. C:\Users\<you>\..., or any folder visible from Windows while you work in WSL), Microsoft Defender and other AV products may quarantine individual files. IOC tables, sample obfuscated PowerShell, malware-analysis snippets, and exploit strings in .claude/skills/, writeups/, and parts of mcp_server/ contain the same byte patterns real attackers use. Common verdicts include Trojan:Script/Wacatac.B!ml, HackTool:*, and generic Heur.*. These are false positives for a security toolkit.
Exclude the repo folder (recommended on a personal dev box), from an elevated PowerShell:
Add-MpPreference -ExclusionPath "C:\path\to\cybersec-toolkit"Restore files from quarantine via Windows Security → Virus & threat protection → Protection history → "Allow on device". This is per file and does not prevent re-detection.
Keep the repo inside the WSL filesystem (e.g.
~/cybersec-toolkit). Defender does not scan WSL2's virtual disk by default.scripts/sync-wsl.shdoes this for the MCP server.
Files removed by Defender show up as D in git status; restore them with git checkout -- <path> once the exclusion is in place.
MCP server
The server gives MCP-capable clients read access to the 670+ tool registry, install status and recommendations, and governed execution of installed tools. The agent knows every tool, which ones are available, and how to chain them.
Supported clients
Client | Integration | Status |
Claude Code |
| Native configuration included |
Claude Desktop |
| Configuration example documented |
OpenCode |
| Live tested |
Codex |
| Native configuration included |
Gemini CLI |
| Native configuration included |
GitHub Copilot |
| CLI live tested; VS Code documented |
Hermes Agent | User | Live tested |
OpenClaw | User | Live tested |
DeepSeek Harness (dsh) |
| Configuration example documented |
Cursor / Cline / Goose | Client MCP settings + Agent Skills | Compatible through MCP; skills supported |
Continue | Client MCP settings; rules/prompts for context | Compatible through MCP |
LM Studio (>=0.3.17) |
| Compatible through MCP |
Ollama | MCP host in front of it | Compatible through an MCP host |
Open WebUI | MCP-to-OpenAPI bridge | Compatible through an MCP host or bridge |
Per-client setup is in docs/AI_CLIENTS.md; coordinating several agents across any MCP client is in docs/ORCHESTRATION.md.
The server is published to the official MCP Registry as io.github.26zl/cybersec-toolkit and listed on Glama:
What the AI can do
Tool | What it does |
| List/filter all 670+ tools by module, method, or install status (includes URLs) |
| Check if a tool is installed (6 detection strategies) |
| Full details: method, module, URL, install/update/remove commands |
| Deep-dive a module: all tools, install status, which profiles use it |
| See every tool a profile installs, grouped by module |
| Curated tool recommendations for 14 CTF challenge categories |
| Bug bounty tool recommendations for 7 target types with methodology and common vulns |
| Companion-first solve assistant for an authorized target: classifies the target/finding, returns triage gates, recommends skills, picks tools from all modules/profiles, and guides step by step; opt-in |
| Map a CVE id or nickname (e.g. |
| Natural-language → profile/module/tool recommendation |
| All 14 profiles with tool counts and install commands |
| Execute installed tools safely (sanitized args, network policy, rate limiting, audit logging); supports remote execution over SSH |
| Pipe tools together without a shell ( |
| Explicit opt-in Python/Bash execution, with per-script venv selection |
| Add, remove, list, and test SSH remote hosts for remote tool execution |
The agent can query every tool and its install state, chain governed tool calls, parse the output, and pivot on what it finds. Script execution requires a separate opt-in.
External recon → attack surface (needs CYBERSEC_MCP_ALLOW_EXTERNAL=1, authorized scope only)
"Enumerate the attack surface for target.com and flag anything exploitable" — fans out
amass/subfinder→ resolves and probes withhttpx→ fingerprints withwhatweb→ runsnucleitemplates → content discovery withffuf, then ranks hosts by exposure and proposes next steps"Found an open redirect on
/go?url=— weaponize it" — verifies withcurl, then builds an SSRF / OAuth-token-theft PoC and probes for an exploitable callback
Web exploitation
"Confirm and exploit the SQLi on the login endpoint" —
sqlmapto confirm and dump (destructive--os-shell/--os-cmdare policy-blocked), thenrun_scriptto automate the auth bypass and pull just enough for a PoC"GraphQL introspection is on — map it and hunt IDOR" — pulls the schema, generates queries, fuzzes object IDs, and diffs authenticated vs unauthenticated responses
Active Directory / internal
"Low-priv creds on 10.10.0.0/24 — find a path to Domain Admin" — collects with
bloodhound, kerberoasts withimpacket(GetUserSPNs.py), cracks the TGS inhashcat, then validates lateral movement withnetexec, all on a Kali box over SSH (manage_remote_hosts, host mode)"Check for DCSync rights and dump if the path exists" — enumerates replication ACLs, then runs
secretsdump.pyagainst the DC
Binary exploitation & reversing
"Build a ret2libc exploit for this 64-bit binary" — triages with
checksec/readelf, finds gadgets withROPgadget, leaks libc via aputs@pltcall, then writes the fullpwntoolschain invenv="pwntools"and pops a shell locally"Recover the algorithm from this stripped binary" —
objdump/radare2disassembly piped into targeted analysis, then arun_scriptreimplementation to verify behavior
Crypto
"Break this RSA — small
e, several ciphertexts" — detects the attack (Håstad / common modulus / Wiener) and solves it withpycryptodome+sympyin a venv, returning plaintext"This JWT is HS256 with a weak key" — cracks the signing secret and forges an admin token
Blue team · detection engineering
"Write a Sigma rule for this technique and convert it to my SIEM" — authors the rule and renders it for the target backend (Splunk / Elastic) via
sigma-cli"Hunt these Windows event logs for lateral movement" — runs
chainsawover the EVTX with Sigma rules, then summarizes hits by host and timeline"Build YARA rules from these samples and scan the tree" — generates
yarasignatures and runs them recursively
DFIR · malware triage
"Timeline this memory dump" — sweeps
volatility3plugins (pslist,netscan,malfind) and chains them into one narrative"Hunt for C2 beaconing in this pcap" —
tshark/tcpdumpextraction →suricatarules → flags periodic callbacks"Statically triage this suspicious file" —
file→strings→capa/yara, then extracts IOCs for enrichment
Cloud · containers · ops
"Audit this AWS account for public S3 and risky IAM" — runs
prowler/scoutsuiteand surfaces only the high-severity findings"Scan this image and k8s manifests before deploy" —
grypeimage scan pluskubescapeconfig checks"What's my redteam coverage — and fix the gaps" — diffs
get_profile_tools("redteam")against install status and emits the exact install commands
Mobile · wireless · blockchain
"Static-analyze this APK for secrets and insecure storage" —
apktool/jadxdecompile → MobSF-style checks, then greps for keys and endpoints"Audit this Wi-Fi capture" — parses the handshake and runs
aircrack-ng/hashcatagainst it"Review this Solidity contract for reentrancy" — runs
slither/mythriland explains the findings
run_tool and run_pipeline are argument-sanitized, network-policed, rate-limited, and audit-logged, and destructive flags (--os-shell, -rf, --exploit, …) are blocked. run_script is off by default: enabling it runs arbitrary code with the server user's filesystem and network permissions (inside the VM by default, on your host with --local), and the external-target policy does not constrain it. Use only against systems you are authorized to test.
Sandbox and host mode
scripts/mcp-launch.sh starts the server inside a Kata Containers VM unless you pass --local or set CYBERSEC_SANDBOX_MODE=local. The VM is created by a Sandcastle isolated-sandbox provider (sandbox/kata.mjs) that pins the Kata runtime; the same provider can also back a Sandcastle agent-orchestration loop instead of an MCP client. Startup fails closed: a missing runtime, sandbox image, KVM device, Node.js, or launcher dependency stops the server with the reason instead of falling back to the host. ./install.sh --doctor reports sandbox readiness. Setup, tuning, and the threat model are in docs/SANDBOX.md.
The VM has its own tool set. The sandbox image ships the MCP server plus
file,git,binutils,nmap,curl, and Python. Tools installed on the host are not visible inside it, andcheck_installedand the advisors report the guest's tools. Bake a profile into the image instead, and rebuild after pulling server changes, because the image carries its own copy of the server:docker build -f sandbox/Dockerfile --build-arg TOOLKIT_PROFILE=ctf -t cybersec-toolkit-sandbox:latest .Files: the VM sees no host path unless
CYBERSEC_SANDBOX_WORKSPACEnames an absolute directory, which appears as/workspace(read-only withCYBERSEC_SANDBOX_WORKSPACE_RO=1). Pass guest paths torun_tool.Network: the VM uses Docker's default bridge. Set
CYBERSEC_SANDBOX_NETWORK=nonefor offline analysis, or a filtered Docker network to scope egress.Privileges: uid 10001, all capabilities dropped,
no-new-privileges, and 2 GB / 2 vCPUs / 512 processes by default. Raw-socket scans (nmap -sS,-sU, OS detection) needCYBERSEC_SANDBOX_CAP_ADD=NET_RAW.Host mode only: remote execution over SSH (
manage_remote_hosts) and venvs under~/.ctf-venvs/that you created on the host.Audit: records leave the VM on a stderr side channel and are appended to the host log, so the trail outlives the VM.
Client setup
Claude Code reads the tracked .mcp.json:
{
"mcpServers": {
"cybersec-tools": {
"command": "bash",
"args": [
"-lc",
"cd \"$(git rev-parse --show-toplevel)\" && exec bash scripts/mcp-launch.sh"
],
"env": {
"CYBERSEC_MCP_ALLOW_EXTERNAL": "0",
"CYBERSEC_MCP_ALLOW_SCRIPTS": "0"
}
}
}
}The login shell (bash -lc) picks up node and uv from your profile, so profile scripts must not print to stdout, which carries the MCP protocol. Other clients use the same launcher; from any directory:
bash /path/to/cybersec-toolkit/scripts/mcp-launch.sh # Kata VM (default)
bash /path/to/cybersec-toolkit/scripts/mcp-launch.sh --local # host, no VM boundaryCodex: the project
.codex/config.tomlresolves the Git root first, so it works from any subdirectory. If Codex ignores project config, copy the[mcp_servers.cybersec-tools]block into~/.codex/config.toml.Cursor / Continue / Cline / Goose: add the launch command in the client's MCP settings, with an absolute path if the client's working directory is not the repo root.
LM Studio (≥0.3.17): an MCP host itself. Add the server to its
mcp.json(samemcpServersshape as.mcp.json) with an absolute path. MCP over LM Studio's API needs ≥0.4.0 and an MCP-capable endpoint such as/api/v1/chator/v1/responses.Ollama and other local models: a model runtime does not speak MCP. Put an MCP-capable host in front of it (OpenCode, Hermes, OpenClaw, Kit, LM Studio, Cline, Continue, Goose, or Open WebUI through an MCP→OpenAPI bridge such as
mcpo) and point that host at the launcher.
Start with this one server and keep scripts and external targets off unless you have an authorized scope; prefer hosts with human-in-the-loop tool approval. Vendor-neutral repository instructions live in AGENTS.md; Claude Code reads CLAUDE.md, Gemini CLI reads GEMINI.md.
The server speaks stdio, so a Windows client can launch it inside WSL. This runs it directly in WSL, without the Kata VM (the equivalent of --local):
{
"mcpServers": {
"cybersec-tools": {
"command": "wsl",
"args": [
"-d", "kali-linux",
"bash", "-lc",
"cd ~/cybersec-toolkit/mcp_server && uv run fastmcp run server.py --transport stdio --no-banner"
]
}
}
}Use a clone inside the WSL filesystem, or run scripts/sync-wsl.sh from a Windows checkout to copy the server to ~/cybersec-toolkit in WSL (uv cannot create a venv on NTFS). Passing CYBERSEC_MCP_* variables from Windows needs WSLENV; see mcp_server/README.md.
Runs the server in the installer image. This is an ordinary container, not the Kata VM, and its user has passwordless sudo:
{
"mcpServers": {
"cybersec-tools": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "CYBERSEC_MCP_ALLOW_EXTERNAL=0", "-e", "CYBERSEC_MCP_ALLOW_SCRIPTS=0",
"--entrypoint", "bash", "cybersec-toolkit",
"-c",
"cd /opt/cybersec-toolkit/mcp_server && uv run fastmcp run server.py --transport stdio --no-banner"
]
}
}
}Script execution
run_script lets the agent write and run Python or Bash. Enable it with "CYBERSEC_MCP_ALLOW_SCRIPTS": "1" in the client's env block. Scripts get the server process's filesystem and network permissions (the VM by default, your user in --local mode), and CYBERSEC_MCP_ALLOW_EXTERNAL does not constrain them. Review generated code and scope before opting in.
Python libraries without console scripts live in named venvs under ~/.ctf-venvs/ (override with CYBERSEC_MCP_VENVS_DIR), selected per script with the venv parameter. ./install.sh --module crypto creates ~/.ctf-venvs/crypto with pycryptodome, sympy, gmpy2, numpy, z3, fpylll (+ cysignals, which fpylll needs at import time but does not declare), and cypari2, so the agent can call run_script(code, venv="crypto"). In the sandbox, venvs are guest paths: build the image with TOOLKIT_PROFILE=ctf to get crypto. Some packages need an older Python, for example pwntools:
python3.12 -m venv ~/.ctf-venvs/pwntools
~/.ctf-venvs/pwntools/bin/pip install pwntools z3-solverWithout a named venv, scripts see FastMCP and the standard library; cd mcp_server && uv sync --extra ctf-core adds requests, pycryptodome, beautifulsoup4, Pillow, and NumPy.
Reusable helpers the agent writes for you (exploits, multi-step solvers, parsers, protocol helpers) are saved under manual_scripts/. In companion mode the agent proposes a script and runs it only after you approve; in autonomous mode it writes and runs scoped helpers when tools and pipelines stop making progress. Plain recon and HTTP requests stay run_tool calls.
Test the server
cd mcp_server && uv run fastmcp dev server.pyThis opens the MCP Inspector for exercising each tool interactively. mcp_server/README.md covers Claude Desktop setup and the full server reference.
Agent Skills
872 Agent Skills live in .claude/skills/, the canonical tree that Claude Code discovers directly. Skills load on demand for the task at hand instead of occupying context permanently. 31 are project-authored and 841 are curated from open-source projects, each attributed in THIRD_PARTY_NOTICES.md:
10 project developer skills (
add-tool,validate-all,module-scaffold,writeup-template,mcp-sync-check,security-wordlists,security-payloads,guided-assessment,skill-dependency-audit,skill-curation-router)4 cross-skill coordinators (
finding-triage,security-comms,authorization-gate,evidence-hygiene) that route findings, communication, authorization checks, and evidence sanitization7 coverage-gap anchors (GRC/privacy, AI/LLM security, IoT/embedded/hardware, mainframe, telecom/5G, SAP/ERP, supply-chain/product security)
6 CTF methodology skills (
ctf-crypto,ctf-pwn,ctf-web,ctf-rev,ctf-forensics,ctf-stego) and 4 bug bounty methodology skills (bounty-recon,bounty-web,bounty-api,bounty-mobile)754 operational how-tos from mukul975/Anthropic-Cybersecurity-Skills (Apache 2.0)
58 offensive methodology skills from SnailSploit Claude-Red (MIT)
14 code audit skills from Trail of Bits (CC-BY-SA 4.0)
10 bug bounty workflow skills from BugHunter (claude-bug-bounty) (MIT)
4 high-level workflows from Transilience (MIT)
1 coding-agent workflow skill from multica-ai/andrej-karpathy-skills (MIT)
Source and category index: .claude/skills/SKILLS.md. Ranking and curation: .claude/skills/CURATION.md and curation.json, regenerated with python3 scripts/curate_claude_skills.py --write.
As a Claude Code plugin. The repo doubles as a plugin marketplace, so any project can pull in the skill library without cloning:
/plugin marketplace add 26zl/cybersec-toolkit
/plugin install cybersec-toolkit@cybersec-toolkitThe plugin carries the skills only; the MCP server is configured separately (see MCP server).
In other clients. OpenCode, Codex, Gemini CLI, GitHub Copilot, Cursor, Cline, Goose, Hermes, and OpenClaw support Agent Skills through their own paths. scripts/sync-skills.sh mirrors .claude/skills/ into the git-ignored .agents/skills/ for clients that read that location (--check reports drift; make setup runs it). Continue and LM Studio do not document automatic skill discovery; give them selected skill content as rules or context instead. Details: docs/AI_CLIENTS.md.
Helper-script dependencies. Some vendored skills include helper scripts with optional Python imports, declared in .claude/skills/requirements.txt and generated from the import inventory. scripts/validate_claude_skills.py checks skill metadata, index counts, curation freshness, and helper-script syntax.
python3 scripts/audit_skill_dependencies.py --check-declared # verify declarations
python3 -m pip install -r .claude/skills/requirements.txt # optional, ideally in a venvDevelopment
Contributions are welcome: testing installs on different distros, adding missing tools, fixing package mappings, improving MCP workflows, writing example use cases, and reporting rough edges from real CTF, lab, bug bounty, pentest, DFIR, or defensive work. Open an issue for bigger changes or send a focused PR for small fixes; CONTRIBUTING.md has the validation checklist.
git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
make setup # submodules, MCP deps, sandbox deps, skill mirror
make check # lint, validators, bats, pytest, sandbox testsmake help lists every target; the raw commands are in AGENTS.md. The MCP Python project resolves dependencies with uv and exclude-newer = "3 days", so fresh releases are ignored for 72 hours to limit the blast radius of a compromised upload; Dependabot and the weekly uv update workflow use the same cooldown. Run the shell tests on Linux, macOS, or WSL: native Windows checkouts can rewrite the Bats submodules with CRLF and fail with $'\r'.
Star History
License
MIT License. See LICENSE.
The repository also redistributes third-party components under their own terms, including some under CC-BY-SA-4.0 (ShareAlike, not relicensable to MIT). If you redistribute or adapt bundled content, follow those terms; see THIRD_PARTY_NOTICES.md.
Contribution workflow: CONTRIBUTING.md. Community expectations: CODE_OF_CONDUCT.md. Vulnerability reporting: SECURITY.md.
Disclaimer
This project is provided for educational, defensive, and explicitly authorized security testing only. Use it only on systems you own or have written permission to assess, and follow all applicable laws, rules of engagement, third-party tool licenses, and service terms.
The toolkit includes dual-use offensive and defensive tools. Some commands can scan networks, execute exploits, modify systems, or trigger security alerts. MCP/AI integrations are guarded by safety policies, but users remain responsible for reviewing scope, prompts, commands, and outputs before running actions.
This repository does not redistribute the security tools themselves; it installs publicly available, open-source projects from their official upstream sources at install time. It is intended for lawful, authorized use only.
The project is provided "as is", without warranty. Maintainers are not responsible for misuse, damage, data loss, service disruption, or legal consequences from using this toolkit.
Third-party content is bundled under its original license — see
THIRD_PARTY_NOTICES.md.
Available Tools
15 toolscheck_installedARead-onlyIdempotent
Check if a specific cybersecurity tool is installed on the system.
Uses multiple detection strategies: .versions tracking, PATH lookup, pipx binary name fallback, /opt directory check, and docker image check.
When host is provided, checks installation on the remote host via SSH using 'which '.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | Optional remote host name (as configured via manage_remote_hosts). | |
| tool_name | Yes | Name of the tool to check (as listed in tools_config.json). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so safety is covered. The description adds real value beyond that by enumerating the five detection strategies and disclosing that a host argument switches to an SSH-based 'which <binary>' check. Only the absence of failure-mode or return-behavior detail keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, followed by the detection strategies and the remote-host behavior. Three short paragraphs, every sentence carrying information, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be recounted, and annotations cover the safety profile. Between purpose, detection strategies, and remote-host semantics, an agent has everything needed to call this correctly in either local or remote mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented and the baseline is 3. The description still adds meaning by explaining what host actually causes (remote SSH installation check) and anchoring tool_name to tools_config.json, going beyond the schema's phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Check if a specific cybersecurity tool is installed on the system.' This is clearly distinct from run_tool, list_tools, or get_tool_info. It does not explicitly name or route away from those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by explaining that supplying host triggers a remote SSH check, which hints at the local-vs-remote decision. However, it never says when to use this versus list_tools or get_tool_info, and offers no exclusions or prerequisites. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cve_infoARead-onlyIdempotent
Map a CVE to the toolkit's tools, skills, and modules, plus live-lookup commands.
Local-first and deterministic: accepts a CVE id (e.g. "CVE-2021-44228") or a common nickname (e.g. "log4shell", "eternalblue", "zerologon", "printnightmare") and returns the curated exploitation skills, mapped registry tools with install status, and relevant modules.
For live CVSS / CISA KEV / EPSS data it returns ready-to-run run_tool("curl", ...) commands rather than fetching itself — those hit external hosts and are subject to the CYBERSEC_MCP_ALLOW_EXTERNAL policy. Always clear the authorization-gate skill before testing.
| Name | Required | Description | Default |
|---|---|---|---|
| cve | Yes | A CVE id (CVE-YYYY-NNNN) or a known vulnerability nickname. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: local-first/deterministic behavior, that it returns ready-to-run run_tool("curl", ...) commands rather than performing network fetches, that external lookups are gated by CYBERSEC_MCP_ALLOW_EXTERNAL, and that an authorization-gate skill must be cleared first. This explains the openWorldHint=false annotation and the safety profile rather than merely restating it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then groups the input contract, return behavior, and safety prerequisite into distinct, information-dense paragraphs. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be enumerated, and the description still previews the return shape (skills, tools with install status, modules). It is complete for normal use, though it omits edge-case behavior such as an unknown or ambiguous CVE id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by listing concrete nickname examples ('log4shell', 'eternalblue', 'zerologon', 'printnightmare') and the CVE-YYYY-NNNN form, which helps the agent normalize input beyond the schema's brief wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('map a CVE to the toolkit's tools, skills, and modules') and specifies the accepted input forms (CVE id or nickname). It is readily distinguishable from siblings like get_tool_info and get_module_info, which operate on the toolkit's own inventory rather than mapping an external CVE.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when it acts locally versus when it defers to curl commands, and names a hard prerequisite ('always clear the authorization-gate skill before testing'). It gives clear operating context but does not name an alternative tool for cases the caller might prefer instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_module_infoBRead-onlyIdempotent
Get full details about a module: description, all tools, and management commands.
| Name | Required | Description | Default |
|---|---|---|---|
| module | Yes | Module name (e.g. "web", "pwn", "forensics"). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it returns the module description, its tools, and management commands, but says nothing about unknown-module behavior or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the resource first and enumerates the payload; nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one documented parameter, read-only annotations, and an output schema that carries the return structure, the description covers everything needed to invoke the tool correctly. Only the installed/prerequisite condition is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, including examples ('web', 'pwn', 'forensics'). The description adds no meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('module') plus the shape of what comes back (description, tools, management commands). It is distinguishable from sibling get_tool_info by operating on a module rather than a tool, though it never explicitly names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite (e.g. must the module be installed first?), and no routing toward alternatives like get_tool_info or list_tools. The agent must infer the context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_profile_toolsARead-onlyIdempotent
List every tool that a specific profile would install.
Given a profile name, returns the complete list of tools grouped by module, with install status for each. This lets you see exactly what you get before running the install command.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | Profile name (e.g. "ctf", "redteam", "web", "full"). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is fully covered. The description adds that results are grouped by module with per-tool install status, which is useful, but does not disclose anything about auth or rate limits beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and efficiently sized at two short sentences. There is mild redundancy, since 'List every tool that a specific profile would install' and 'returns the complete list of tools' carry nearly the same information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a one-parameter, read-only tool with an output schema present and annotations covering safety, the description supplies everything needed to invoke it correctly. Return-value details are appropriately left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'profile' parameter already documents example values ('ctf', 'redteam', 'web', 'full'). The description only restates 'Given a profile name,' adding no format or validation detail beyond the schema — the baseline 3 for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'List every tool that a specific profile would install' — and adds scope detail, returning tools grouped by module. This clearly separates it from siblings like list_tools (all tools) and get_tool_info (single tool).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a concrete use context: 'This lets you see exactly what you get before running the install command,' which implies previewing before install. It stops short of naming explicit alternatives (e.g., list_profiles for discovering valid names), so it is clear but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tool_infoBRead-onlyIdempotent
Get detailed information about a cybersecurity tool.
Returns the tool's install method, module, URL, installation status, module description, and management commands (install, update, remove).
| Name | Required | Description | Default |
|---|---|---|---|
| tool_name | Yes | Name of the tool to look up. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety profile is covered. The description adds that the result includes install method, module, URL, install status and management commands, which is useful behavioral context, but it says nothing about what happens for an unknown tool_name or any rate limits. Given annotations carry the main burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded, no filler. The second sentence enumerating return fields is slightly redundant given an output schema exists, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since an output schema exists, the description need not explain return values, so its core obligation is purpose plus routing. Purpose is covered; what's missing is any signal about when to prefer this over list_tools, check_installed, or get_module_info, and what happens on a bad tool name.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented parameter, so the baseline 3 applies. The description adds no format, casing, or matching-rule detail (e.g., exact name vs alias) beyond the schema's 'Name of the tool to look up.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get detailed information about a cybersecurity tool') and even enumerates the fields returned, so an agent knows exactly what comes back. It does not, however, differentiate itself from nearby siblings like get_module_info or check_installed, which also return tool/module metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use statement. The description never mentions list_tools (browse the catalog) or check_installed (installation state), even though this tool overlaps heavily with both, leaving the agent to infer routing from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
guided_assessmentADestructive
Plan, guide, or autonomously solve a security task over the MCP toolchain.
An orchestrator on top of the registry, advisors, install checks, audit logging, and execution policy. Bootstrap commands use the governed execute_tool() path, so target scope, external-network, shell-injection, and blocked-flag checks apply.
By DEFAULT it auto-detects the right workflow + tools for the problem (workflow/ target_type="auto") and acts as a companion: it returns classification, triage gates, recommended skills, reporting next steps, a plan, tool install status, next actions, and the full MCP toolchain surface WITHOUT auto-running commands in this initial call. The agent can then run tools step by step as the user approves. The heaviest mode (autonomous) starts the auto-solver contract: it bootstraps triage, then the client agent continues with the full MCP toolchain (registry/advisors/install checks/run_tool/run_pipeline and separately gated run_script). When registry tools and pipelines are not enough, autonomous mode may create, save, and run scoped helper scripts for the user, persisting reusable ones under manual_scripts/. Simple recon/HTTP commands such as curl remain run_tool calls.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | "companion" (default) or "autonomous" (opt-in). | companion |
| target | Yes | URL, hostname/IP, or local file path to assess. | |
| finding | No | Optional short finding summary to classify for triage/report routing. Raw finding text is used locally but not echoed in the result. | |
| workflow | No | "auto" (default — inferred), "bounty", "ctf", or "generic". | auto |
| intensity | No | "low" (default) or "medium". Medium may include low-volume nmap. | low |
| max_steps | No | Maximum number of bootstrap steps autonomous mode auto-executes. | |
| target_type | No | "auto" (default — inferred from the target) or an explicit type: bounty type (web_app/api/cloud/network/iot/mobile_app) or CTF category. | auto |
| authorization_confirmed | No | Required before any network step executes. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (destructiveHint=true, openWorldHint=true), the description discloses governed execution checks ('target scope, external-network, shell-injection, and blocked-flag checks'), confirms no auto-run in the default call, and explains that autonomous mode may create, save, and run scoped helper scripts, persisting reusable ones under manual_scripts/. It also states that authorization_confirmed is required before any network step executes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and is dense with useful operational detail. It is a bit paragraph-heavy and could be more scannable, but given the complexity of the orchestration behavior, most sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be fully explained, yet the description still summarizes what companion mode returns and what autonomous mode bootstraps. Combined with the annotations and 100% schema coverage, the definition gives an agent enough context to choose and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter semantics already live in the input schema. The description largely repeats default values and enum-like options (mode companion/autonomous, workflow auto, target_type auto) without adding substantial syntax or format guidance beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific compound verb+resource: 'Plan, guide, or autonomously solve a security task over the MCP toolchain.' It immediately positions the tool as an 'orchestrator on top of the registry, advisors, install checks, audit logging, and execution policy,' which cleanly distinguishes it from sibling tools like run_tool, run_pipeline, and suggest_for_*.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the default companion mode versus opt-in autonomous mode, and explicitly routes simple recon/HTTP commands to run_tool instead. It does not, however, give a full when-not-to-use rule or spell out the exact conditions under which autonomous mode should be preferred over companion mode, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_profilesARead-onlyIdempotent
List all 14 available installation profiles with details.
Each profile is a curated set of modules targeting a specific use case. Shows module count, tool count, and install command for each profile. Profiles range from 'osint' (2 modules) to 'full' (18 modules, 670+ tools).
Returns: All profiles with descriptions, module lists, tool counts, and install commands.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and closed-world behavior, so the safety profile is covered. The description adds genuine context beyond that: exactly 14 profiles, the osint-to-full size range, and that each is a curated module set. The 'Returns' sentence overlaps the output schema and is somewhat redundant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and size, and the supporting sentences are short. The trailing 'Returns:' block restates what the earlier sentence ('Shows module count, tool count, and install command') already established, which is mild duplication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not strictly required, and the description supplies scope, size, and content shape. It is essentially complete for a zero-param read tool, with the only gap being guidance on when to prefer it over adjacent discovery tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. Schema coverage is 100% and the description correctly implies no filtering options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all 14 available installation profiles') with explicit scope and quantitative detail. Clear enough to distinguish from siblings like get_profile_tools or recommend_install, which deal with a single profile or a recommendation rather than the full catalog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the discovery framing but never stated: it does not say when to call this versus get_profile_tools, recommend_install, or check_installed. No exclusions or prerequisites are given, though the context is reasonably inferable for a list tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_toolsARead-onlyIdempotent
List and filter the 670+ cybersecurity tools in the registry.
Returns the tools drawn from tools_config.json — the same registry the installer and the advisors share — with the total count, the filters still available to narrow the results, and one entry per tool. Combine the filters to scope the list: module="web" for web tools, method="pipx" for Python-packaged tools, installed_only=True for only what is on this host. Start here to discover what exists before check_installed or get_tool_info.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | Filter by install method. One of apt, pipx, go, cargo, gem, git, binary, docker, snap, special, source, npm. | |
| module | No | Filter by module (e.g. "web", "pwn", "forensics"). 18 modules available. | |
| installed_only | No | If true, only return tools that are currently installed. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, closed-world behavior, so the safety profile is covered. The description adds valuable context about data provenance (drawn from tools_config.json, shared with installer/advisors) and return shape (total count, available filters, one entry per tool), which goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then provides provenance, filter examples, and entry-point guidance. Efficient and well-ordered, though the provenance sentence is slightly verbose for a list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety, the description fills remaining gaps: data source, filter combinations, and the tool's role in the workflow relative to siblings. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so baseline is 3, but the description adds semantic value by giving concrete examples of the three filters and showing combinations, which helps an agent reason about filter interaction beyond the schema's individual parameter docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list/filter) and resource (670+ cybersecurity tools in the registry), and grounds it in the shared tools_config.json registry. Distinguishes from siblings by naming check_installed and get_tool_info as downstream steps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Start here to discover what exists before check_installed or get_tool_info,' giving a clear entry-point role, and offers concrete filter combinations (module="web", method="pipx", installed_only=True) to guide use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_remote_hostsADestructive
Add, list, test, or remove the SSH hosts that run_tool can target remotely.
Manages the remote-host registry that lets run_tool (and check_installed) run a tool on a remote Kali/Linux box over SSH instead of locally, so the tool only has to be installed on the remote. The action selects the operation: "list" shows every configured host; "add" registers or updates a host (needs name and hostname, plus optional user, port, ssh_key, and a tool_allowlist that restricts which tools may run there); "remove" deletes a host by name; "test" opens an SSH connection to confirm the host is reachable.
Connections use StrictHostKeyChecking=accept-new, so the key presented on the first connection is pinned in ~/.ssh/known_hosts and any later change is rejected. Verify that first fingerprint out-of-band for a host you do not control, or add the key to known_hosts before "test".
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Host name (required for add/remove/test). | |
| port | No | SSH port (default 22). | |
| user | No | SSH username (default "kali"). | kali |
| action | Yes | Operation to perform: "list", "add", "remove", or "test". | |
| ssh_key | No | Path to SSH private key (e.g. "~/.ssh/id_kali"). | |
| hostname | No | IP address or hostname of the remote machine (required for add). | |
| description | No | Human-readable description of the host. | |
| tool_allowlist | No | Comma-separated list of allowed tool names (e.g. "nmap,gobuster,sqlmap"). None means all tools allowed. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive/openWorld/non-idempotent, and the description adds genuinely new behavioral context: the StrictHostKeyChecking=accept-new policy, that the first-seen key is pinned in ~/.ssh/known_hosts and later changes are rejected, and the out-of-band fingerprint verification advice. That is security-relevant disclosure beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then enumerates actions, then ends with the SSH key-pinning caveat. Every paragraph adds information, though the per-action enumeration is dense and slightly overlaps the schema's own descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action mutation tool with side effects, an output schema present (so returns need no explanation), and full annotation coverage, the description supplies the missing pieces: action semantics, credential requirements, and the connection-security caveat. An agent has enough to select actions and call them correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds real meaning: which parameters each action requires (name for add/remove/test, hostname for add) and what tool_allowlist actually does (restricts which tools may run on that host, with None meaning all). It stops short of restating defaults like user='kali' or port=22.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb set (add/list/test/remove) and a concrete resource (SSH remote hosts), and names the dependent sibling tools run_tool and check_installed. An agent can immediately distinguish this registry-management tool from the execution tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains that the action parameter selects the operation and gives per-action conditions (add needs name+hostname, remove takes a name, test opens a connection), plus the prerequisite context that hosts enable remote execution via run_tool. It does not, however, state explicit when-not-to-use cases or an ordering rule such as 'add before test before run_tool'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_installARead-onlyIdempotent
Recommend which profile, modules, or individual tools to install.
Analyzes a natural-language description of what the user wants to do and recommends the best installation approach — from a full profile down to just a few individual tools. Avoids installing everything when only a subset is needed.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Natural-language description of what the user wants to do. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds useful context by clarifying this is an advisory recommendation engine rather than an installer, despite the 'install' in the name, and explains the over-install avoidance behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first sentence, followed by scope and rationale. All sentences are relevant, though the second sentence could be slightly trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a present output schema, full parameter documentation, and complete annotations, the description covers everything an agent needs to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'task' parameter, so the schema already explains it. The description's reference to a 'natural-language description of what the user wants to do' merely restates the schema, adding no syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (recommend) and resource (profile, modules, or individual tools to install), and scopes the output range from a full profile down to a few tools. It does not explicitly name or contrast with siblings like list_profiles or get_profile_tools, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage ('Analyzes a natural-language description of what the user wants to do') and gives a rationale ('Avoids installing everything when only a subset is needed'), but provides no explicit when-to-use guidance or alternatives among the many sibling tools (list_profiles, get_profile_tools, check_installed).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_pipelineADestructive
Execute a pipeline of tools, piping stdout from each step into stdin of the next.
Replaces shell piping (e.g. strings binary | grep flag) with a safe,
no-shell alternative. Each step is validated individually (allowlist,
argument sanitization, policy checks) before any process starts.
Each step's stdout and stderr are bounded to 200KB as they are read, and an
intermediate step's bounded stdout is what gets piped into the next step. If
any step hits that cap, the returned truncated flag is set and the final
stdout carries a truncation marker.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | Reserved for future use. Currently only local execution is supported. | |
| steps | Yes | List of dicts, each with 'tool' (required) and 'args' (optional) keys. Max 10 steps per pipeline. | |
| timeout | No | Global timeout for entire pipeline in seconds (default 120, max 300). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, but the description still adds real value: per-step validation before any process starts, a 200KB per-step stdout/stderr cap, a truncation marker, and the returned `truncated` flag. It omits what happens to earlier steps if a later step fails, which matters for a destructive pipeline, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by the motivating example, then the safety/bounding guarantees. Every sentence carries concrete information (validation order, 200KB cap, truncation marker) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-described, and the description covers the semantics an agent needs: step structure, validation, bounding, and truncation signaling. The main residual gap is failure/partial-execution behavior across steps in a destructive multi-step run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema itself documents `host` (reserved), the step-list shape with the max-10 constraint, and the timeout default/max. The description adds no further parameter-level syntax or semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Execute a pipeline of tools') plus the precise mechanism (stdout piped to stdin of the next step), which implicitly separates it from the single-invocation siblings run_tool and run_script. It does not explicitly name those alternatives, so an agent still has to infer the routing, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for it: replacing shell piping such as `strings binary | grep flag` with a safe, no-shell alternative. That is a concrete usage trigger, though it stops short of explicitly contrasting with run_tool/run_script or stating when not to use a pipeline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_scriptADestructive
Write and execute a Python or Bash script, returning its output.
Writes the code to a temporary file, executes it via python3/bash, and returns stdout/stderr. The temp file is deleted after execution. Requires CYBERSEC_MCP_ALLOW_SCRIPTS=1. This is an explicit full-code execution opt-in: scripts are not OS-sandboxed and are not constrained by CYBERSEC_MCP_ALLOW_EXTERNAL.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The script source code to execute. | |
| venv | No | Optional Python venv name from ~/.ctf-venvs/ (e.g. "pwntools"). Allows using a different Python with specific packages installed. Ignored for language="bash". If not set, uses the MCP server's Python. | |
| timeout | No | Maximum execution time in seconds (default 120, max 300). | |
| language | No | "python" (default) or "bash". | python |
| working_dir | No | Working directory for the script (default: system temp dir). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive/openWorld/non-idempotent, and the description meaningfully extends them: it discloses the temp-file write-then-delete lifecycle, the interpreter used, and two independent gating/containment flags. The non-sandboxing warning is exactly the kind of consequence an agent needs before calling a mutating execution tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with what the tool does and the effect, followed by execution mechanics and then the opt-in/containment caveats. Every sentence contributes information an agent would not otherwise have.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value explanation is unnecessary, and the description covers the remaining risk surface: execution mechanism, cleanup, and the two environment flags governing permission and containment. Nothing material is missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the venv, timeout, language, and working_dir semantics are already fully documented in the schema, including defaults and the venv/bash interaction. The description adds no parameter-level detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb+resource pair ('Write and execute a Python or Bash script') plus the scope of the effect ('returning its output'), which is enough to separate it from sibling run_tool and run_pipeline. Nothing about the action is left ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the precondition for use (CYBERSEC_MCP_ALLOW_SCRIPTS=1) and the security posture (not OS-sandboxed, not bound by CYBERSEC_MCP_ALLOW_EXTERNAL). It stops short of naming when to reach for run_tool/run_pipeline instead of scripting an ad-hoc program, so it lacks an explicit alternative-routing clause.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_toolADestructive
Execute an installed cybersecurity tool or system utility and return its output.
Runs tools from the 670+ registry as well as ~120 standard system utilities (strings, file, curl, grep, base64, xxd, jq, etc.) that are allowed without being in the registry. Arguments are sanitized to prevent shell injection. Timeout is clamped to 1-300s. Output is truncated at 200KB.
Network tools (including curl, wget, ping, etc.) are restricted to local/private targets by default. Set CYBERSEC_MCP_ALLOW_EXTERNAL=1 to allow external targets.
When host is provided, the tool is executed on the remote host via SSH. The tool does not need to be installed locally — only on the remote host.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Command-line arguments as a string (e.g. "--version" or "-sV 10.0.0.1"). | |
| host | No | Optional remote host name (as configured via manage_remote_hosts). | |
| timeout | No | Maximum execution time in seconds (default 120, max 300). | |
| tool_name | Yes | Name of the tool to run (registry tool or system utility). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructive/openWorld/non-idempotent, and the description adds substantial context on top: argument sanitization against shell injection, timeout clamping to 1-300s, 200KB output truncation, default local/private network restriction with an env-var escape hatch, and SSH semantics where the tool need only exist on the remote host.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layered constraints in short, well-separated paragraphs. Every sentence carries operational information (sanitization, timeout clamp, truncation, network policy, SSH) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a high-risk execution tool: safety profile is covered by annotations, argument and host semantics by the description, and return values by the existing output schema. An agent has everything needed to call it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning for 'host' (executed remotely via SSH, tool need not be installed locally) and for 'args' (sanitized to prevent shell injection). Timeout and tool_name semantics are largely left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Execute') and resource ('an installed cybersecurity tool or system utility') plus the outcome ('return its output'). It clearly separates itself from siblings like run_script and run_pipeline by scoping to registry tools and a fixed set of allowed system utilities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives rich context: registry vs. ~120 built-in system utilities, network tools restricted to local/private targets unless CYBERSEC_MCP_ALLOW_EXTERNAL=1, and remote execution when host is supplied. It never explicitly says when to prefer run_script/run_pipeline or the suggest/recommend siblings, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_for_bountyARead-onlyIdempotent
Suggest cybersecurity tools for a bug bounty target type.
Provides curated tool recommendations with installation status, methodology steps (starting with scope verification), common vulnerabilities, and quick wins for 7 target types: web_app, api, mobile_app, cloud, network, iot, llm.
Also accepts aliases: web/webapp (web_app), rest/graphql (api), android/ios/mobile (mobile_app), aws/azure/gcp/k8s (cloud), infra/infrastructure (network), firmware/embedded (iot).
| Name | Required | Description | Default |
|---|---|---|---|
| target_type | Yes | Type of bug bounty target (e.g. "web_app", "api", "cloud"). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so safety is covered. The description adds real behavioral value by disclosing what the response contains: curated tools with installation status, methodology steps starting with scope verification, common vulnerabilities, and quick wins.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose in the first sentence, then supporting detail. The alias list is somewhat long but each entry maps useful input strings, so it earns its place; no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, yet the description still sketches the response shape. Combined with the enum-free parameter being fully enumerated, the definition is complete enough for an agent to invoke correctly, missing only explicit sibling routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% but the schema has no enum for target_type; the description supplies the seven valid values (web_app, api, mobile_app, cloud, network, iot, llm) and the accepted aliases, which the schema alone does not convey. This meaningfully extends parameter semantics beyond the structured field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Suggest cybersecurity tools') plus the scope ('for a bug bounty target type'), immediately distinguishing it from the sibling suggest_for_ctf. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context of use is implied by 'bug bounty target type' and the enumerated target list, but there is no explicit when-to-use / when-not guidance and no routing to alternatives such as suggest_for_ctf or guided_assessment. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_for_ctfARead-onlyIdempotent
Suggest cybersecurity tools for a CTF challenge category.
Provides curated tool recommendations with installation status for 14 challenge types: web, crypto, pwn, reversing, forensics, stego, misc, networking, wireless, osint, cloud, mobile, blockchain, llm.
Also accepts aliases: re/rev (reversing), binary/exploitation (pwn), steganography (stego), network (networking), recon (osint), etc.
| Name | Required | Description | Default |
|---|---|---|---|
| challenge_type | Yes | Type of CTF challenge (e.g. "web", "crypto", "pwn"). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so safety is covered. The description adds useful operational context beyond them: that recommendations are curated and include installation status, and that 14 named categories are supported. It does not discuss result size or ordering, but the output schema presumably covers returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the category list and alias list in separate, scannable sentences. Slightly list-heavy, but each element earns its place by supporting valid parameter values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, output-schema-backed read tool, the description covers what it does, the full valid value space, and the fact that install status is included. Nothing critical is missing; only cross-tool routing is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter, so the baseline is 3, but the description goes further by enumerating all 14 accepted challenge_type values and a set of aliases (re/rev, binary/exploitation, steganography, network, recon). Since the schema declares no enums, this enumeration is genuinely additive and prevents wrong-value calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Suggest cybersecurity tools for a CTF challenge category'), which the agent can immediately act on. It does not name or contrast with the obvious sibling suggest_for_bounty, so differentiation from that near-twin is left implicit via the word 'CTF'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a CTF challenge category' implies the use context, but there is no explicit when-to-use / when-not guidance and no reference to the adjacent suggest_for_bounty or recommend_install tools. The agent must infer that this is the CTF-scoped path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.2.1- Changed
list_tools2 fields changed- changed
Input schema / properties / method / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "enum": [ + "apt", + "pipx", + "go", + "cargo", + "gem", + "git", + "binary", + "docker", + "snap", + "special", + "source", + "npm" + ], + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / method / descriptionPrevious value: -"Filter by install method (apt, pipx, go, cargo, gem, git, binary, docker, snap, special, source, npm)."New value: +"Filter by install method. One of apt, pipx, go, cargo, gem, git,\nbinary, docker, snap, special, source, npm."
- Changed
manage_remote_hosts2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"One of \"list\", \"add\", \"remove\", \"test\"."New value: +"Operation to perform: \"list\", \"add\", \"remove\", or \"test\"." - added
Input schema / properties / action / enumAdded value: +[ + "list", + "add", + "remove", + "test" +]
15 tool updates
v1.2.0- First observed
check_installed - First observed
get_cve_info - First observed
get_module_info - First observed
get_profile_tools - First observed
get_tool_info - First observed
guided_assessment - First observed
list_profiles - First observed
list_tools - First observed
manage_remote_hosts - First observed
recommend_install - First observed
run_pipeline - First observed
run_script - First observed
run_tool - First observed
suggest_for_bounty - First observed
suggest_for_ctf
TDQS
Scored across 15 tools
Most tools have distinct roles (registry discovery, install planning, advisory, execution, remote management), but there is some overlap between advisory/orchestration tools like guided_assessment, recommend_install, and suggest_for_* that could cause hesitation. Descriptions help clarify, so boundaries are mostly clear.
Names are consistently snake_case and mostly follow a verb_noun pattern (get_tool_info, list_tools, run_tool). Minor deviations like guided_assessment (adjective_noun) and suggest_for_ctf (verb_preposition_noun) are readable but slightly break the pattern.
15 tools is well within the ideal range and each tool serves a distinct purpose across discovery, installation, execution, and orchestration. No tool feels redundant or unnecessary for the toolkit's scope.
The surface covers registry discovery, profile/module insight, install recommendation, tool execution (single, pipeline, script), remote host management, and high-level orchestration. Minor gaps exist: no direct install/uninstall/update tool (though commands are surfaced via get_tool_info) and no explicit audit-log viewer.
Maintenance
Related MCP Connectors
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
Search, inspect and invoke every public tool on Invokera through one MCP connection.
Scans remote MCP servers for protocol, security, and TLS issues; exposes scan tools via MCP.
Agent-native catalogue of Baseframe Labs dev tools and MCP servers.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables aggregation, filtering, transformation, and composition of tools from multiple MCP servers through a single proxy with tool views.5AGPL 3.0
- AlicenseNot gradedqualityAmaintenanceMCP server wrapping the toolgovern CLI as a single generic run tool for agent-tool policy validation.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceProvides a security and context-control layer that multiplexes multiple MCP servers behind a single endpoint, scanning tool definitions and results, enforcing authorization, rate limiting, and audit logging, and dynamically retrieving tools to manage context window usage.MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to run server-side security audits through callable tools that sweep attack surfaces, trace source-to-sink reachability, verify live findings, and manage scope, recon, chains, manuals, intelligence, and memory from any MCP client.3,832 npm512MIT