Skip to main content
Glama

CI Integration Security OpenSSF Scorecard Last commit License: MIT Docker image Glama score

    /\   /\        ______      __              _____
   (o ) ( o)      / ____/_  __/ /_  ___  _____/ ___/___  _____
    \ \_/ /      / /   / / / / __ \/ _ \/ ___/\__ \/ _ \/ ___/
  <==\   /==>   / /___/ /_/ / /_/ /  __/ /   ___/ /  __/ /__
     \ V /      \____/\__, /_.___/\___/_/   /____/\___/\___/
     /_ _\           /____/                          by 26zl
      |_|                     Toolkit

A security toolkit that AI agents can drive, under rules you set. One command installs 670+ security tools on Linux or Termux. An MCP (Model Context Protocol) server lets Claude Code, Codex, Gemini CLI, OpenCode, and other MCP clients discover, recommend, and run those tools through a governed execution path, inside a disposable Kata Containers VM by default. 872 Agent Skills supply the methodology for CTF, pentest, bug bounty, DFIR, and blue-team work.

Component

What you get

Installer

670+ tools in 18 modules, 14 profiles, and 12 install methods for Debian/Ubuntu/Kali/Parrot, Fedora/RHEL, Arch, openSUSE, and Termux

MCP server

15 tools for discovery, advice, and governed execution; external targets and script execution are off by default

Sandbox

One Kata Containers VM per session by default; --local runs the server on the host

Agent Skills

872 skills (31 project-authored, 841 curated), also installable as a Claude Code plugin

What makes it different: most toolkits stop at installing tools. Here an AI can also drive them: infer the problem type, pick tools from every module and profile, and work through the problem with you step by step. When you explicitly authorize it, the same toolchain runs an autonomous solver loop. Companion by default; autonomous only when you ask. It complements Kali, Parrot, or BlackArch rather than replacing them: it runs on the machine you already have, Termux included.

Quick start · How it works · Trust & safety · Installer · MCP server · Agent Skills · Development · License

Quick start

1. Install the tools on a supported Linux distro or on Termux:

git clone --depth 1 --branch v1.3.0 https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
./install.sh --doctor             # read-only preflight: distro, prerequisites, MCP server, sandbox
sudo ./install.sh --profile ctf   # one profile; with no flags, all 18 modules

2. Connect an AI client. The tracked configs for Claude Code (.mcp.json), Codex, Gemini CLI, and OpenCode start the server through scripts/mcp-launch.sh, with external targets and script execution disabled. Choose where the tools run:

Mode

Tools run

Host needs

One-time setup

Sandbox (default)

In a disposable Kata VM, from the sandbox image

Linux with KVM, Docker 23+ with a Kata runtime, Node.js 22+

make sandbox-image, then npm --prefix sandbox ci --ignore-scripts

Host (--local)

As your user, with the tools install.sh installed

uv

Add --local to the launcher, or set CYBERSEC_SANDBOX_MODE=local in the client's env

Kata needs KVM: on macOS, on Windows, and in VMs without nested virtualization, use host mode.

3. Restart the client. The 15 tools appear (/mcp in Claude Code). Then ask for what you need, for example "triage this binary" or "map the attack surface of my lab at 10.10.0.0/24".

Related MCP server: toolgovern

How it works

Two entry points share one tool registry. An operator runs the bash installer to put tools on disk: on the host, or into the sandbox image at build time. An AI agent talks to the MCP server, which runs inside a Kata VM by default, to discover, recommend, and execute those tools through one governed path. tools_config.json is the single source of truth that the modules define and the MCP advisors read; CI validators keep the Python and bash sides in sync.

How it works: an operator runs the bash installer to put tools on disk; an AI agent drives the MCP server, which runs in a Kata VM by default, to discover, recommend, and execute them. The installer and MCP server meet at the tools_config.json registry and the installed tools, with security.py governing tool execution and CI validators keeping the Python and bash sides in sync.

Solid arrows are runtime or installation actions; dashed arrows are validation and context relationships. Client configurations enter through the root-aware launcher, which boots the VM through a Sandcastle Kata provider (sandbox/kata.mjs) — or, with --local, starts the server on the host. security.py governs run_tool and run_pipeline through the allowlist, argument checks, and network policy without invoking a shell. run_script is a separate, disabled-by-default capability that runs arbitrary code. Agent Skills stay outside the execution path: .claude/skills/ is canonical, and scripts/sync-skills.sh generates .agents/skills/ for clients that read the portable mirror. Mermaid source: assets/how-it-works.mmd.

Trust & safety

What runs, and what is gated:

  • Default-safe MCP. CYBERSEC_MCP_ALLOW_EXTERNAL=0 rejects network targets that do not resolve to private or loopback ranges, and CYBERSEC_MCP_ALLOW_SCRIPTS=0 disables run_script. External scopes and scripting are explicit opt-ins.

  • One gate for governed execution (mcp_server/security.py): registry allowlist, no shell (create_subprocess_exec, never shell=True), argument sanitization, a per-tool blocked-flag denylist (e.g. sqlmap --os-shell, nmap -iL, file-list and target-injection flags), target and network policy, rate limiting, output caps, and timeouts. The policy knows enough CLI grammar to tell a target from a header, wordlist, output path, or target-list flag. Tool output reaches the model without terminal escape sequences or LLM control markers, and lines addressed to an AI reader are flagged.

  • Tools run in a disposable VM by default. The launcher boots a Kata Containers VM with its own kernel, no host filesystem beyond an optional CYBERSEC_SANDBOX_WORKSPACE mount, no Docker socket, a non-root user with every capability dropped, and memory, CPU, and process limits. It is destroyed when the client disconnects. Startup fails closed instead of falling back to the host; --local is the explicit opt-out.

  • Know the limits. The VM is not a network boundary: it reaches whatever its Docker network reaches, and CYBERSEC_MCP_ALLOW_EXTERNAL is a preflight check on resolved addresses, not a firewall (set CYBERSEC_SANDBOX_NETWORK=none or use a filtered network). Inside the VM, and on your host in --local mode, allowed tools run with the server user's permissions, and some of them spawn child processes or load plugins.

  • Audit trail without leaks. Actions are logged as JSON lines to an owner-only (0600), rotating log under the user's state directory (~/.local/state/cybersec-tools-mcp/audit.log by default). Script bodies are never stored, only their SHA256 and length, and credential-shaped strings are redacted from tool arguments. A sandboxed server mirrors its records to the same host log, and records are hash-chained, so an edited or deleted record inside a chain shows up in make audit-verify (limits).

  • Least privilege in the installer. It runs as root but drops to the invoking user ($SUDO_USER) for cloned-repo builds and pip/cargo/gem installs; binary releases are SHA256-verified when checksums are published.

  • Dual-use tooling is gated. C2 and phishing frameworks (Sliver, Caldera, gophish, evilginx, …) are off by default and install only with --include-c2 (the redteam and full profiles set it); the MCP layer reflects this and never auto-runs them.

  • Authorized use only. See SECURITY.md, the supply chain model, and the disclaimer.

Installer

Requirements

A supported Linux distro (Debian/Ubuntu/Kali/Parrot, Fedora/RHEL, Arch, openSUSE) or Termux. Runtimes (Python, Go, Ruby, Java, Rust, Node.js), dev libraries, pipx, and build tools are installed automatically. The installer does not run on macOS or native Windows; use WSL or Docker.

  • Docker is needed only for --enable-docker (C2 frameworks, MobSF, BeEF, BloodHound, TheHive, Cortex). Install it yourself: Docker Engine docs.

  • GitHub authentication is recommended. The installer downloads ~50 release binaries and makes ~50+ GitHub API calls; unauthenticated requests are limited to 60 per hour, authenticated ones to 5,000. It uses your gh auth login session automatically, also under sudo. Alternatively, pass a personal access token (no scopes needed) through sudo, which otherwise drops it: sudo --preserve-env=GITHUB_TOKEN ./install.sh.

Install

From the latest release (pinned; recommended):

git clone --depth 1 --branch v1.3.0 https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit && sudo ./install.sh

From main (newest tools and fixes, including unreleased work):

git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit && sudo ./install.sh

With no flags, the installer covers the standard tools of all 18 modules. C2/phishing tools and Docker images stay opt-in; --profile full --enable-docker installs everything the current platform supports. To install a subset:

sudo ./install.sh --profile ctf                      # CTF tools only
sudo ./install.sh --profile redteam --enable-docker  # Red team + Docker C2
sudo ./install.sh --module web --module recon        # Specific modules
sudo ./install.sh --tool sqlmap --tool nmap          # Individual tools
sudo ./install.sh --dry-run --profile ctf            # Preview without installing
./install.sh --doctor                                # Read-only preflight; no root needed
sudo ./install.sh --help                # Full help
sudo ./install.sh --list-profiles       # Show profiles
sudo ./install.sh --list-modules        # Show modules
sudo ./install.sh --skip-heavy          # Skip large/slow packages
sudo ./install.sh --skip-pipx           # Skip all pipx (Python) installs
sudo ./install.sh --skip-go             # Skip all Go tool installs
sudo ./install.sh --skip-cargo          # Skip all Cargo (Rust) installs
sudo ./install.sh --skip-gems           # Skip all Ruby gem installs
sudo ./install.sh --skip-git            # Skip all git clone installs
sudo ./install.sh --skip-binary         # Skip all binary release downloads
sudo ./install.sh --skip-source         # Skip build-from-source, snap, npm, and curl-pipe installs
sudo ./install.sh --fast                # Skip checksum verification (see Supply chain model)
sudo ./install.sh --require-checksums   # Fail if a binary release has no checksum file
sudo ./install.sh --production          # Strict checksum preset for release downloads
sudo ./install.sh --upgrade-system      # Upgrade system packages before installing
sudo ./install.sh --list-sessions       # List install sessions and exit
sudo ./install.sh --rollback <id|last>  # Roll back tools installed in a session
sudo ./install.sh --version             # Show installer version and exit
sudo ./install.sh --enable-docker       # Pull Docker images
sudo ./install.sh --include-c2          # Include C2/phishing frameworks (Empire also needs --enable-docker)
sudo ./install.sh -j 8                  # 8 parallel install jobs (default: 4)
sudo ./install.sh -v                    # Verbose / debug output

--tool installs only the named tool, without the full dependency setup. Dry-run time estimates count install entries across methods, so they can exceed the de-duplicated 670+ tool registry.

The time goes to I/O-bound work that no scripting language can speed up:

What takes time

Why

System packages (apt/dnf)

Downloading and unpacking .deb/.rpm packages and resolving dependencies

Go tools

Downloading modules and compiling each binary

pipx (Python)

Creating one isolated venv per tool and downloading wheels

Cargo (Rust) crates

Compiling from source

Git clones

Cloning each repository

Binary releases

Downloading pre-built binaries from GitHub

./install.sh --dry-run prints the live per-method breakdown. pipx, Go, git, and binary installs run in parallel (-j 4 by default); the system package manager and Cargo run sequentially. To go faster: install a profile or --module instead of everything, --skip-cargo to avoid Rust compilation, -j 8 for more parallel jobs, and an apt-cacher-ng proxy for repeated installs.

Docker and Podman

The prebuilt image is the installer on Ubuntu, not a sandbox. Its default command is a dry run of the full profile:

docker run --rm ghcr.io/26zl/cybersec-toolkit                                  # preview only
docker run -it --name cybersec --entrypoint bash ghcr.io/26zl/cybersec-toolkit  # keep a container
sudo ./install.sh --profile ctf                                                # inside it; reattach with: docker start -ai cybersec

Build it yourself with docker build -t cybersec-toolkit ., or run a throwaway install through the bundled Compose file with docker compose run --rm installer --profile ctf. On Apple silicon, add --platform linux/amd64 to docker build and docker run. Podman works as a drop-in replacement (podman compose needs a compose provider). The image grants the toolkit user passwordless sudo so the installer can manage packages; treat code inside it as root-capable.

Termux (Android)

pkg install git
git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
./install.sh --profile lightweight

Profiles

Profile

Modules

Description

full

All 18

Complete security toolkit

ctf

misc, crypto, pwn, reversing, stego, forensics, cracking, web, mobile, blockchain

CTF competitions

redteam

misc, networking, recon, web, enterprise, pwn, mobile, cracking, cloud, wireless, reversing, crypto

Offensive security

web

misc, networking, recon, web, llm

Web application testing

osint

misc, recon

OSINT gathering

forensics

misc, forensics, blueteam, reversing, stego, cracking

Digital forensics and incident response

pwn

misc, pwn, reversing, crypto

Binary exploitation and reverse engineering

mobile

misc, mobile, web, reversing

Mobile application security testing

cloud

misc, cloud, containers, networking, recon

Cloud and container security auditing

blockchain

misc, blockchain, web, crypto

Smart contract auditing and blockchain security

wireless

misc, wireless, networking

WiFi, Bluetooth, and SDR security

lightweight

misc, networking, recon, web, cracking

Hobby ethical hacking essentials (HTB, THM, bug bounty)

crackstation

misc, cracking, crypto

Hash cracking

blueteam

misc, blueteam, forensics, reversing, mobile, containers, networking, cloud, recon

Defensive security, IR, malware analysis

Modules

Module

Tools

Description

misc

40

Post-exploitation, social engineering, wordlists, resources, C2 (Docker + Loki)

networking

57

Port scanning, packet capture, tunneling, MITM, protocol tools

recon

82

Subdomain enumeration, OSINT, DNS, automated recon frameworks

web

60

Vulnerability scanning, fuzzing, SQLi, XSS, CMS scanners, API testing

crypto

15

RSA attacks, cipher analysis, hash attacks, constraint solving

pwn

36

Exploit frameworks, binary exploitation, fuzzing, payload generation

reversing

33

Disassemblers, debuggers, emulation, Java/Python reversing

forensics

57

Disk/memory forensics, file carving, timeline analysis, log analysis, hardware/serial

enterprise

80

Active Directory, Kerberos, Azure AD, credential harvesting, lateral movement

wireless

41

WiFi cracking, Bluetooth, SDR, rogue AP

cracking

34

Hash cracking (john, hashcat), brute force, wordlist generation

stego

15

Image/audio steganography, detection, StegCracker

cloud

22

AWS/Azure/GCP security auditing, Checkov

containers

15

Docker/Kubernetes security (Grype, Syft, Kubescape, kubeaudit)

blueteam

36

IDS/IPS, SIEM, incident response, threat intelligence, hardening, malware analysis (YARA, ClamAV, FLOSS, Capa, Loki)

mobile

18

Android/iOS app testing, APK analysis, MobSF (Docker)

blockchain

15

Smart contract auditing (Slither, Mythril, Foundry, Aderyn), blockchain forensics, Echidna (Docker)

llm

14

LLM red teaming, prompt injection, jailbreak testing, AI vulnerability scanning

Method

Count

Examples

Git clone

193

GitHub repos with auto-setup, resources, wordlists

System packages (apt/dnf/pacman/zypper)

166

nmap, wireshark, john, hashcat

pipx

137

sqlmap, impacket, bloodhound, volatility3

Go install

62

nuclei, subfinder, ffuf, httpx

Binary release

51

gitleaks, chainsaw, findomain, FLOSS, Capa, Loki, Syft, Kubescape

Build from source

23

massdns, duplicut, AFLplusplus, honggfuzz

Docker

13

Empire, MobSF, BeEF, BloodHound, TheHive, Cortex, PentAGI

Cargo (Rust)

8

feroxbuster, RustScan, pwninit, yara-x-cli

Ruby gem

6

wpscan, evil-winrm, brakeman

npm

5

promptfoo, apk-mitm, surya, solgraph

Special

5

Metasploit, Foundry, Steampipe, patator (curl-pipe installers), crypto venv bootstrap

Snap

1

zaproxy

Post-install scripts

All scripts support --help and need root on Linux (sudo); Termux needs no root.

Script

Purpose

Example

scripts/verify.sh

Check which tools are installed

sudo ./scripts/verify.sh --module web --skip-heavy

scripts/update.sh

Update all installed tools

sudo ./scripts/update.sh --skip-system

scripts/remove.sh

Remove tools by module

sudo ./scripts/remove.sh --module enterprise --yes

scripts/remove.sh --deep-clean

Purge caches and build artifacts

sudo ./scripts/remove.sh --deep-clean --yes

scripts/backup.sh

Back up and restore tool configs

sudo ./scripts/backup.sh backup

--deep-clean removes the Go module/build cache, the Cargo registry, pip/pipx/npm/gem caches, orphaned pipx venvs, stale symlinks, and log files. Add --remove-deps to also purge Rustup toolchains.

Non-system tools land in /usr/local/bin/ on Linux and $PREFIX/bin on Termux; system packages use their default location (/usr/bin/).

Method

Binary location (Linux)

Binary location (Termux)

Data location

pipx

/usr/local/bin/

$PREFIX/bin/

/opt/pipx/ or ~/.local/pipx/

Go

/usr/local/bin/

$PREFIX/bin/

/opt/go/ or ~/.go/

Cargo

/usr/local/bin/ (symlinked)

$PREFIX/bin/ (symlinked)

~/.cargo/

Git repos

/usr/local/bin/ (symlinked)

$PREFIX/bin/ (symlinked)

/opt/<repo>/ or ~/tools/<repo>/

Binary releases

/usr/local/bin/

Skipped (glibc incompatible with Bionic)

--

The .versions file records what was installed, how, and when.

If --enable-docker is set and Docker is missing, the installer stops and asks you to install Docker first.

Image

Module

Flag

Description

bcsecurity/empire

misc

--enable-docker --include-c2

Empire C2

spiderfoot/spiderfoot

misc

--enable-docker

SpiderFoot OSINT

beefproject/beef

web

--enable-docker

BeEF browser exploitation

opensecurity/mobile-security-framework-mobsf

mobile

--enable-docker

MobSF

specterops/bloodhound

enterprise

--enable-docker

BloodHound CE

trailofbits/echidna

blockchain

--enable-docker

Echidna smart contract fuzzer

checkmarx/kics:latest

cloud

--enable-docker

KICS infrastructure-as-code scanner

sagemath/sagemath:latest

crypto

--enable-docker

SageMath (Coppersmith, Groebner bases, curve arithmetic)

strangebee/thehive:latest

blueteam

--enable-docker

TheHive IR platform

thehiveproject/cortex:latest

blueteam

--enable-docker

Cortex analysis

zeek/zeek:latest

blueteam

--enable-docker

Zeek network analysis

wagga40/zircolite:latest

blueteam

--enable-docker

Zircolite EVTX detection

vxcontrol/pentagi:latest

llm

--enable-docker

PentAGI autonomous pentesting

Distro support

Debian/Ubuntu/Kali is the primary target and has the strongest test coverage: Kali and Parrot get the full apt set, while plain Debian/Ubuntu skip a handful of Kali-only packages. Fedora, Arch, and openSUSE skip the packages their repositories do not carry (the - entries in lib/distro_compat.tsv, several dozen per distro) and run in the integration workflow.

Platform

Status

WSL

Supported for installs and MCP use; the wireless module and kernel-level packages are skipped. No dedicated CI job, so validate release-critical changes in a local WSL distro. See Windows Defender false positives if the repo lives on a Windows-mounted path.

ARM (aarch64/armv7)

Supported, with automatic skips for x86-only binary releases and build-from-source tools. No dedicated CI job.

Termux (Android)

Supported without sudo. Docker, snap, binary releases, and build-from-source are skipped (Bionic incompatible). No dedicated CI job.

Windows (native)

Not supported. Use WSL.

macOS

The installer is not supported; use the Docker image. The MCP server runs on macOS in host mode (--local).

Other Linux distros

Distros without apt/dnf/pacman/zypper (NixOS, Gentoo, Void, Alpine, Slackware, …) are detected and blocked with a clear error. Use the Docker image.

Supply chain model

The installer downloads and runs code from the internet. On Linux it runs as root (sudo); on Termux it runs in the app's user sandbox.

  • System packages: signed by your distro's repositories (apt, dnf, pacman, zypper, pkg).

  • pipx/Go/Cargo/Gem/npm: fetched from their registries without signature verification; pipx tools are isolated in venvs.

  • Binary releases: SHA256-verified when the release publishes a checksum file, with a hard failure on mismatch. About half of the upstream releases publish none; --require-checksums or the --production preset fails those tools instead of installing them unverified. --fast disables all checksum verification, including for releases that publish checksums, and cannot be combined with the strict flags; keep it out of CI and production.

  • Runtime bootstraps: the rustup, uv, and cargo-binstall installers and NodeSource's setup script (run as root) are fetched over HTTPS and sanity-checked before they run, but not signature-verified. cargo-binstall then downloads prebuilt Rust binaries where available; --skip-source skips its installer on a fresh host, so Rust tools compile from crates.io.

  • Go SDK: SHA256-verified against go.dev when the API is reachable; strict mode fails if it is not.

  • Git repos: cloned at HEAD; dependencies go into isolated venvs, and setup.py is not executed.

  • Build from source: runs make, as root on Linux. Review what you build.

  • MCP Python dependencies: resolved by uv with a 3-day exclude-newer window, and Dependabot waits the same 3 days. This does not apply to the security tools themselves, which follow their upstream release channels.

  • Toolkit Docker images: Ubuntu and uv are digest-pinned; apt packages resolve from the current signed Ubuntu repositories, so rebuilds are not bit-for-bit reproducible.

  • Optional tool images: pulled by mutable tags, several of them latest. --production does not pin or verify them.

--production does not pin Git clones, language-package registries, or build-from-source tools; those track their upstream release channels.

Windows Defender false positives

On a Windows-mounted path (e.g. C:\Users\<you>\..., or any folder visible from Windows while you work in WSL), Microsoft Defender and other AV products may quarantine individual files. IOC tables, sample obfuscated PowerShell, malware-analysis snippets, and exploit strings in .claude/skills/, writeups/, and parts of mcp_server/ contain the same byte patterns real attackers use. Common verdicts include Trojan:Script/Wacatac.B!ml, HackTool:*, and generic Heur.*. These are false positives for a security toolkit.

  1. Exclude the repo folder (recommended on a personal dev box), from an elevated PowerShell:

    Add-MpPreference -ExclusionPath "C:\path\to\cybersec-toolkit"
  2. Restore files from quarantine via Windows Security → Virus & threat protection → Protection history → "Allow on device". This is per file and does not prevent re-detection.

  3. Keep the repo inside the WSL filesystem (e.g. ~/cybersec-toolkit). Defender does not scan WSL2's virtual disk by default. scripts/sync-wsl.sh does this for the MCP server.

Files removed by Defender show up as D in git status; restore them with git checkout -- <path> once the exclusion is in place.

MCP server

The server gives MCP-capable clients read access to the 670+ tool registry, install status and recommendations, and governed execution of installed tools. The agent knows every tool, which ones are available, and how to chain them.

Supported clients

Client

Integration

Status

Claude Code

.mcp.json (tracked) + .claude/skills/

Native configuration included

Claude Desktop

claude_desktop_config.json

Configuration example documented

OpenCode

opencode.jsonc (tracked) + .agents/skills/

Live tested

Codex

.codex/config.toml (tracked)

Native configuration included

Gemini CLI

GEMINI.md + .gemini/settings.json (tracked)

Native configuration included

GitHub Copilot

.mcp.json (CLI) + .github/copilot-instructions.md

CLI live tested; VS Code documented

Hermes Agent

User ~/.hermes/config.yaml

Live tested

OpenClaw

User ~/.openclaw/openclaw.json + .agents/skills/

Live tested

DeepSeek Harness (dsh)

$DSH_HOME/settings.yaml + .agents/skills/

Configuration example documented

Cursor / Cline / Goose

Client MCP settings + Agent Skills

Compatible through MCP; skills supported

Continue

Client MCP settings; rules/prompts for context

Compatible through MCP

LM Studio (>=0.3.17)

mcp.json; manual or MCP-provided context

Compatible through MCP

Ollama

MCP host in front of it

Compatible through an MCP host

Open WebUI

MCP-to-OpenAPI bridge

Compatible through an MCP host or bridge

Per-client setup is in docs/AI_CLIENTS.md; coordinating several agents across any MCP client is in docs/ORCHESTRATION.md.

The server is published to the official MCP Registry as io.github.26zl/cybersec-toolkit and listed on Glama:

What the AI can do

Tool

What it does

list_tools

List/filter all 670+ tools by module, method, or install status (includes URLs)

check_installed

Check if a tool is installed (6 detection strategies)

get_tool_info

Full details: method, module, URL, install/update/remove commands

get_module_info

Deep-dive a module: all tools, install status, which profiles use it

get_profile_tools

See every tool a profile installs, grouped by module

suggest_for_ctf

Curated tool recommendations for 14 CTF challenge categories

suggest_for_bounty

Bug bounty tool recommendations for 7 target types with methodology and common vulns

guided_assessment

Companion-first solve assistant for an authorized target: classifies the target/finding, returns triage gates, recommends skills, picks tools from all modules/profiles, and guides step by step; opt-in autonomous starts an auto-solver loop over run_tool, run_pipeline, and the separately gated run_script

get_cve_info

Map a CVE id or nickname (e.g. log4shell) to curated skills, registry tools, modules, and live NVD/KEV/EPSS lookup commands

recommend_install

Natural-language → profile/module/tool recommendation

list_profiles

All 14 profiles with tool counts and install commands

run_tool

Execute installed tools safely (sanitized args, network policy, rate limiting, audit logging); supports remote execution over SSH

run_pipeline

Pipe tools together without a shell (strings binary | grep flag)

run_script

Explicit opt-in Python/Bash execution, with per-script venv selection

manage_remote_hosts

Add, remove, list, and test SSH remote hosts for remote tool execution

The agent can query every tool and its install state, chain governed tool calls, parse the output, and pivot on what it finds. Script execution requires a separate opt-in.

External recon → attack surface (needs CYBERSEC_MCP_ALLOW_EXTERNAL=1, authorized scope only)

  • "Enumerate the attack surface for target.com and flag anything exploitable" — fans out amass / subfinder → resolves and probes with httpx → fingerprints with whatweb → runs nuclei templates → content discovery with ffuf, then ranks hosts by exposure and proposes next steps

  • "Found an open redirect on /go?url= — weaponize it" — verifies with curl, then builds an SSRF / OAuth-token-theft PoC and probes for an exploitable callback

Web exploitation

  • "Confirm and exploit the SQLi on the login endpoint" — sqlmap to confirm and dump (destructive --os-shell/--os-cmd are policy-blocked), then run_script to automate the auth bypass and pull just enough for a PoC

  • "GraphQL introspection is on — map it and hunt IDOR" — pulls the schema, generates queries, fuzzes object IDs, and diffs authenticated vs unauthenticated responses

Active Directory / internal

  • "Low-priv creds on 10.10.0.0/24 — find a path to Domain Admin" — collects with bloodhound, kerberoasts with impacket (GetUserSPNs.py), cracks the TGS in hashcat, then validates lateral movement with netexec, all on a Kali box over SSH (manage_remote_hosts, host mode)

  • "Check for DCSync rights and dump if the path exists" — enumerates replication ACLs, then runs secretsdump.py against the DC

Binary exploitation & reversing

  • "Build a ret2libc exploit for this 64-bit binary" — triages with checksec / readelf, finds gadgets with ROPgadget, leaks libc via a puts@plt call, then writes the full pwntools chain in venv="pwntools" and pops a shell locally

  • "Recover the algorithm from this stripped binary" — objdump / radare2 disassembly piped into targeted analysis, then a run_script reimplementation to verify behavior

Crypto

  • "Break this RSA — small e, several ciphertexts" — detects the attack (Håstad / common modulus / Wiener) and solves it with pycryptodome + sympy in a venv, returning plaintext

  • "This JWT is HS256 with a weak key" — cracks the signing secret and forges an admin token

Blue team · detection engineering

  • "Write a Sigma rule for this technique and convert it to my SIEM" — authors the rule and renders it for the target backend (Splunk / Elastic) via sigma-cli

  • "Hunt these Windows event logs for lateral movement" — runs chainsaw over the EVTX with Sigma rules, then summarizes hits by host and timeline

  • "Build YARA rules from these samples and scan the tree" — generates yara signatures and runs them recursively

DFIR · malware triage

  • "Timeline this memory dump" — sweeps volatility3 plugins (pslist, netscan, malfind) and chains them into one narrative

  • "Hunt for C2 beaconing in this pcap" — tshark / tcpdump extraction → suricata rules → flags periodic callbacks

  • "Statically triage this suspicious file" — file → strings → capa / yara, then extracts IOCs for enrichment

Cloud · containers · ops

  • "Audit this AWS account for public S3 and risky IAM" — runs prowler / scoutsuite and surfaces only the high-severity findings

  • "Scan this image and k8s manifests before deploy" — grype image scan plus kubescape config checks

  • "What's my redteam coverage — and fix the gaps" — diffs get_profile_tools("redteam") against install status and emits the exact install commands

Mobile · wireless · blockchain

  • "Static-analyze this APK for secrets and insecure storage" — apktool / jadx decompile → MobSF-style checks, then greps for keys and endpoints

  • "Audit this Wi-Fi capture" — parses the handshake and runs aircrack-ng / hashcat against it

  • "Review this Solidity contract for reentrancy" — runs slither / mythril and explains the findings

run_tool and run_pipeline are argument-sanitized, network-policed, rate-limited, and audit-logged, and destructive flags (--os-shell, -rf, --exploit, …) are blocked. run_script is off by default: enabling it runs arbitrary code with the server user's filesystem and network permissions (inside the VM by default, on your host with --local), and the external-target policy does not constrain it. Use only against systems you are authorized to test.

Sandbox and host mode

scripts/mcp-launch.sh starts the server inside a Kata Containers VM unless you pass --local or set CYBERSEC_SANDBOX_MODE=local. The VM is created by a Sandcastle isolated-sandbox provider (sandbox/kata.mjs) that pins the Kata runtime; the same provider can also back a Sandcastle agent-orchestration loop instead of an MCP client. Startup fails closed: a missing runtime, sandbox image, KVM device, Node.js, or launcher dependency stops the server with the reason instead of falling back to the host. ./install.sh --doctor reports sandbox readiness. Setup, tuning, and the threat model are in docs/SANDBOX.md.

  • The VM has its own tool set. The sandbox image ships the MCP server plus file, git, binutils, nmap, curl, and Python. Tools installed on the host are not visible inside it, and check_installed and the advisors report the guest's tools. Bake a profile into the image instead, and rebuild after pulling server changes, because the image carries its own copy of the server:

    docker build -f sandbox/Dockerfile --build-arg TOOLKIT_PROFILE=ctf -t cybersec-toolkit-sandbox:latest .
  • Files: the VM sees no host path unless CYBERSEC_SANDBOX_WORKSPACE names an absolute directory, which appears as /workspace (read-only with CYBERSEC_SANDBOX_WORKSPACE_RO=1). Pass guest paths to run_tool.

  • Network: the VM uses Docker's default bridge. Set CYBERSEC_SANDBOX_NETWORK=none for offline analysis, or a filtered Docker network to scope egress.

  • Privileges: uid 10001, all capabilities dropped, no-new-privileges, and 2 GB / 2 vCPUs / 512 processes by default. Raw-socket scans (nmap -sS, -sU, OS detection) need CYBERSEC_SANDBOX_CAP_ADD=NET_RAW.

  • Host mode only: remote execution over SSH (manage_remote_hosts) and venvs under ~/.ctf-venvs/ that you created on the host.

  • Audit: records leave the VM on a stderr side channel and are appended to the host log, so the trail outlives the VM.

Client setup

Claude Code reads the tracked .mcp.json:

{
  "mcpServers": {
    "cybersec-tools": {
      "command": "bash",
      "args": [
        "-lc",
        "cd \"$(git rev-parse --show-toplevel)\" && exec bash scripts/mcp-launch.sh"
      ],
      "env": {
        "CYBERSEC_MCP_ALLOW_EXTERNAL": "0",
        "CYBERSEC_MCP_ALLOW_SCRIPTS": "0"
      }
    }
  }
}

The login shell (bash -lc) picks up node and uv from your profile, so profile scripts must not print to stdout, which carries the MCP protocol. Other clients use the same launcher; from any directory:

bash /path/to/cybersec-toolkit/scripts/mcp-launch.sh           # Kata VM (default)
bash /path/to/cybersec-toolkit/scripts/mcp-launch.sh --local   # host, no VM boundary
  • Codex: the project .codex/config.toml resolves the Git root first, so it works from any subdirectory. If Codex ignores project config, copy the [mcp_servers.cybersec-tools] block into ~/.codex/config.toml.

  • Cursor / Continue / Cline / Goose: add the launch command in the client's MCP settings, with an absolute path if the client's working directory is not the repo root.

  • LM Studio (≥0.3.17): an MCP host itself. Add the server to its mcp.json (same mcpServers shape as .mcp.json) with an absolute path. MCP over LM Studio's API needs ≥0.4.0 and an MCP-capable endpoint such as /api/v1/chat or /v1/responses.

  • Ollama and other local models: a model runtime does not speak MCP. Put an MCP-capable host in front of it (OpenCode, Hermes, OpenClaw, Kit, LM Studio, Cline, Continue, Goose, or Open WebUI through an MCP→OpenAPI bridge such as mcpo) and point that host at the launcher.

Start with this one server and keep scripts and external targets off unless you have an authorized scope; prefer hosts with human-in-the-loop tool approval. Vendor-neutral repository instructions live in AGENTS.md; Claude Code reads CLAUDE.md, Gemini CLI reads GEMINI.md.

The server speaks stdio, so a Windows client can launch it inside WSL. This runs it directly in WSL, without the Kata VM (the equivalent of --local):

{
  "mcpServers": {
    "cybersec-tools": {
      "command": "wsl",
      "args": [
        "-d", "kali-linux",
        "bash", "-lc",
        "cd ~/cybersec-toolkit/mcp_server && uv run fastmcp run server.py --transport stdio --no-banner"
      ]
    }
  }
}

Use a clone inside the WSL filesystem, or run scripts/sync-wsl.sh from a Windows checkout to copy the server to ~/cybersec-toolkit in WSL (uv cannot create a venv on NTFS). Passing CYBERSEC_MCP_* variables from Windows needs WSLENV; see mcp_server/README.md.

Runs the server in the installer image. This is an ordinary container, not the Kata VM, and its user has passwordless sudo:

{
  "mcpServers": {
    "cybersec-tools": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "CYBERSEC_MCP_ALLOW_EXTERNAL=0", "-e", "CYBERSEC_MCP_ALLOW_SCRIPTS=0",
        "--entrypoint", "bash", "cybersec-toolkit",
        "-c",
        "cd /opt/cybersec-toolkit/mcp_server && uv run fastmcp run server.py --transport stdio --no-banner"
      ]
    }
  }
}

Script execution

run_script lets the agent write and run Python or Bash. Enable it with "CYBERSEC_MCP_ALLOW_SCRIPTS": "1" in the client's env block. Scripts get the server process's filesystem and network permissions (the VM by default, your user in --local mode), and CYBERSEC_MCP_ALLOW_EXTERNAL does not constrain them. Review generated code and scope before opting in.

Python libraries without console scripts live in named venvs under ~/.ctf-venvs/ (override with CYBERSEC_MCP_VENVS_DIR), selected per script with the venv parameter. ./install.sh --module crypto creates ~/.ctf-venvs/crypto with pycryptodome, sympy, gmpy2, numpy, z3, fpylll (+ cysignals, which fpylll needs at import time but does not declare), and cypari2, so the agent can call run_script(code, venv="crypto"). In the sandbox, venvs are guest paths: build the image with TOOLKIT_PROFILE=ctf to get crypto. Some packages need an older Python, for example pwntools:

python3.12 -m venv ~/.ctf-venvs/pwntools
~/.ctf-venvs/pwntools/bin/pip install pwntools z3-solver

Without a named venv, scripts see FastMCP and the standard library; cd mcp_server && uv sync --extra ctf-core adds requests, pycryptodome, beautifulsoup4, Pillow, and NumPy.

Reusable helpers the agent writes for you (exploits, multi-step solvers, parsers, protocol helpers) are saved under manual_scripts/. In companion mode the agent proposes a script and runs it only after you approve; in autonomous mode it writes and runs scoped helpers when tools and pipelines stop making progress. Plain recon and HTTP requests stay run_tool calls.

Test the server

cd mcp_server && uv run fastmcp dev server.py

This opens the MCP Inspector for exercising each tool interactively. mcp_server/README.md covers Claude Desktop setup and the full server reference.

Agent Skills

872 Agent Skills live in .claude/skills/, the canonical tree that Claude Code discovers directly. Skills load on demand for the task at hand instead of occupying context permanently. 31 are project-authored and 841 are curated from open-source projects, each attributed in THIRD_PARTY_NOTICES.md:

  • 10 project developer skills (add-tool, validate-all, module-scaffold, writeup-template, mcp-sync-check, security-wordlists, security-payloads, guided-assessment, skill-dependency-audit, skill-curation-router)

  • 4 cross-skill coordinators (finding-triage, security-comms, authorization-gate, evidence-hygiene) that route findings, communication, authorization checks, and evidence sanitization

  • 7 coverage-gap anchors (GRC/privacy, AI/LLM security, IoT/embedded/hardware, mainframe, telecom/5G, SAP/ERP, supply-chain/product security)

  • 6 CTF methodology skills (ctf-crypto, ctf-pwn, ctf-web, ctf-rev, ctf-forensics, ctf-stego) and 4 bug bounty methodology skills (bounty-recon, bounty-web, bounty-api, bounty-mobile)

  • 754 operational how-tos from mukul975/Anthropic-Cybersecurity-Skills (Apache 2.0)

  • 58 offensive methodology skills from SnailSploit Claude-Red (MIT)

  • 14 code audit skills from Trail of Bits (CC-BY-SA 4.0)

  • 10 bug bounty workflow skills from BugHunter (claude-bug-bounty) (MIT)

  • 4 high-level workflows from Transilience (MIT)

  • 1 coding-agent workflow skill from multica-ai/andrej-karpathy-skills (MIT)

Source and category index: .claude/skills/SKILLS.md. Ranking and curation: .claude/skills/CURATION.md and curation.json, regenerated with python3 scripts/curate_claude_skills.py --write.

As a Claude Code plugin. The repo doubles as a plugin marketplace, so any project can pull in the skill library without cloning:

/plugin marketplace add 26zl/cybersec-toolkit
/plugin install cybersec-toolkit@cybersec-toolkit

The plugin carries the skills only; the MCP server is configured separately (see MCP server).

In other clients. OpenCode, Codex, Gemini CLI, GitHub Copilot, Cursor, Cline, Goose, Hermes, and OpenClaw support Agent Skills through their own paths. scripts/sync-skills.sh mirrors .claude/skills/ into the git-ignored .agents/skills/ for clients that read that location (--check reports drift; make setup runs it). Continue and LM Studio do not document automatic skill discovery; give them selected skill content as rules or context instead. Details: docs/AI_CLIENTS.md.

Helper-script dependencies. Some vendored skills include helper scripts with optional Python imports, declared in .claude/skills/requirements.txt and generated from the import inventory. scripts/validate_claude_skills.py checks skill metadata, index counts, curation freshness, and helper-script syntax.

python3 scripts/audit_skill_dependencies.py --check-declared   # verify declarations
python3 -m pip install -r .claude/skills/requirements.txt      # optional, ideally in a venv

Development

Contributions are welcome: testing installs on different distros, adding missing tools, fixing package mappings, improving MCP workflows, writing example use cases, and reporting rough edges from real CTF, lab, bug bounty, pentest, DFIR, or defensive work. Open an issue for bigger changes or send a focused PR for small fixes; CONTRIBUTING.md has the validation checklist.

git clone https://github.com/26zl/cybersec-toolkit.git && cd cybersec-toolkit
make setup    # submodules, MCP deps, sandbox deps, skill mirror
make check    # lint, validators, bats, pytest, sandbox tests

make help lists every target; the raw commands are in AGENTS.md. The MCP Python project resolves dependencies with uv and exclude-newer = "3 days", so fresh releases are ignored for 72 hours to limit the blast radius of a compromised upload; Dependabot and the weekly uv update workflow use the same cooldown. Run the shell tests on Linux, macOS, or WSL: native Windows checkouts can rewrite the Bats submodules with CRLF and fail with $'\r'.

Star History

License

MIT License. See LICENSE.

The repository also redistributes third-party components under their own terms, including some under CC-BY-SA-4.0 (ShareAlike, not relicensable to MIT). If you redistribute or adapt bundled content, follow those terms; see THIRD_PARTY_NOTICES.md.

Contribution workflow: CONTRIBUTING.md. Community expectations: CODE_OF_CONDUCT.md. Vulnerability reporting: SECURITY.md.

Disclaimer

This project is provided for educational, defensive, and explicitly authorized security testing only. Use it only on systems you own or have written permission to assess, and follow all applicable laws, rules of engagement, third-party tool licenses, and service terms.

The toolkit includes dual-use offensive and defensive tools. Some commands can scan networks, execute exploits, modify systems, or trigger security alerts. MCP/AI integrations are guarded by safety policies, but users remain responsible for reviewing scope, prompts, commands, and outputs before running actions.

This repository does not redistribute the security tools themselves; it installs publicly available, open-source projects from their official upstream sources at install time. It is intended for lawful, authorized use only.

The project is provided "as is", without warranty. Maintainers are not responsible for misuse, damage, data loss, service disruption, or legal consequences from using this toolkit.

Third-party content is bundled under its original license — see THIRD_PARTY_NOTICES.md.

Available Tools

15 tools
check_installedA
Read-onlyIdempotent

Check if a specific cybersecurity tool is installed on the system.

Uses multiple detection strategies: .versions tracking, PATH lookup, pipx binary name fallback, /opt directory check, and docker image check.

When host is provided, checks installation on the remote host via SSH using 'which '.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNoOptional remote host name (as configured via manage_remote_hosts).
tool_nameYesName of the tool to check (as listed in tools_config.json).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so safety is covered. The description adds real value beyond that by enumerating the five detection strategies and disclosing that a host argument switches to an SSH-based 'which <binary>' check. Only the absence of failure-mode or return-behavior detail keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, followed by the detection strategies and the remote-host behavior. Three short paragraphs, every sentence carrying information, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be recounted, and annotations cover the safety profile. Between purpose, detection strategies, and remote-host semantics, an agent has everything needed to call this correctly in either local or remote mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented and the baseline is 3. The description still adds meaning by explaining what host actually causes (remote SSH installation check) and anchoring tool_name to tools_config.json, going beyond the schema's phrasing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Check if a specific cybersecurity tool is installed on the system.' This is clearly distinct from run_tool, list_tools, or get_tool_info. It does not explicitly name or route away from those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by explaining that supplying host triggers a remote SSH check, which hints at the local-vs-remote decision. However, it never says when to use this versus list_tools or get_tool_info, and offers no exclusions or prerequisites. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cve_infoA
Read-onlyIdempotent

Map a CVE to the toolkit's tools, skills, and modules, plus live-lookup commands.

Local-first and deterministic: accepts a CVE id (e.g. "CVE-2021-44228") or a common nickname (e.g. "log4shell", "eternalblue", "zerologon", "printnightmare") and returns the curated exploitation skills, mapped registry tools with install status, and relevant modules.

For live CVSS / CISA KEV / EPSS data it returns ready-to-run run_tool("curl", ...) commands rather than fetching itself — those hit external hosts and are subject to the CYBERSEC_MCP_ALLOW_EXTERNAL policy. Always clear the authorization-gate skill before testing.

ParametersJSON Schema
NameRequiredDescriptionDefault
cveYesA CVE id (CVE-YYYY-NNNN) or a known vulnerability nickname.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: local-first/deterministic behavior, that it returns ready-to-run run_tool("curl", ...) commands rather than performing network fetches, that external lookups are gated by CYBERSEC_MCP_ALLOW_EXTERNAL, and that an authorization-gate skill must be cleared first. This explains the openWorldHint=false annotation and the safety profile rather than merely restating it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the one-line purpose, then groups the input contract, return behavior, and safety prerequisite into distinct, information-dense paragraphs. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be enumerated, and the description still previews the return shape (skills, tools with install status, modules). It is complete for normal use, though it omits edge-case behavior such as an unknown or ambiguous CVE id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by listing concrete nickname examples ('log4shell', 'eternalblue', 'zerologon', 'printnightmare') and the CVE-YYYY-NNNN form, which helps the agent normalize input beyond the schema's brief wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('map a CVE to the toolkit's tools, skills, and modules') and specifies the accepted input forms (CVE id or nickname). It is readily distinguishable from siblings like get_tool_info and get_module_info, which operate on the toolkit's own inventory rather than mapping an external CVE.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when it acts locally versus when it defers to curl commands, and names a hard prerequisite ('always clear the authorization-gate skill before testing'). It gives clear operating context but does not name an alternative tool for cases the caller might prefer instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_module_infoB
Read-onlyIdempotent

Get full details about a module: description, all tools, and management commands.

ParametersJSON Schema
NameRequiredDescriptionDefault
moduleYesModule name (e.g. "web", "pwn", "forensics").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it returns the module description, its tools, and management commands, but says nothing about unknown-module behavior or result size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the resource first and enumerates the payload; nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one documented parameter, read-only annotations, and an output schema that carries the return structure, the description covers everything needed to invoke the tool correctly. Only the installed/prerequisite condition is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, including examples ('web', 'pwn', 'forensics'). The description adds no meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('module') plus the shape of what comes back (description, tools, management commands). It is distinguishable from sibling get_tool_info by operating on a module rather than a tool, though it never explicitly names that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite (e.g. must the module be installed first?), and no routing toward alternatives like get_tool_info or list_tools. The agent must infer the context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_profile_toolsA
Read-onlyIdempotent

List every tool that a specific profile would install.

Given a profile name, returns the complete list of tools grouped by module, with install status for each. This lets you see exactly what you get before running the install command.

ParametersJSON Schema
NameRequiredDescriptionDefault
profileYesProfile name (e.g. "ctf", "redteam", "web", "full").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is fully covered. The description adds that results are grouped by module with per-tool install status, which is useful, but does not disclose anything about auth or rate limits beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and efficiently sized at two short sentences. There is mild redundancy, since 'List every tool that a specific profile would install' and 'returns the complete list of tools' carry nearly the same information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a one-parameter, read-only tool with an output schema present and annotations covering safety, the description supplies everything needed to invoke it correctly. Return-value details are appropriately left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'profile' parameter already documents example values ('ctf', 'redteam', 'web', 'full'). The description only restates 'Given a profile name,' adding no format or validation detail beyond the schema — the baseline 3 for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'List every tool that a specific profile would install' — and adds scope detail, returning tools grouped by module. This clearly separates it from siblings like list_tools (all tools) and get_tool_info (single tool).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies a concrete use context: 'This lets you see exactly what you get before running the install command,' which implies previewing before install. It stops short of naming explicit alternatives (e.g., list_profiles for discovering valid names), so it is clear but not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_tool_infoB
Read-onlyIdempotent

Get detailed information about a cybersecurity tool.

Returns the tool's install method, module, URL, installation status, module description, and management commands (install, update, remove).

ParametersJSON Schema
NameRequiredDescriptionDefault
tool_nameYesName of the tool to look up.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so the safety profile is covered. The description adds that the result includes install method, module, URL, install status and management commands, which is useful behavioral context, but it says nothing about what happens for an unknown tool_name or any rate limits. Given annotations carry the main burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, no filler. The second sentence enumerating return fields is slightly redundant given an output schema exists, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since an output schema exists, the description need not explain return values, so its core obligation is purpose plus routing. Purpose is covered; what's missing is any signal about when to prefer this over list_tools, check_installed, or get_module_info, and what happens on a bad tool name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single documented parameter, so the baseline 3 applies. The description adds no format, casing, or matching-rule detail (e.g., exact name vs alias) beyond the schema's 'Name of the tool to look up.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Get detailed information about a cybersecurity tool') and even enumerates the fields returned, so an agent knows exactly what comes back. It does not, however, differentiate itself from nearby siblings like get_module_info or check_installed, which also return tool/module metadata.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use statement. The description never mentions list_tools (browse the catalog) or check_installed (installation state), even though this tool overlaps heavily with both, leaving the agent to infer routing from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guided_assessmentA
Destructive

Plan, guide, or autonomously solve a security task over the MCP toolchain.

An orchestrator on top of the registry, advisors, install checks, audit logging, and execution policy. Bootstrap commands use the governed execute_tool() path, so target scope, external-network, shell-injection, and blocked-flag checks apply.

By DEFAULT it auto-detects the right workflow + tools for the problem (workflow/ target_type="auto") and acts as a companion: it returns classification, triage gates, recommended skills, reporting next steps, a plan, tool install status, next actions, and the full MCP toolchain surface WITHOUT auto-running commands in this initial call. The agent can then run tools step by step as the user approves. The heaviest mode (autonomous) starts the auto-solver contract: it bootstraps triage, then the client agent continues with the full MCP toolchain (registry/advisors/install checks/run_tool/run_pipeline and separately gated run_script). When registry tools and pipelines are not enough, autonomous mode may create, save, and run scoped helper scripts for the user, persisting reusable ones under manual_scripts/. Simple recon/HTTP commands such as curl remain run_tool calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"companion" (default) or "autonomous" (opt-in).companion
targetYesURL, hostname/IP, or local file path to assess.
findingNoOptional short finding summary to classify for triage/report routing. Raw finding text is used locally but not echoed in the result.
workflowNo"auto" (default — inferred), "bounty", "ctf", or "generic".auto
intensityNo"low" (default) or "medium". Medium may include low-volume nmap.low
max_stepsNoMaximum number of bootstrap steps autonomous mode auto-executes.
target_typeNo"auto" (default — inferred from the target) or an explicit type: bounty type (web_app/api/cloud/network/iot/mobile_app) or CTF category.auto
authorization_confirmedNoRequired before any network step executes.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (destructiveHint=true, openWorldHint=true), the description discloses governed execution checks ('target scope, external-network, shell-injection, and blocked-flag checks'), confirms no auto-run in the default call, and explains that autonomous mode may create, save, and run scoped helper scripts, persisting reusable ones under manual_scripts/. It also states that authorization_confirmed is required before any network step executes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and is dense with useful operational detail. It is a bit paragraph-heavy and could be more scannable, but given the complexity of the orchestration behavior, most sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be fully explained, yet the description still summarizes what companion mode returns and what autonomous mode bootstraps. Combined with the annotations and 100% schema coverage, the definition gives an agent enough context to choose and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter semantics already live in the input schema. The description largely repeats default values and enum-like options (mode companion/autonomous, workflow auto, target_type auto) without adding substantial syntax or format guidance beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific compound verb+resource: 'Plan, guide, or autonomously solve a security task over the MCP toolchain.' It immediately positions the tool as an 'orchestrator on top of the registry, advisors, install checks, audit logging, and execution policy,' which cleanly distinguishes it from sibling tools like run_tool, run_pipeline, and suggest_for_*.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly explains the default companion mode versus opt-in autonomous mode, and explicitly routes simple recon/HTTP commands to run_tool instead. It does not, however, give a full when-not-to-use rule or spell out the exact conditions under which autonomous mode should be preferred over companion mode, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_profilesA
Read-onlyIdempotent

List all 14 available installation profiles with details.

Each profile is a curated set of modules targeting a specific use case. Shows module count, tool count, and install command for each profile. Profiles range from 'osint' (2 modules) to 'full' (18 modules, 670+ tools).

Returns: All profiles with descriptions, module lists, tool counts, and install commands.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and closed-world behavior, so the safety profile is covered. The description adds genuine context beyond that: exactly 14 profiles, the osint-to-full size range, and that each is a curated module set. The 'Returns' sentence overlaps the output schema and is somewhat redundant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and size, and the supporting sentences are short. The trailing 'Returns:' block restates what the earlier sentence ('Shows module count, tool count, and install command') already established, which is mild duplication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not strictly required, and the description supplies scope, size, and content shape. It is essentially complete for a zero-param read tool, with the only gap being guidance on when to prefer it over adjacent discovery tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. Schema coverage is 100% and the description correctly implies no filtering options.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all 14 available installation profiles') with explicit scope and quantitative detail. Clear enough to distinguish from siblings like get_profile_tools or recommend_install, which deal with a single profile or a recommendation rather than the full catalog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the discovery framing but never stated: it does not say when to call this versus get_profile_tools, recommend_install, or check_installed. No exclusions or prerequisites are given, though the context is reasonably inferable for a list tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_toolsA
Read-onlyIdempotent

List and filter the 670+ cybersecurity tools in the registry.

Returns the tools drawn from tools_config.json — the same registry the installer and the advisors share — with the total count, the filters still available to narrow the results, and one entry per tool. Combine the filters to scope the list: module="web" for web tools, method="pipx" for Python-packaged tools, installed_only=True for only what is on this host. Start here to discover what exists before check_installed or get_tool_info.

ParametersJSON Schema
NameRequiredDescriptionDefault
methodNoFilter by install method. One of apt, pipx, go, cargo, gem, git, binary, docker, snap, special, source, npm.
moduleNoFilter by module (e.g. "web", "pwn", "forensics"). 18 modules available.
installed_onlyNoIf true, only return tools that are currently installed.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, closed-world behavior, so the safety profile is covered. The description adds valuable context about data provenance (drawn from tools_config.json, shared with installer/advisors) and return shape (total count, available filters, one entry per tool), which goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then provides provenance, filter examples, and entry-point guidance. Efficient and well-ordered, though the provenance sentence is slightly verbose for a list tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description fills remaining gaps: data source, filter combinations, and the tool's role in the workflow relative to siblings. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds semantic value by giving concrete examples of the three filters and showing combinations, which helps an agent reason about filter interaction beyond the schema's individual parameter docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list/filter) and resource (670+ cybersecurity tools in the registry), and grounds it in the shared tools_config.json registry. Distinguishes from siblings by naming check_installed and get_tool_info as downstream steps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Start here to discover what exists before check_installed or get_tool_info,' giving a clear entry-point role, and offers concrete filter combinations (module="web", method="pipx", installed_only=True) to guide use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_remote_hostsA
Destructive

Add, list, test, or remove the SSH hosts that run_tool can target remotely.

Manages the remote-host registry that lets run_tool (and check_installed) run a tool on a remote Kali/Linux box over SSH instead of locally, so the tool only has to be installed on the remote. The action selects the operation: "list" shows every configured host; "add" registers or updates a host (needs name and hostname, plus optional user, port, ssh_key, and a tool_allowlist that restricts which tools may run there); "remove" deletes a host by name; "test" opens an SSH connection to confirm the host is reachable.

Connections use StrictHostKeyChecking=accept-new, so the key presented on the first connection is pinned in ~/.ssh/known_hosts and any later change is rejected. Verify that first fingerprint out-of-band for a host you do not control, or add the key to known_hosts before "test".

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHost name (required for add/remove/test).
portNoSSH port (default 22).
userNoSSH username (default "kali").kali
actionYesOperation to perform: "list", "add", "remove", or "test".
ssh_keyNoPath to SSH private key (e.g. "~/.ssh/id_kali").
hostnameNoIP address or hostname of the remote machine (required for add).
descriptionNoHuman-readable description of the host.
tool_allowlistNoComma-separated list of allowed tool names (e.g. "nmap,gobuster,sqlmap"). None means all tools allowed.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-idempotent, and the description adds genuinely new behavioral context: the StrictHostKeyChecking=accept-new policy, that the first-seen key is pinned in ~/.ssh/known_hosts and later changes are rejected, and the out-of-band fingerprint verification advice. That is security-relevant disclosure beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then enumerates actions, then ends with the SSH key-pinning caveat. Every paragraph adds information, though the per-action enumeration is dense and slightly overlaps the schema's own descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-action mutation tool with side effects, an output schema present (so returns need no explanation), and full annotation coverage, the description supplies the missing pieces: action semantics, credential requirements, and the connection-security caveat. An agent has enough to select actions and call them correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real meaning: which parameters each action requires (name for add/remove/test, hostname for add) and what tool_allowlist actually does (restricts which tools may run on that host, with None meaning all). It stops short of restating defaults like user='kali' or port=22.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb set (add/list/test/remove) and a concrete resource (SSH remote hosts), and names the dependent sibling tools run_tool and check_installed. An agent can immediately distinguish this registry-management tool from the execution tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains that the action parameter selects the operation and gives per-action conditions (add needs name+hostname, remove takes a name, test opens a connection), plus the prerequisite context that hosts enable remote execution via run_tool. It does not, however, state explicit when-not-to-use cases or an ordering rule such as 'add before test before run_tool'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_installA
Read-onlyIdempotent

Recommend which profile, modules, or individual tools to install.

Analyzes a natural-language description of what the user wants to do and recommends the best installation approach — from a full profile down to just a few individual tools. Avoids installing everything when only a subset is needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesNatural-language description of what the user wants to do.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds useful context by clarifying this is an advisory recommendation engine rather than an installer, despite the 'install' in the name, and explains the over-install avoidance behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action in the first sentence, followed by scope and rationale. All sentences are relevant, though the second sentence could be slightly trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a present output schema, full parameter documentation, and complete annotations, the description covers everything an agent needs to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'task' parameter, so the schema already explains it. The description's reference to a 'natural-language description of what the user wants to do' merely restates the schema, adding no syntax or format detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (recommend) and resource (profile, modules, or individual tools to install), and scopes the output range from a full profile down to a few tools. It does not explicitly name or contrast with siblings like list_profiles or get_profile_tools, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage ('Analyzes a natural-language description of what the user wants to do') and gives a rationale ('Avoids installing everything when only a subset is needed'), but provides no explicit when-to-use guidance or alternatives among the many sibling tools (list_profiles, get_profile_tools, check_installed).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pipelineA
Destructive

Execute a pipeline of tools, piping stdout from each step into stdin of the next.

Replaces shell piping (e.g. strings binary | grep flag) with a safe, no-shell alternative. Each step is validated individually (allowlist, argument sanitization, policy checks) before any process starts.

Each step's stdout and stderr are bounded to 200KB as they are read, and an intermediate step's bounded stdout is what gets piped into the next step. If any step hits that cap, the returned truncated flag is set and the final stdout carries a truncation marker.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNoReserved for future use. Currently only local execution is supported.
stepsYesList of dicts, each with 'tool' (required) and 'args' (optional) keys. Max 10 steps per pipeline.
timeoutNoGlobal timeout for entire pipeline in seconds (default 120, max 300).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, but the description still adds real value: per-step validation before any process starts, a 200KB per-step stdout/stderr cap, a truncation marker, and the returned `truncated` flag. It omits what happens to earlier steps if a later step fails, which matters for a destructive pipeline, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the motivating example, then the safety/bounding guarantees. Every sentence carries concrete information (validation order, 200KB cap, truncation marker) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-described, and the description covers the semantics an agent needs: step structure, validation, bounding, and truncation signaling. The main residual gap is failure/partial-execution behavior across steps in a destructive multi-step run.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself documents `host` (reserved), the step-list shape with the max-10 constraint, and the timeout default/max. The description adds no further parameter-level syntax or semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Execute a pipeline of tools') plus the precise mechanism (stdout piped to stdin of the next step), which implicitly separates it from the single-invocation siblings run_tool and run_script. It does not explicitly name those alternatives, so an agent still has to infer the routing, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for it: replacing shell piping such as `strings binary | grep flag` with a safe, no-shell alternative. That is a concrete usage trigger, though it stops short of explicitly contrasting with run_tool/run_script or stating when not to use a pipeline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_scriptA
Destructive

Write and execute a Python or Bash script, returning its output.

Writes the code to a temporary file, executes it via python3/bash, and returns stdout/stderr. The temp file is deleted after execution. Requires CYBERSEC_MCP_ALLOW_SCRIPTS=1. This is an explicit full-code execution opt-in: scripts are not OS-sandboxed and are not constrained by CYBERSEC_MCP_ALLOW_EXTERNAL.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe script source code to execute.
venvNoOptional Python venv name from ~/.ctf-venvs/ (e.g. "pwntools"). Allows using a different Python with specific packages installed. Ignored for language="bash". If not set, uses the MCP server's Python.
timeoutNoMaximum execution time in seconds (default 120, max 300).
languageNo"python" (default) or "bash".python
working_dirNoWorking directory for the script (default: system temp dir).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/openWorld/non-idempotent, and the description meaningfully extends them: it discloses the temp-file write-then-delete lifecycle, the interpreter used, and two independent gating/containment flags. The non-sandboxing warning is exactly the kind of consequence an agent needs before calling a mutating execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does and the effect, followed by execution mechanics and then the opt-in/containment caveats. Every sentence contributes information an agent would not otherwise have.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return-value explanation is unnecessary, and the description covers the remaining risk surface: execution mechanism, cleanup, and the two environment flags governing permission and containment. Nothing material is missing for a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the venv, timeout, language, and working_dir semantics are already fully documented in the schema, including defaults and the venv/bash interaction. The description adds no parameter-level detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb+resource pair ('Write and execute a Python or Bash script') plus the scope of the effect ('returning its output'), which is enough to separate it from sibling run_tool and run_pipeline. Nothing about the action is left ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the precondition for use (CYBERSEC_MCP_ALLOW_SCRIPTS=1) and the security posture (not OS-sandboxed, not bound by CYBERSEC_MCP_ALLOW_EXTERNAL). It stops short of naming when to reach for run_tool/run_pipeline instead of scripting an ad-hoc program, so it lacks an explicit alternative-routing clause.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_toolA
Destructive

Execute an installed cybersecurity tool or system utility and return its output.

Runs tools from the 670+ registry as well as ~120 standard system utilities (strings, file, curl, grep, base64, xxd, jq, etc.) that are allowed without being in the registry. Arguments are sanitized to prevent shell injection. Timeout is clamped to 1-300s. Output is truncated at 200KB.

Network tools (including curl, wget, ping, etc.) are restricted to local/private targets by default. Set CYBERSEC_MCP_ALLOW_EXTERNAL=1 to allow external targets.

When host is provided, the tool is executed on the remote host via SSH. The tool does not need to be installed locally — only on the remote host.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoCommand-line arguments as a string (e.g. "--version" or "-sV 10.0.0.1").
hostNoOptional remote host name (as configured via manage_remote_hosts).
timeoutNoMaximum execution time in seconds (default 120, max 300).
tool_nameYesName of the tool to run (registry tool or system utility).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/openWorld/non-idempotent, and the description adds substantial context on top: argument sanitization against shell injection, timeout clamping to 1-300s, 200KB output truncation, default local/private network restriction with an env-var escape hatch, and SSH semantics where the tool need only exist on the remote host.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then layered constraints in short, well-separated paragraphs. Every sentence carries operational information (sanitization, timeout clamp, truncation, network policy, SSH) with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a high-risk execution tool: safety profile is covered by annotations, argument and host semantics by the description, and return values by the existing output schema. An agent has everything needed to call it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning for 'host' (executed remotely via SSH, tool need not be installed locally) and for 'args' (sanitized to prevent shell injection). Timeout and tool_name semantics are largely left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Execute') and resource ('an installed cybersecurity tool or system utility') plus the outcome ('return its output'). It clearly separates itself from siblings like run_script and run_pipeline by scoping to registry tools and a fixed set of allowed system utilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives rich context: registry vs. ~120 built-in system utilities, network tools restricted to local/private targets unless CYBERSEC_MCP_ALLOW_EXTERNAL=1, and remote execution when host is supplied. It never explicitly says when to prefer run_script/run_pipeline or the suggest/recommend siblings, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_for_bountyA
Read-onlyIdempotent

Suggest cybersecurity tools for a bug bounty target type.

Provides curated tool recommendations with installation status, methodology steps (starting with scope verification), common vulnerabilities, and quick wins for 7 target types: web_app, api, mobile_app, cloud, network, iot, llm.

Also accepts aliases: web/webapp (web_app), rest/graphql (api), android/ios/mobile (mobile_app), aws/azure/gcp/k8s (cloud), infra/infrastructure (network), firmware/embedded (iot).

ParametersJSON Schema
NameRequiredDescriptionDefault
target_typeYesType of bug bounty target (e.g. "web_app", "api", "cloud").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so safety is covered. The description adds real behavioral value by disclosing what the response contains: curated tools with installation status, methodology steps starting with scope verification, common vulnerabilities, and quick wins.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose in the first sentence, then supporting detail. The alias list is somewhat long but each entry maps useful input strings, so it earns its place; no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, yet the description still sketches the response shape. Combined with the enum-free parameter being fully enumerated, the definition is complete enough for an agent to invoke correctly, missing only explicit sibling routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% but the schema has no enum for target_type; the description supplies the seven valid values (web_app, api, mobile_app, cloud, network, iot, llm) and the accepted aliases, which the schema alone does not convey. This meaningfully extends parameter semantics beyond the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Suggest cybersecurity tools') plus the scope ('for a bug bounty target type'), immediately distinguishing it from the sibling suggest_for_ctf. An agent can tell what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context of use is implied by 'bug bounty target type' and the enumerated target list, but there is no explicit when-to-use / when-not guidance and no routing to alternatives such as suggest_for_ctf or guided_assessment. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_for_ctfA
Read-onlyIdempotent

Suggest cybersecurity tools for a CTF challenge category.

Provides curated tool recommendations with installation status for 14 challenge types: web, crypto, pwn, reversing, forensics, stego, misc, networking, wireless, osint, cloud, mobile, blockchain, llm.

Also accepts aliases: re/rev (reversing), binary/exploitation (pwn), steganography (stego), network (networking), recon (osint), etc.

ParametersJSON Schema
NameRequiredDescriptionDefault
challenge_typeYesType of CTF challenge (e.g. "web", "crypto", "pwn").

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so safety is covered. The description adds useful operational context beyond them: that recommendations are curated and include installation status, and that 14 named categories are supported. It does not discuss result size or ordering, but the output schema presumably covers returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the category list and alias list in separate, scannable sentences. Slightly list-heavy, but each element earns its place by supporting valid parameter values.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, output-schema-backed read tool, the description covers what it does, the full valid value space, and the fact that install status is included. Nothing critical is missing; only cross-tool routing is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter, so the baseline is 3, but the description goes further by enumerating all 14 accepted challenge_type values and a set of aliases (re/rev, binary/exploitation, steganography, network, recon). Since the schema declares no enums, this enumeration is genuinely additive and prevents wrong-value calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Suggest cybersecurity tools for a CTF challenge category'), which the agent can immediately act on. It does not name or contrast with the obvious sibling suggest_for_bounty, so differentiation from that near-twin is left implicit via the word 'CTF'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a CTF challenge category' implies the use context, but there is no explicit when-to-use / when-not guidance and no reference to the adjacent suggest_for_bounty or recommend_install tools. The agent must infer that this is the CTF-scoped path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.2.1
    • Changedlist_tools2 fields changed
      • changedInput schema / properties / method / anyOf
        Previous value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "enum": [
        +      "apt",
        +      "pipx",
        +      "go",
        +      "cargo",
        +      "gem",
        +      "git",
        +      "binary",
        +      "docker",
        +      "snap",
        +      "special",
        +      "source",
        +      "npm"
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / method / description
        Previous value: -"Filter by install method (apt, pipx, go, cargo, gem, git, binary, docker, snap, special, source, npm)."New value: +"Filter by install method. One of apt, pipx, go, cargo, gem, git,\nbinary, docker, snap, special, source, npm."
    • Changedmanage_remote_hosts2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"One of \"list\", \"add\", \"remove\", \"test\"."New value: +"Operation to perform: \"list\", \"add\", \"remove\", or \"test\"."
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "list",
        +  "add",
        +  "remove",
        +  "test"
        +]
  2. 15 tool updatesv1.2.0
    • First observedcheck_installed
    • First observedget_cve_info
    • First observedget_module_info
    • First observedget_profile_tools
    • First observedget_tool_info
    • First observedguided_assessment
    • First observedlist_profiles
    • First observedlist_tools
    • First observedmanage_remote_hosts
    • First observedrecommend_install
    • First observedrun_pipeline
    • First observedrun_script
    • First observedrun_tool
    • First observedsuggest_for_bounty
    • First observedsuggest_for_ctf

TDQS

A3.9/5.0

Scored across 15 tools

Disambiguation4/5

Most tools have distinct roles (registry discovery, install planning, advisory, execution, remote management), but there is some overlap between advisory/orchestration tools like guided_assessment, recommend_install, and suggest_for_* that could cause hesitation. Descriptions help clarify, so boundaries are mostly clear.

Naming Consistency4/5

Names are consistently snake_case and mostly follow a verb_noun pattern (get_tool_info, list_tools, run_tool). Minor deviations like guided_assessment (adjective_noun) and suggest_for_ctf (verb_preposition_noun) are readable but slightly break the pattern.

Tool Count5/5

15 tools is well within the ideal range and each tool serves a distinct purpose across discovery, installation, execution, and orchestration. No tool feels redundant or unnecessary for the toolkit's scope.

Completeness4/5

The surface covers registry discovery, profile/module insight, install recommendation, tool execution (single, pipeline, script), remote host management, and high-level orchestration. Minor gaps exist: no direct install/uninstall/update tool (though commands are surfaced via get_tool_info) and no explicit audit-log viewer.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables aggregation, filtering, transformation, and composition of tools from multiple MCP servers through a single proxy with tool views.
    5
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a security and context-control layer that multiplexes multiple MCP servers behind a single endpoint, scanning tool definitions and results, enforcing authorization, rate limiting, and audit logging, and dynamically retrieving tools to manage context window usage.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to run server-side security audits through callable tools that sweep attack surfaces, trace source-to-sink reachability, verify live findings, and manage scope, recon, chains, manuals, intelligence, and memory from any MCP client.
    3,832 npm
    512
    MIT