HTB MCP Server
Provides live control of Hack The Box labs through the HTB Labs API, enabling AI agents to list, search and spawn machines, stop/reset/extend and submit flags, browse and submit challenges, manage VPN servers and download OVPN configs, view user profile, rank progression and recent activity, and call arbitrary HTB API endpoints directly.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@HTB MCP Serververify my HTB session, then spawn BoardLight and give me the IP"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
htb-mcp
MCP server that gives AI agents live control of Hack The Box labs — machines, challenges, flags and VPN — by wrapping the bundled, zero-dependency HTB Agent Toolkit CLI.
Unlike index-based servers, nothing here is cached content: every tool call drives the real HTB Labs API for your account. The server is a thin, safe adapter: it spawns toolkit/htb.py (Python 3.10+, standard library only) with argument arrays (no shell), returns the CLI's compact JSON as structured content, and maps its structured stderr errors (config, auth, rate_limit, conflict, state, ...) onto MCP errors.
Quick start
Requirements: Node.js 22.13 or newer and Python 3.10 or newer on PATH (set HTB_PYTHON if yours is elsewhere).
Claude Code:
claude mcp add htb -- npx -y @zebbern/htbCodex CLI:
codex mcp add htb -- npx -y @zebbern/htbAny MCP client (Claude Desktop, Cursor, Kimi, etc.), config JSON — HTB_API_KEY is optional if you instead create toolkit/.env.local inside the installed package:
{
"mcpServers": {
"htb": {
"command": "npx",
"args": ["-y", "@zebbern/htb"],
"env": { "HTB_API_KEY": "your-htb-app-token-here" }
}
}
}Create an App Token at https://app.hackthebox.com/profile/settings (App Tokens section). Then call htb_doctor once to verify the session.
As a plugin (bundles the agent skill that teaches safe, effective usage): this repo is a valid plugin for Claude Code (.claude-plugin/), Codex (.codex-plugin/) and Kimi (kimi-plugin/). Add it from your client's plugin marketplace flow pointing at zebbern/htb-mcp, or for Kimi Work use this plugin link.
Then ask things like:
"Check my HTB session and show which machine is active"
"Spawn BoardLight and give me the IP"
"Submit this flag for machine 444 as difficulty 50: HTB{...}"
"Switch my VPN to us-free-1 and download the config"
Related MCP server: CTFd MCP Server
Tools at a glance
Tool | What it does |
| Session health check: API key source, whoami, active machine, VPN assignment — run it first |
| List (playable/retired/todo/unreleased/SP-tier) and ranked search over machines |
| Machine details by id or name; what's running right now (live) |
| Spawn a box; |
| Lifecycle control of the active (or named) machine |
| Submit a user/root flag with a difficulty rating (10–100) |
| Browse and search challenges |
| Submit a challenge flag |
| List usable VPN servers ( |
| Switch servers; download OVPN configs |
| Profile, rank progression, recent owns |
| Escape hatch: call any HTB API endpoint directly (v4 by default, v5 via |
Read-only tools are annotated readOnlyHint; stop, reset and both submit tools are annotated destructiveHint. Full reference: docs/tools.md.
Documentation
docs/tools.md: complete tool reference with parameters and error types
docs/architecture.md: the wrapper design (process model, timeouts, error mapping)
docs/development.md: local setup, smoke test, releasing
docs/agents.md: agent skill, MCP registry and plugin packaging
Security
Token handling: the App Token is read by the bundled Python toolkit from
HTB_API_KEY(inherited from your MCP client config) ortoolkit/.env.local. The MCP server never reads, stores, logs or prints the key — it only passes the environment through to the child process.Never commit
.env.local. It is covered by.gitignore, and by both the root.npmignoreandtoolkit/.npmignoresonpm packcannot ship it; the release workflow hard-fails if a tarball ever contains it or the.cachedirectory.No shell injection: commands are built as argument arrays and run with
execFile,shell: false— machine names, flags and raw JSON bodies are passed verbatim as argv entries.State-changing tools are annotated: read-only vs destructive hints let MCP hosts gate
stop/reset/submit. Annotations don't replace consent — the bundled agent skill instructs agents to confirm before destructive or irreversible actions.This controls your real HTB account against live labs. Use it only for your own authorized HTB activity.
Credits
HTB Agent Toolkit — the bundled CLI (MIT), vendored under
toolkit/Hack The Box — this is an unofficial client of the public Labs API, not affiliated with or endorsed by Hack The Box Ltd
Available Tools
22 toolshtb_challenge_infoHTB Challenge InfoARead-onlyIdempotent
Show a challenge by id or name (names resolve case-insensitively). Returns an info block: id, name, category, difficulty, points, stars, retired, state, solved, release, description, solves, maker. Set full=true for the raw API payload.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Return the raw API payload instead of the curated info block | |
| target | Yes | Challenge id or name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description adds real behavioral detail on top: name resolution is case-insensitive, and full=true switches to the raw API payload rather than the curated block.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and lookup key, then the return contract. The field enumeration is long but earns its place given there is no output schema; the full=true note is placed last, where an optional flag belongs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates well by enumerating the returned info block fields. It does not cover what happens on an unknown or ambiguous name, which is the one remaining gap for a lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so baseline is 3, but the description adds semantics the schema lacks: case-insensitive name matching for target and the concrete effect of full=true (raw payload vs curated block), which is more explicit than the schema's phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (show) and resource (a challenge) with the lookup key (id or name), which cleanly separates it from the list/search siblings. It never explicitly names those siblings, so the differentiation is inferable rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the singular lookup semantics: use this when you already have an id or name. There is no explicit when-to-use/when-not guidance and no sibling (htb_challenge_list, htb_challenge_search) named as an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_challenge_listList HTB ChallengesARead-onlyIdempotent
List Hack The Box challenges (active by default; retired=true for retired ones). Returns compact JSON rows (id, name, category, difficulty, points, retired, state, solved, solves). Cached for 1h; noCache bypasses the cache.
| Name | Required | Description | Default |
|---|---|---|---|
| noCache | No | Bypass the local list cache for this call | |
| retired | No | List retired challenges instead of active ones |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so safety is covered. The description adds genuinely useful operational context the annotations lack: a 1-hour cache and that noCache bypasses it, plus the exact shape of returned rows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with purpose and scope, then caching behavior. Every clause carries information (default, override, cache policy, return fields) with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with full annotation coverage and no output schema, the description compensates well by enumerating the returned row fields and documenting the cache. Minor gaps remain, such as result-count limits or pagination behavior for large challenge sets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, and the description's restatement of retired/noCache semantics is largely redundant. Baseline 3 is appropriate given the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Hack The Box challenges') and immediately clarifies the default scope (active) versus the retired variant. It doesn't explicitly name the sibling it differs from (htb_challenge_search, htb_challenge_info), though the listing-versus-search distinction is reasonably inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells the agent the default behavior and how to get retired items, which implies usage, but there is no explicit when-to-use-this vs htb_challenge_search or htb_challenge_info guidance, and no exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_challenge_searchSearch HTB ChallengesARead-onlyIdempotent
Client-side search over the (cached) challenge lists, ranked by relevance. Matches id, name, category, difficulty, points, state. Use all to also scan retired challenges. Follow up with htb_challenge_info on a candidate.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | Also scan retired challenges | |
| limit | No | Max results (default 20) | |
| query | Yes | Search terms, e.g. 'crypto' or 'web easy' | |
| retired | No | Search only retired challenges |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds meaningful behavior beyond that: search is client-side over a cached list, results are relevance-ranked, and matches span six fields. This is a real behavioral disclosure, not a restatement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core operation and mechanism, then parameters, then the follow-up workflow. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with full schema coverage and no output schema, this covers mechanism, scope, matching fields, and the next step. It stops short of explaining result shape or ranking tie-breaking, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description exceeds it by enumerating which fields the query matches (id, name, category, difficulty, points, state) and clarifying the `all` flag's effect on retired challenges, information the schema does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (challenge lists) with scope qualifiers (cached, ranked by relevance). It clearly differentiates from siblings htb_challenge_list and htb_challenge_info by naming info as the follow-up step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit guidance on the `all` flag and a clear next-step workflow ('Follow up with htb_challenge_info on a candidate'). It does not, however, contrast itself against htb_challenge_list, so the agent must infer the split between listing and searching.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_challenge_submitSubmit Challenge FlagADestructive
Submit a challenge flag. difficulty is a required rating in steps of 10 from 10 to 100 — ask the user if they did not provide one. Submissions are irreversible: confirm with the user first.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | The flag, e.g. HTB{...} | |
| target | Yes | Challenge id or name | |
| difficulty | Yes | Difficulty rating 10-100 in steps of 10 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and readOnlyHint=false, but the description adds material context beyond them: submissions are irreversible and a user confirmation should precede the call. That is genuine added behavioral value, though it does not explain what a failed/invalid flag submission does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the difficulty requirement, then the irreversibility warning. Every sentence carries an instruction the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-param mutation with no output schema and annotations covering the safety profile, the description supplies the key missing operational detail (irreversibility plus confirmation). It omits only minor edges such as flag-format validation or what a rejected submission returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the flag/target/difficulty semantics and the 10-100 step-of-10 constraint are already documented in the schema. The description reinforces the difficulty requirement and the ask-the-user fallback but adds no format details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Submit a challenge flag'), which cleanly separates it from htb_machine_submit by resource. It stops short of explicitly naming the sibling alternative, so it is clear but not fully differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operational guidance: prompt the user for difficulty if omitted, and confirm before submitting because the action is irreversible. It does not describe when this tool is preferred over other challenge tools (e.g. after htb_challenge_info), but the invocation conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_doctorHTB Session Health CheckARead-onlyIdempotent
Session health check: verifies API key resolution (and where the key came from), API connectivity (whoami), the currently active machine, and the VPN assignment in one call. Run this first in any session; exits with an error only when the key or API check fails. Needs HTB_API_KEY (env var or toolkit/.env.local).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds value beyond them by disclosing a prerequisite (HTB_API_KEY via env var or toolkit/.env.local) and failure semantics (errors only when the key or API check fails), though it does not describe the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then when to run it, then the prerequisite. Every sentence earns its place with no filler, though the parenthetical about key origin adds slight density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic with no output schema, the description covers what is checked, when to run it, prerequisites, and failure conditions. An agent has everything needed to call it correctly; only the exact response format is left implicit, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to compensate for, and it correctly implies a no-argument invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (health check / verifies) and enumerates exactly what is checked: API key resolution and source, API connectivity via whoami, active machine, and VPN assignment. This clearly distinguishes it from action-oriented siblings like htb_machine_start or htb_vpn_status, which each cover a single concern.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Run this first in any session" gives explicit, actionable timing guidance that no sibling provides. There is no when-not-to-use or alternative named, but no sibling competes for the same diagnostic role, so the guidance is nearly complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_activeActive HTB MachineARead-onlyIdempotent
Show the currently active machine: id, name, ip, os, difficulty, expires_at/expires_in, lab server. Always live, never cached. An empty info block means nothing is running. details=true adds the synopsis and linked academy modules.
| Name | Required | Description | Default |
|---|---|---|---|
| details | No | Add synopsis and academy modules to the response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds real behavioral context beyond that: the response is always live and never cached, and an empty info block is a meaningful 'nothing running' signal rather than an error. It doesn't cover latency or rate-limit behavior, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: identity and return shape first, then the freshness guarantee and empty-result semantics, then the optional flag. No filler, and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing the return payload, and it does so by enumerating the fields returned. It also covers the empty case and the details flag, leaving nothing an agent needs in order to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'details' parameter, whose schema text already says 'Add synopsis and academy modules to the response'. The description's 'details=true adds the synopsis and linked academy modules' is essentially a restatement, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show the currently active machine') and immediately enumerates the returned fields (id, name, ip, os, difficulty, expires_at/expires_in, lab server). The scope ('currently active') cleanly separates it from siblings like htb_machine_list, htb_machine_search and htb_machine_profile without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Always live, never cached' tells the agent this is the authoritative source for current runtime state, implicitly steering it here over a cached listing. 'An empty info block means nothing is running' gives a concrete interpretation rule. It stops short of naming an explicit alternative or when-not-to-use condition, so it isn't a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_extendExtend HTB MachineA
Extend a machine's expiry time; defaults to the currently active machine when target is omitted. Non-destructive: the running box and its state are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Machine id or name; omit to extend the active machine |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false, and the description largely restates this as 'Non-destructive,' adding the mild detail that the running box and its state are untouched. It does not cover auth requirements, rate limits, or what happens to a machine that is not running, so the added value over the annotations is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded clauses with no filler; the core action comes first and the default-target rule follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with annotations covering the safety profile and no output schema, the description supplies the essential default-target behavior and safety note. It is nearly complete, missing only edge-case behavior (e.g., what happens with an invalid or expired machine).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is already documented in the schema, including the 'omit to extend the active machine' default. The description repeats that default behavior rather than adding format or constraint detail, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (extend a machine's expiry time) that clearly distinguishes it from siblings like htb_machine_stop, htb_machine_reset, or htb_machine_start. It does not explicitly name those siblings, so it falls short of the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete usage rule — omit target to act on the active machine — but says nothing about when to extend versus reset, restart, or stop a machine. Usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_listList HTB MachinesARead-onlyIdempotent
List Hack The Box machines. Default is the playable list (page 1); use filter for retired, todo, or unreleased lists, and spTier for Seasonal/starting-point tier lists. Returns compact JSON rows (id, name, os, difficulty, points, active, spawned, free). Lists are cached for 1h; noCache bypasses the cache.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page number (playable and retired lists only) | |
| filter | No | Which list to fetch (default playable) | playable |
| spTier | No | Starting-point tier 1-3 (overrides filter) | |
| noCache | No | Bypass the local list cache for this call |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the description earns credit for adding real context beyond them: the 1h list cache and that noCache bypasses it. It also discloses the returned row shape, which matters because there is no output schema. It omits any auth/rate-limit or pagination-size details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences: purpose and default first, parameter routing second, return shape and caching last. Every clause carries information and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates returned fields (id, name, os, difficulty, points, active, spawned, free) and explains caching, making the tool callable without surprises. It does not say how many rows a page returns or the total page count, a minor remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (page, filter, spTier, noCache) are already documented in the schema. The description restates the filter/spTier behavior and the page-1 default without adding syntax or format detail beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List Hack The Box machines') and immediately clarifies the default scope (playable list, page 1). It does not explicitly differentiate itself from the sibling htb_machine_search, so it falls short of a 5, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear parameter-selection guidance: 'use filter for retired, todo, or unreleased lists, and spTier for Seasonal/starting-point tier lists.' This tells the agent which argument selects which list. It stops short of a 5 because it never states when to prefer this over htb_machine_search or htb_machine_active.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_profileHTB Machine ProfileARead-onlyIdempotent
Show a machine profile by id or name (e.g. 444 or 'BoardLight'). Returns an info block: id, name, os, difficulty, points, stars, retired, free, release, maker, user/system owns, tags. Set full=true for the complete raw API payload.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Return the raw API payload instead of the curated info block | |
| target | Yes | Machine id or name (e.g. '444' or 'BoardLight') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive behaviour, so the safety profile is covered. The description adds real value beyond that by disclosing the default curated output shape and the full=true switch to the raw payload, though it says nothing about error behaviour when a name is ambiguous or missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and identifier examples before the return-value list. The field enumeration is slightly long but earns its place given no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the info block fields and explains the full flag, which is enough for correct invocation. It stops short of covering lookup-failure or ambiguous-name behaviour, but the core is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters are documented there, including the example id/name form and the full=true meaning. The description largely restates those schema descriptions rather than adding new semantics (no format, no ambiguity or case-sensitivity rules), so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Show a machine profile') with two accepted identifier forms, and enumerates the returned fields. An agent can distinguish it from htb_machine_list, htb_machine_search, and htb_machine_active without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the id/name lookup phrasing, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to siblings like htb_machine_search when the user only has a partial name. The agent must infer the lookup context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_resetReset HTB MachineADestructive
Reset a machine to a fresh state; defaults to the currently active machine when target is omitted. Wipes the box back to its initial state — confirm with the user before calling.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Machine id or name; omit to reset the active machine |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds real value on top: it spells out what destruction means ('wipes the box back to its initial state') and imposes a user-confirmation requirement, though it says nothing about duration, quotas, or whether the reset can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences: the operation and default target come first, then the destructive warning. Nothing is padded and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive action with no output schema, the description covers the operation, the default target, and the confirmation requirement. It omits only minor operational details such as expected duration or post-reset state, which are not essential to invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter's own description already states 'omit to reset the active machine.' The description's default-target sentence largely repeats that, adding no new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reset a machine') plus the resulting scope ('to a fresh state'), which is unambiguous and easily separable from siblings like htb_machine_start, htb_machine_stop, or htb_machine_extend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the default behavior when 'target' is omitted and adds a clear pre-call condition ('confirm with the user before calling'). It does not name an alternative tool for cases where a reset is not what the agent wants, but the context for use is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_searchSearch HTB MachinesARead-onlyIdempotent
Client-side search over the HTB machine lists, ranked by relevance. Matches id, name, OS, difficulty, tags, makers, points, IP. Use all to also scan retired machines. profiles additionally fetches every scanned machine's profile to match description text — powerful but SLOW (one API request per machine, can take minutes), use it only when list fields are not enough. Follow up with htb_machine_profile on a candidate.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | Also scan retired machines | |
| limit | No | Max results (default 20) | |
| query | Yes | Search terms, e.g. 'kerberos', 'windows easy', or a machine name | |
| retired | No | Search only retired machines | |
| maxPages | No | Max list pages to scan (default 10) | |
| profiles | No | Also match machine description text (slow: one request per machine) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, yet the description still adds real behavioral context: search is client-side, results are relevance-ranked, and `profiles` triggers one API request per machine that 'can take minutes'. It stops short of describing result shape or pagination behavior, but the cost disclosure is genuinely useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what it does, then the two flags that change cost, then the follow-up action. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description tells the agent how to proceed after a hit (htb_machine_profile) and what the search surface is. It does not explain result ordering details or what happens when the scan hits maxPages, which are minor gaps for a read-only search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema for `all` (scan retired machines) and `profiles` (matches description text, and why it is slow). limit/maxPages/retired are left entirely to the schema, which is acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Client-side search over the HTB machine lists, ranked by relevance') and enumerates exactly which fields are matched, which cleanly distinguishes it from the sibling htb_machine_list. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance for the two expensive flags ('Use `all` to also scan retired machines', 'use it only when list fields are not enough'), warns about cost, and routes the agent to the next step ('Follow up with htb_machine_profile on a candidate'). Alternatives and exclusions are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_startStart HTB MachineA
Start (spawn) a machine by id or name. mode 'auto' (default) tries play then falls back to spawn. With wait=true (default) the CLI retries while spawn capacity is full, then waits for the IP — this can take several minutes at peak times; the response contains {id, name, ip, spawn}. Only one machine can be active at a time: a conflict error names the blocker, stop it (with the user's consent) and retry.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Start mode (default auto: play, then spawn) | |
| wait | No | Retry a full spawn server and wait for the machine IP (default true) | |
| target | Yes | Machine id or name to start | |
| interval | No | Seconds between retries/polls (default 15) | |
| retryFor | No | Seconds to keep retrying a full spawn server (default 600) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the read/write/open-world profile; the description adds the operational traits that matter: retries while spawn capacity is full, potentially several minutes of latency at peak times, the response shape {id, name, ip, spawn}, and the single-active-machine conflict constraint with a consent requirement before stopping the blocker.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and default mode, then the latency/retry caveat, then the response shape, then the conflict constraint. Dense and mostly waste-free, though the mode restatement overlaps with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description supplies the return shape ({id, name, ip, spawn}) along with the retry/latency behavior an agent needs to set expectations. Combined with the annotations, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so mode, wait, interval, and retryFor defaults are already documented in the schema. The description restates the auto and wait semantics but adds no format or edge-case detail beyond them (e.g. it never mentions interval/retryFor), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start (spawn) a machine by id or name') and immediately distinguishes the spawn/play modes, which separates it cleanly from siblings like htb_machine_stop, htb_machine_reset, and htb_machine_extend. An agent can identify the tool's action without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains the mode default ('auto' tries play then falls back to spawn) and the wait behavior, plus the concrete conflict path: only one machine active, a conflict error names the blocker, stop it with user consent and retry. This is genuine when-to-use and what-to-do-next guidance rather than a restatement of the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_stopStop HTB MachineADestructive
Stop (terminate) a machine; defaults to the currently active machine when target is omitted. This kills the running box and its state — confirm with the user before calling.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Machine id or name; omit to stop the active machine |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds real substance beyond that: it specifies what is destroyed ('kills the running box and its state') and mandates user confirmation. It stops short of describing post-stop state or error behavior if no active machine exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero filler, with the destructive nature front-loaded so the agent sees the risk before the mechanics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the destructive effect, the parameter default, and the safety gate — adequate for a one-parameter mutation tool with no output schema. Minor gap: no mention of failure behavior when no active machine exists or a target is invalid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter's default behavior is already documented in the schema ('omit to stop the active machine'). The description restates that default rather than adding syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Stop (terminate) a machine') with a clarifying synonym, and is unambiguously distinct from siblings like htb_machine_start, htb_machine_reset, and htb_machine_extend. An agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the condition for the omitted parameter ('defaults to the currently active machine when target is omitted') and a clear pre-call precondition ('confirm with the user before calling'). It does not explicitly name alternative tools, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_machine_submitSubmit Machine FlagADestructive
Submit a user or root flag for a machine (HTB infers which from the flag value). difficulty is a required rating in steps of 10 from 10 (piece of cake) to 100 (brainfuck) — ask the user if they did not provide one. Submissions are irreversible: confirm with the user first.
| Name | Required | Description | Default |
|---|---|---|---|
| flag | Yes | The flag, e.g. HTB{...} | |
| target | Yes | Machine id or name | |
| difficulty | Yes | Difficulty rating 10-100 in steps of 10 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds genuinely new context: the operation is irreversible, user confirmation is required, and the difficulty value must be solicited from the user. It does not describe error behavior or what a successful submission returns, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct information: the action, the difficulty constraint plus user-prompt rule, and the irreversibility warning. Nothing is repeated from the schema and the critical warning is not buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation with full schema coverage and annotations covering the safety profile, nothing essential is missing. There is no output schema to explain, and the description closes the main behavioral gap (irreversibility and confirmation) on its own.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains the 10-100 step-10 difficulty scale with human labels and clarifies that the flag value itself determines whether a user or root flag is being submitted. That is more than the schema's bare 'difficulty rating 10-100' text conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Submit a user or root flag for a machine') and notes that the flag type is inferred, which separates it from htb_challenge_submit. An agent can identify the operation and its target without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operational guidance: ask the user for a difficulty if one wasn't provided, and confirm before submitting because submissions are irreversible. It stops short of naming an alternative tool (e.g. htb_challenge_submit for challenges), so it is clear context rather than full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_rawRaw HTB API CallA
Escape hatch: call any HTB API endpoint directly. Hits the v4 base (https://labs.hackthebox.com/api/v4) by default — pass baseUrl 'https://labs.hackthebox.com/api/v5' for v5 endpoints. path like '/machine/active' (leading slash optional); data takes a raw JSON string body. Set output to save a binary response (for example, GET /challenge/download/) instead of decoding it as text. Relative output paths resolve under the toolkit directory. Use when a typed tool is missing or an endpoint moved.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | Raw JSON request body string | |
| path | Yes | API path, e.g. '/machine/active' | |
| method | Yes | HTTP method | |
| output | No | Save the response body to this file path | |
| baseUrl | No | Override API base URL (use the v5 URL for v5 endpoints) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true; the description adds real behavioral context beyond that — default v4 base URL, v5 override, leading-slash tolerance, raw JSON body format, binary-output saving, and relative-path resolution under the toolkit directory. It stops short of stating auth requirements or error/status handling for a tool that can issue DELETE/POST, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the escape-hatch purpose and the routing condition, then packs concrete usage details. It is dense and multi-clause, but each sentence carries actionable information; the baseUrl and output sentences in particular are not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, untyped passthrough with no output schema, the description covers base URL selection, path/data formats, and binary-output behavior — enough to invoke it correctly. Remaining gaps (auth expectations, behavior on non-2xx responses) are real but secondary given the openWorld annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description genuinely adds meaning: baseUrl is explained as the v5 override, data as a raw JSON string, output as a binary-save target with relative paths resolved under the toolkit directory, and path's optional leading slash. These are semantics not present in the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('call any HTB API endpoint directly') and frames itself as an 'escape hatch', which cleanly separates it from the typed siblings like htb_machine_list or htb_challenge_info. An agent can identify its role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: 'Use when a typed tool is missing or an endpoint moved.' That is precisely the when-to-use condition that distinguishes this fallback from the ~21 typed siblings, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_user_activityHTB User ActivityARead-onlyIdempotent
Show recent owns (machine user/root flags and challenges), newest first. Rows: date, type (user/root/challenge), id, name, points, blood.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max entries (default 20) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive). The description adds genuinely useful behavior: results are ordered newest-first, and it enumerates the returned row fields, which matters since no output schema exists. It says nothing about the 100-entry cap or rate limits, but that is minor here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the operation and its output shape are front-loaded, and the second sentence earns its place by describing columns absent from any other structured field.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description supplies the missing return shape (date, type, id, name, points, blood) and sort order. Only the limit/pagination behavior is left to the schema, which is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'limit' parameter is fully documented there with default and bounds. The description adds no paging guidance, so it sits at the baseline for a schema that does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Show") plus a well-defined resource ("recent owns" = machine user/root flags and challenges), with the parenthetical clarifying the domain-specific term 'owns'. An agent can distinguish it from htb_user_progress or htb_user_info, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to call this versus htb_user_progress or htb_user_info, and no prerequisites noted. The 'recent owns, newest first' framing implies a feed-style usage, but the agent must infer it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_user_infoHTB User ProfileARead-onlyIdempotent
Show the user's HTB profile: id, name, rank, points, ranking, user/system owns, respects, country, team, vip. Set full=true for the raw profile payload.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Return the raw API payload instead of the curated info block |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is fully covered without the description. The description adds one genuinely useful behavioral fact beyond that: by default it returns a curated block rather than the raw payload. It says nothing about auth requirements, rate limits, or failure modes, so it clears the (lower) annotated bar but does not exceed it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the resource and field list front-loaded and the flag note last; nothing is wasted. The 11-field enumeration is slightly list-heavy, but it substitutes for a missing output schema and so earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, listing the curated fields is exactly what the description should do, and the full=true escape hatch is explained. Safety is covered by annotations, so for a zero-required-parameter read tool the definition is close to complete; only the sibling boundary is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'full' parameter is already documented in the schema as returning the raw API payload. The description restates the same semantics rather than adding format, size, or cost implications, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Show') and resource ('the user's HTB profile') and enumerates the returned fields, so an agent immediately knows this is a read of the caller's own profile. It does not, however, explicitly distinguish itself from near-siblings like htb_user_progress or htb_user_activity, which an agent must infer from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the field list and the 'Set full=true for the raw profile payload' note hint at when the raw variant is wanted, but there is no explicit when-to-use guidance or any pointer to alternative sibling tools. An agent can call it correctly, but must guess the boundary against htb_user_progress/htb_user_activity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_user_progressHTB Rank ProgressARead-onlyIdempotent
Show rank progression: current rank and points, next rank and its points requirement, current_rank_progress / rank_requirement (percent), rank_ownership, and own counts. Use this to answer 'how close am I to ranking up?'.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds value by disclosing the concrete content of the response (rank, points, percent progress, ownership), which matters because no output schema exists. It does not mention authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the field inventory and closing with the usage trigger. The enumerated field list is dense but earns its place given there is no output schema to document returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description correctly shoulders the burden of describing return content, which it does thoroughly. Only auth/permission expectations are left implicit, which is a minor gap for a read-only, open-world tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('rank progression') and enumerates the concrete fields returned (current rank/points, next rank, rank_requirement, rank_ownership). It is clearly narrower than htb_user_info or htb_user_activity, though it never names those siblings explicitly to draw the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage trigger: 'Use this to answer "how close am I to ranking up?"', which tells an agent exactly when to reach for it. It stops short of stating when NOT to use it or naming htb_user_info as the alternative for general profile data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_vpn_downloadDownload HTB VPN ConfigARead-onlyIdempotent
Download an OVPN config file for a server (id, alias, or live name). Default output is lab-vpn.ovpn in the toolkit directory; variant 0 = UDP. The toolkit does NOT run OpenVPN — the user connects themselves with the downloaded file.
| Name | Required | Description | Default |
|---|---|---|---|
| output | No | Output file path (default lab-vpn.ovpn) | |
| server | Yes | Server id, alias like 'us-free-1', or live name | |
| variant | No | Config variant (default 0 = UDP) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld, so the key addition is behavioral: the call writes a file to a default path and the user must connect themselves. That 'does NOT run OpenVPN' note is exactly the kind of side-effect clarification annotations do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the download action, then defaults, then the critical caveat about not launching OpenVPN. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains where the file lands and what variant means, which is enough for a simple download tool. Lacks any note on error behavior for an invalid server identifier, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all three params are documented in-schema, so the baseline is 3. The description restates the defaults (lab-vpn.ovpn, variant 0 = UDP) without adding format or syntax detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Download an OVPN config file') and names the three accepted server identifier forms (id, alias, live name). An agent can distinguish this from htb_vpn_servers (listing) and htb_vpn_switch (switching) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful context about the default output location and that the toolkit does not run OpenVPN, but never says when to use this versus htb_vpn_servers (to discover a server) or htb_vpn_switch. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_vpn_serversList HTB VPN ServersARead-onlyIdempotent
List the VPN servers the account can use, live (includes VIP/VIP+/dedicated pools), assigned server first. Rows: id, name, group, location, clients, full, assigned. Default product pool is 'labs'. static=true shows only the built-in offline alias table and needs NO API key.
| Name | Required | Description | Default |
|---|---|---|---|
| static | No | Show only the offline alias table (works without an API key) | |
| product | No | Product pool to list (default labs) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the safety profile is covered. The description adds genuine behavioral context beyond them: the results are live, VIP/VIP+/dedicated pools are included, the assigned server is ordered first, and the static mode requires no API key (an auth requirement). It stops short of covering pagination or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose, then the returned columns, then the parameter caveats. The compact 'Rows: id, name, group, location, clients, full, assigned' uses no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden and does so by enumerating the row fields. Combined with the default pool, static-mode auth note, and result ordering, an agent has everything needed to call this 2-param read tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters (static, product enum) are already documented with defaults and enum values in the schema. The description's 'needs NO API key' and product-pool defaults only lightly extend the schema, so baseline 3 is appropriate for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb + resource + scope: 'List the VPN servers the account can use', and further narrows it with 'live (includes VIP/VIP+/dedicated pools)'. This is clearly distinguishable from sibling tools like htb_vpn_status, htb_vpn_switch, and htb_vpn_download, and it even names the columns returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the 'static=true' mode condition ('offline alias table... needs NO API key') and the default pool, which implies when each mode applies. However, it never states when to choose this tool over sibling tools such as htb_vpn_status or htb_vpn_download, so usage routing is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_vpn_statusHTB VPN StatusARead-onlyIdempotent
Show the VPN server currently assigned to the account: id, name, location, product — or assigned=false when none.
| Name | Required | Description | Default |
|---|---|---|---|
| product | No | Product pool, e.g. 'labs' (default) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds that the result may be assigned=false when no server exists, which is genuinely useful negative-case behavior. It does not address auth requirements or rate limits, so against a lowered annotation bar this is a moderate 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence front-loading the verb and resource, then listing return fields. No filler, and the assigned=false fallback is appended efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only status tool with a one-param optional schema, the description covers purpose and the empty-result case. It doesn't clarify the product parameter behavior or distinguish explicitly from htb_vpn_servers, but an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'product' is fully documented in the schema (100% coverage) with an example ('labs') and default. The description doesn't mention the product parameter at all, offering no meaning beyond the schema. Baseline 3 applies when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Show) and resource (VPN server currently assigned to the account), and enumerates the returned fields (id, name, location, product), which clearly distinguishes it from htb_vpn_servers (list all available) and htb_vpn_switch (change assignment).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through 'currently assigned' but never explicitly says when to use this versus htb_vpn_servers or htb_vpn_switch. The distinction is inferable from the sibling names, but no explicit when/when-not guidance or alternative is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
htb_vpn_switchSwitch HTB VPN ServerA
Switch the account to a different VPN server. Accepts a numeric id (289), a static alias (us-free-1, eu-sp-1, ...) or a live name from htb_vpn_servers (e.g. 'EU Machines VIP+ 1'). Affects the user's whole HTB connection — mention it before switching.
| Name | Required | Description | Default |
|---|---|---|---|
| server | Yes | Server id, alias like 'us-free-1', or live name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds genuinely new behavioral context beyond that: the change affects the user's whole HTB connection and should be surfaced to the user beforehand. It still omits what happens on failure or whether an active session is dropped.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then input formats, then the impact warning. Every sentence carries information an agent needs and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema and annotations covering the safety profile, the description is nearly complete: it explains accepted inputs and the blast radius of the change. Only the outcome of the call (success/failure, reconnection behavior) is unaddressed, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description enriches the single parameter by enumerating the three accepted input forms (numeric id, static alias, live name) with concrete examples and naming htb_vpn_servers as the authoritative source for live names, which the schema description only gestures at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Switch the account to a different VPN server') and clearly separates itself from the sibling htb_vpn_servers/htb_vpn_status tools by naming the server-listing tool as a source of valid values.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context ('mention it before switching') and points to htb_vpn_servers as the place to obtain live names, so the agent knows how to source input. It does not state when-not to use this tool or name explicit alternatives, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
v0.1.1- First observed
htb_challenge_info - First observed
htb_challenge_list - First observed
htb_challenge_search - First observed
htb_challenge_submit - First observed
htb_doctor - First observed
htb_machine_active - First observed
htb_machine_extend - First observed
htb_machine_list - First observed
htb_machine_profile - First observed
htb_machine_reset - First observed
htb_machine_search - First observed
htb_machine_start - First observed
htb_machine_stop - First observed
htb_machine_submit - First observed
htb_raw - First observed
htb_user_activity - First observed
htb_user_info - First observed
htb_user_progress - First observed
htb_vpn_download - First observed
htb_vpn_servers - First observed
htb_vpn_status - First observed
htb_vpn_switch
TDQS
Scored across 22 tools
Most tools target a clearly distinct resource+action (machine_start vs machine_stop vs machine_reset vs machine_extend, vpn_servers vs vpn_status vs vpn_switch vs vpn_download). A few near-neighbors exist (machine_list vs machine_search, challenge_list vs challenge_search, machine_active vs machine_profile), but the descriptions explicitly differentiate scope and use cases, so an agent can reliably pick the right one.
Every tool uses the same `htb_<resource>_<action>` snake_case pattern (htb_machine_start, htb_challenge_submit, htb_vpn_switch, htb_user_info). The two outliers (htb_doctor, htb_raw) still follow the consistent lowercase htb_ prefix and read as intentional singletons.
22 tools is on the heavy side, but the server spans genuinely distinct sub-domains (machines, challenges, VPN, user profile, session health, raw escape hatch), and each tool covers a real operation. There is little obvious redundancy, so the count is justified by the broad surface rather than bloated.
The set covers the full lifecycle: machine list/search/profile/start/stop/reset/extend/submit, challenge list/search/info/submit, VPN list/status/switch/download, user info/progress/activity, plus a health check and htb_raw escape hatch for anything missing. No obvious dead ends for the HTB domain.
Maintenance
Related MCP Connectors
Real Linux labs your AI agent deploys, routes and runs, with domains, TLS, DBs and an audit log.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
AI pentesting: run scans, triage vulnerabilities, review PRs, manage schedules and assets.
Verified, pay-per-use API tools for AI agents through one authenticated connection.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceEnables AI assistants to execute penetration testing commands and security tools on Kali Linux remotely. Supports automated reconnaissance, vulnerability scanning, and CTF solving through integration with 25+ offensive security tools like nmap, gobuster, and nuclei.16-
- AlicenseBqualityCmaintenanceEnables interaction with CTFd platforms for Capture The Flag competitions, allowing users to list challenges, read details, manage dynamic Docker containers, and submit flags through natural language.538 PyPI24Apache 2.0
- FlicenseNot gradedqualityFmaintenanceEnables AI agents to interact with Exegol pentesting containers to execute commands and manage container status. It includes seven predefined workflows for automated security tasks such as web reconnaissance, port scanning, and vulnerability assessment.2-
- AlicenseCqualityCmaintenanceEnables interaction with Hack The Box App services including machines, challenges, sherlocks, and more through the HTB API v4.601MIT