@darkmoon_ai/mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@darkmoon_ai/mcp-serverstart an authorized pentest against app.example.com and show me the run id"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@darkmoon_ai/mcp-server
A Model Context Protocol server that lets an MCP client (Claude Desktop, Goose, Continue, LibreChat, ...) drive Darkmoon, an open source (GPL-3.0) autonomous AI penetration testing platform.
Requires Darkmoon Pro
The Darkmoon engine and CLI are open source. This server talks to the Darkmoon Dashboard API, which is part of Darkmoon Pro and always self-hosted: there is no public hosted endpoint, so you supply the base URL of your own instance. It does not work against the open source CLI alone.
Related MCP server: Cockpit Lite MCP Server
Tools
Tool | Description |
| Start an autonomous pentest against one authorized target and return the |
| Report |
| List campaigns visible to the dashboard user (read only) |
| Vulnerabilities and severity statistics for a campaign (read only) |
Only run assessments against systems you own or are explicitly authorized in writing to test. Findings can include false positives and must be reviewed by a qualified human.
Configuration
Variable | Description |
| Base URL of your Darkmoon Pro Dashboard API (required) |
| Dashboard credentials; a JWT is requested on each call and never cached |
| Alternative to username/password: a pre-issued JWT |
| Optional per-request timeout, default 60000 |
Client configuration
Claude Desktop (claude_desktop_config.json), Continue and LibreChat use the same mcpServers shape:
{
"mcpServers": {
"darkmoon": {
"command": "npx",
"args": ["-y", "@darkmoon_ai/mcp-server"],
"env": {
"DARKMOON_BASE_URL": "https://darkmoon.example.internal",
"DARKMOON_USERNAME": "your-dashboard-user",
"DARKMOON_PASSWORD": "your-dashboard-password"
}
}
}
}Goose (~/.config/goose/config.yaml):
extensions:
darkmoon:
type: stdio
enabled: true
name: darkmoon
cmd: npx
args: ["-y", "@darkmoon_ai/mcp-server"]
envs:
DARKMOON_BASE_URL: https://darkmoon.example.internal
DARKMOON_USERNAME: your-dashboard-user
DARKMOON_PASSWORD: your-dashboard-passwordContinue (.continue/mcpServers/darkmoon.yaml):
name: Darkmoon
version: 0.1.0
schema: v1
mcpServers:
- name: darkmoon
command: npx
args: ["-y", "@darkmoon_ai/mcp-server"]
env:
DARKMOON_BASE_URL: https://darkmoon.example.internal
DARKMOON_USERNAME: your-dashboard-user
DARKMOON_PASSWORD: your-dashboard-passwordDevelop
npm install
npm run build
npm test # mocked Dashboard API, in-memory MCP client and a real stdio processLicense
GPL-3.0-only, same as Darkmoon.
Available Tools
4 toolsget_findingsGet campaign findingsARead-only
Return the vulnerabilities and aggregated severity statistics for a Darkmoon campaign (read only). Each finding carries title, severity, CVSS score, category, status (exploited, confirmed or unconfirmed), endpoint and remediation guidance. Findings may contain false positives and require human review.
| Name | Required | Description | Default |
|---|---|---|---|
| campaign_id | Yes | Darkmoon campaign id, e.g. camp_20260922_abc123 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=true), so the description's '(read only)' is redundant. However, it adds genuine context beyond the annotations: the shape of each finding (title, severity, CVSS, category, status, endpoint, remediation), the status enum values, and the caveat that findings may be false positives and need human review.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences: purpose first, then payload contents, then the reliability caveat. Every sentence earns its place; only the redundant '(read only)' could be trimmed since annotations already declare it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing return contents, and it does so well (fields, status values, severity statistics). The main remaining gap is that it does not describe pagination or result-size behavior for what could be a large findings list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and schema description coverage is 100% with a clear example format ('camp_20260922_abc123') in the schema itself. The description adds nothing about campaign_id semantics, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return the vulnerabilities and aggregated severity statistics for a Darkmoon campaign'), so an agent knows exactly what it retrieves. It does not explicitly contrast itself with siblings like list_campaigns or get_run_status, but the resource is distinct enough that differentiation is largely implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: call this to inspect the findings of a known campaign id. There is no explicit when-to-use vs alternatives guidance (e.g., when to reach for get_findings instead of get_run_status or list_campaigns), and no prerequisites stated beyond the required campaign_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_statusGet run statusARead-only
Report whether a Darkmoon run is 'running', 'completed', 'error' or 'unknown' (run log not found), with the event count and the 5 most recent events.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run_id returned by run_pentest |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no output schema, the description usefully discloses the exact return payload: the status vocabulary plus event count and 5 most recent events. It also explains the 'unknown' edge case as 'run log not found', which is a genuine failure-mode disclosure beyond the readOnly/openWorld annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, and every clause earns its place by specifying the return content and the edge-case state. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-param status-check tool with annotations covering safety, the description is nearly complete: it names the states and the events returned. The main gap is the polling/usage workflow and any freshness semantics, which the description omits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single run_id parameter is fully documented there as the value returned by run_pentest. The description adds no syntax, format, or sourcing detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Report") and resource (Darkmoon run status) and enumerates the four possible states the tool returns. It does not differentiate itself from siblings like run_pentest or get_findings, but the purpose is unambiguous on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives, nor any mention that it is typically used to poll after run_pentest finishes. The only hint of context ('run_id returned by run_pentest') lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_campaignsList campaignsARead-only
List the Darkmoon campaigns visible to the dashboard user, with ids and status. Use a campaign id with get_findings.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds real value on top: it discloses that results are filtered to the calling user's visibility (an authorization scope) and that the payload carries ids and status, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the purpose and scope front-loaded ahead of the workflow hint. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema in place, the description does the right thing by naming the returned fields (ids, status). It omits ordering, pagination, or truncation behavior, a minor gap for a listing tool but not one that blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema cannot be misinterpreted and there is nothing for the description to disambiguate. Baseline for a no-parameter tool is 4; it is not a 5 because there are no argument semantics to enrich.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (Darkmoon campaigns) plus the scope ('visible to the dashboard user') and what is returned (ids and status). The handoff to get_findings further anchors it against the sibling listing/retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives explicit downstream context: take a campaign id and feed it to get_findings. It does not state when NOT to use this tool (e.g. versus get_findings directly), so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_pentestStart a Darkmoon pentestA
Start an autonomous Darkmoon penetration test against one authorized target. The run executes in the background and can take a long time. Returns the run_id; poll it with get_run_status and read results with get_findings once a campaign exists (list_campaigns). Only use against systems the user owns or has explicit written authorization to test. Findings can include false positives and must be reviewed by a qualified human.
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | Optional focus areas, e.g. ['auth', 'injection'] | |
| target | Yes | Host, URL or scope to assess. Only targets you are authorized to test. | |
| program | No | Optional program name or rules-of-engagement note | |
| severity | No | Optional minimum severity to report |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the operation is non-read-only, open-world, non-idempotent and non-destructive, but the description adds material behavior beyond them: the run is asynchronous and long-running, it returns a run_id rather than findings, and results are noisy and require qualified human review. The authorization precondition and false-positive caveat are exactly the kind of disclosure an agent cannot infer from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences with zero redundancy; the core action and target scope come first, then the asynchronous lifecycle, then the safety caveat. Every sentence contributes a distinct fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async kickoff tool with no output schema, the description supplies everything needed to invoke and follow up correctly: what it returns (run_id), the polling/reading path, the runtime characteristics, and the authorization and verification requirements. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters (target, focus, program, severity). The description only reinforces the target constraint with 'one authorized target' and adds no format, syntax or defaulting detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start an autonomous Darkmoon penetration test') plus the scope constraint ('against one authorized target'). The lifecycle sentence implicitly separates it from siblings by assigning get_run_status to polling, get_findings to result reading, and list_campaigns to campaign discovery, so an agent can position this as the entry point without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternatives and the conditions that select them ('poll it with get_run_status', 'read results with get_findings once a campaign exists (list_campaigns)'). It also states an explicit when-not-to-use condition: only systems the user owns or has written authorization to test.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_findings - First observed
get_run_status - First observed
list_campaigns - First observed
run_pentest
TDQS
Scored across 4 tools
Each tool targets a distinct action (start, status, list campaigns, get findings), but the relationship between a 'run' and a 'campaign' is not fully clear, which could cause slight misselection between get_run_status and list_campaigns when checking progress.
All names follow a consistent snake_case verb_noun pattern (run_pentest, get_run_status, list_campaigns, get_findings), with clear verb prefixes that are easy to predict.
Four tools is well-scoped for a focused pentest service; each tool (start, monitor, list campaigns, read findings) earns its place without redundancy or bloat.
Core start-monitor-results workflow is covered, but notable gaps exist: no cancel/stop for long-running runs, and no explicit way to map a run_id to its resulting campaign, forcing agents to guess from list_campaigns.
Maintenance
Related MCP Connectors
MCP server for Pentest-Tools.com: run scans, manage findings and reports via your preffered LLM.
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
Authenticated MCP server for ClearPolicy policy and compliance workflows.
A paid remote MCP for developer endpoint scanner MCP, built to return verdicts, receipts, usage logs
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAutonomous pentests from one command: real security tools, working PoCs, and audit-ready reports, all driven via MCP.271 PyPI1,709MIT
- AlicenseNot gradedqualityCmaintenanceEnables authorized penetration testing through MCP, providing parallel reconnaissance, vulnerability scanning, attack path analysis, and self-contained HTML reporting with compliance tagging.MIT
- AlicenseNot gradedqualityBmaintenanceEnables authorized bug bounty automation via a scope-enforced MCP bridge, supporting web, secrets, mobile, and LLM red-team scanning, with reporting and advisory.MIT
- AlicenseCqualityBmaintenanceEnables authorized pentest and bug bounty workflows from any MCP client, with scoped recon, per-host rate limits, and scanner output turned into deduplicated, triaged finding cards.932 PyPI1MIT