Skip to main content
Glama
ASCIT31

@darkmoon_ai/mcp-server

by ASCIT31

@darkmoon_ai/mcp-server

A Model Context Protocol server that lets an MCP client (Claude Desktop, Goose, Continue, LibreChat, ...) drive Darkmoon, an open source (GPL-3.0) autonomous AI penetration testing platform.

Requires Darkmoon Pro

The Darkmoon engine and CLI are open source. This server talks to the Darkmoon Dashboard API, which is part of Darkmoon Pro and always self-hosted: there is no public hosted endpoint, so you supply the base URL of your own instance. It does not work against the open source CLI alone.

Related MCP server: Cockpit Lite MCP Server

Tools

Tool

Description

run_pentest

Start an autonomous pentest against one authorized target and return the run_id

get_run_status

Report running, completed, error or unknown for a run, from its run log

list_campaigns

List campaigns visible to the dashboard user (read only)

get_findings

Vulnerabilities and severity statistics for a campaign (read only)

Only run assessments against systems you own or are explicitly authorized in writing to test. Findings can include false positives and must be reviewed by a qualified human.

Configuration

Variable

Description

DARKMOON_BASE_URL

Base URL of your Darkmoon Pro Dashboard API (required)

DARKMOON_USERNAME, DARKMOON_PASSWORD

Dashboard credentials; a JWT is requested on each call and never cached

DARKMOON_TOKEN

Alternative to username/password: a pre-issued JWT

DARKMOON_TIMEOUT_MS

Optional per-request timeout, default 60000

Client configuration

Claude Desktop (claude_desktop_config.json), Continue and LibreChat use the same mcpServers shape:

{
  "mcpServers": {
    "darkmoon": {
      "command": "npx",
      "args": ["-y", "@darkmoon_ai/mcp-server"],
      "env": {
        "DARKMOON_BASE_URL": "https://darkmoon.example.internal",
        "DARKMOON_USERNAME": "your-dashboard-user",
        "DARKMOON_PASSWORD": "your-dashboard-password"
      }
    }
  }
}

Goose (~/.config/goose/config.yaml):

extensions:
  darkmoon:
    type: stdio
    enabled: true
    name: darkmoon
    cmd: npx
    args: ["-y", "@darkmoon_ai/mcp-server"]
    envs:
      DARKMOON_BASE_URL: https://darkmoon.example.internal
      DARKMOON_USERNAME: your-dashboard-user
      DARKMOON_PASSWORD: your-dashboard-password

Continue (.continue/mcpServers/darkmoon.yaml):

name: Darkmoon
version: 0.1.0
schema: v1
mcpServers:
  - name: darkmoon
    command: npx
    args: ["-y", "@darkmoon_ai/mcp-server"]
    env:
      DARKMOON_BASE_URL: https://darkmoon.example.internal
      DARKMOON_USERNAME: your-dashboard-user
      DARKMOON_PASSWORD: your-dashboard-password

Develop

npm install
npm run build
npm test      # mocked Dashboard API, in-memory MCP client and a real stdio process

License

GPL-3.0-only, same as Darkmoon.

Available Tools

4 tools
get_findingsGet campaign findingsA
Read-only

Return the vulnerabilities and aggregated severity statistics for a Darkmoon campaign (read only). Each finding carries title, severity, CVSS score, category, status (exploited, confirmed or unconfirmed), endpoint and remediation guidance. Findings may contain false positives and require human review.

ParametersJSON Schema
NameRequiredDescriptionDefault
campaign_idYesDarkmoon campaign id, e.g. camp_20260922_abc123

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=true), so the description's '(read only)' is redundant. However, it adds genuine context beyond the annotations: the shape of each finding (title, severity, CVSS, category, status, endpoint, remediation), the status enum values, and the caveat that findings may be false positives and need human review.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, front-loaded sentences: purpose first, then payload contents, then the reliability caveat. Every sentence earns its place; only the redundant '(read only)' could be trimmed since annotations already declare it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing return contents, and it does so well (fields, status values, severity statistics). The main remaining gap is that it does not describe pagination or result-size behavior for what could be a large findings list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, and schema description coverage is 100% with a clear example format ('camp_20260922_abc123') in the schema itself. The description adds nothing about campaign_id semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the vulnerabilities and aggregated severity statistics for a Darkmoon campaign'), so an agent knows exactly what it retrieves. It does not explicitly contrast itself with siblings like list_campaigns or get_run_status, but the resource is distinct enough that differentiation is largely implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this to inspect the findings of a known campaign id. There is no explicit when-to-use vs alternatives guidance (e.g., when to reach for get_findings instead of get_run_status or list_campaigns), and no prerequisites stated beyond the required campaign_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_statusGet run statusA
Read-only

Report whether a Darkmoon run is 'running', 'completed', 'error' or 'unknown' (run log not found), with the event count and the 5 most recent events.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesThe run_id returned by run_pentest

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no output schema, the description usefully discloses the exact return payload: the status vocabulary plus event count and 5 most recent events. It also explains the 'unknown' edge case as 'run log not found', which is a genuine failure-mode disclosure beyond the readOnly/openWorld annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, and every clause earns its place by specifying the return content and the edge-case state. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-param status-check tool with annotations covering safety, the description is nearly complete: it names the states and the events returned. The main gap is the polling/usage workflow and any freshness semantics, which the description omits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single run_id parameter is fully documented there as the value returned by run_pentest. The description adds no syntax, format, or sourcing detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Report") and resource (Darkmoon run status) and enumerates the four possible states the tool returns. It does not differentiate itself from siblings like run_pentest or get_findings, but the purpose is unambiguous on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives, nor any mention that it is typically used to poll after run_pentest finishes. The only hint of context ('run_id returned by run_pentest') lives in the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_campaignsList campaignsA
Read-only

List the Darkmoon campaigns visible to the dashboard user, with ids and status. Use a campaign id with get_findings.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds real value on top: it discloses that results are filtered to the calling user's visibility (an authorization scope) and that the payload carries ids and status, which the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the purpose and scope front-loaded ahead of the workflow hint. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema in place, the description does the right thing by naming the returned fields (ids, status). It omits ordering, pagination, or truncation behavior, a minor gap for a listing tool but not one that blocks correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema cannot be misinterpreted and there is nothing for the description to disambiguate. Baseline for a no-parameter tool is 4; it is not a 5 because there are no argument semantics to enrich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (Darkmoon campaigns) plus the scope ('visible to the dashboard user') and what is returned (ids and status). The handoff to get_findings further anchors it against the sibling listing/retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives explicit downstream context: take a campaign id and feed it to get_findings. It does not state when NOT to use this tool (e.g. versus get_findings directly), so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_pentestStart a Darkmoon pentestA

Start an autonomous Darkmoon penetration test against one authorized target. The run executes in the background and can take a long time. Returns the run_id; poll it with get_run_status and read results with get_findings once a campaign exists (list_campaigns). Only use against systems the user owns or has explicit written authorization to test. Findings can include false positives and must be reviewed by a qualified human.

ParametersJSON Schema
NameRequiredDescriptionDefault
focusNoOptional focus areas, e.g. ['auth', 'injection']
targetYesHost, URL or scope to assess. Only targets you are authorized to test.
programNoOptional program name or rules-of-engagement note
severityNoOptional minimum severity to report

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the operation is non-read-only, open-world, non-idempotent and non-destructive, but the description adds material behavior beyond them: the run is asynchronous and long-running, it returns a run_id rather than findings, and results are noisy and require qualified human review. The authorization precondition and false-positive caveat are exactly the kind of disclosure an agent cannot infer from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tightly packed sentences with zero redundancy; the core action and target scope come first, then the asynchronous lifecycle, then the safety caveat. Every sentence contributes a distinct fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async kickoff tool with no output schema, the description supplies everything needed to invoke and follow up correctly: what it returns (run_id), the polling/reading path, the runtime characteristics, and the authorization and verification requirements. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters (target, focus, program, severity). The description only reinforces the target constraint with 'one authorized target' and adds no format, syntax or defaulting detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start an autonomous Darkmoon penetration test') plus the scope constraint ('against one authorized target'). The lifecycle sentence implicitly separates it from siblings by assigning get_run_status to polling, get_findings to result reading, and list_campaigns to campaign discovery, so an agent can position this as the entry point without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternatives and the conditions that select them ('poll it with get_run_status', 'read results with get_findings once a campaign exists (list_campaigns)'). It also states an explicit when-not-to-use condition: only systems the user owns or has written authorization to test.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedget_findings
    • First observedget_run_status
    • First observedlist_campaigns
    • First observedrun_pentest

TDQS

A4/5.0

Scored across 4 tools

Disambiguation4/5

Each tool targets a distinct action (start, status, list campaigns, get findings), but the relationship between a 'run' and a 'campaign' is not fully clear, which could cause slight misselection between get_run_status and list_campaigns when checking progress.

Naming Consistency5/5

All names follow a consistent snake_case verb_noun pattern (run_pentest, get_run_status, list_campaigns, get_findings), with clear verb prefixes that are easy to predict.

Tool Count5/5

Four tools is well-scoped for a focused pentest service; each tool (start, monitor, list campaigns, read findings) earns its place without redundancy or bloat.

Completeness3/5

Core start-monitor-results workflow is covered, but notable gaps exist: no cancel/stop for long-running runs, and no explicit way to map a run_id to its resulting campaign, forcing agents to guess from list_campaigns.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Autonomous pentests from one command: real security tools, working PoCs, and audit-ready reports, all driven via MCP.
    271 PyPI
    1,709
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables authorized penetration testing through MCP, providing parallel reconnaissance, vulnerability scanning, attack path analysis, and self-contained HTML reporting with compliance tagging.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables authorized bug bounty automation via a scope-enforced MCP bridge, supporting web, secrets, mobile, and LLM red-team scanning, with reporting and advisory.
    MIT
  • A
    license
    C
    quality
    B
    maintenance
    Enables authorized pentest and bug bounty workflows from any MCP client, with scoped recon, per-host rate limits, and scanner output turned into deduplicated, triaged finding cards.
    9
    32 PyPI
    1
    MIT