Skip to main content
Glama
Mhdd-24

@mhdd_24/ai-red-team-mcp

by Mhdd-24

@mhdd_24/ai-red-team-mcp

MCP server for Structured robustness/security testing.

Same architecture as @mhdd_24/sublime-mcp.

Full documentation: docs/WIKI.md


How it works (30 seconds)

You (chat) → MCP client → ai-red-team-mcp → AI Red Team APIs / CLIs / local tools

Related MCP server: PromptWall MCP Server

Prerequisites

Requirement

Notes

Node.js 18+

ESM TypeScript MCP server

Credentials / CLIs

See environment variables below


Install

Option A — npm (after publish)

npm install -g @mhdd_24/ai-red-team-mcp

Option B — npx

npx @mhdd_24/ai-red-team-mcp

Option C — clone and build

git clone https://github.com/Mhdd-24/AI-Red-Team-MCP.git
cd AI-Red-Team-MCP
npm install
npm run build
node dist/index.js

Configure Cursor

Edit ~/.cursor/mcp.json:

{
  "mcpServers": {
    "airedteam": {
      "command": "npx",
      "args": ["-y", "@mhdd_24/ai-red-team-mcp"],
      "env": {
        "_": "optional"
      }
    }
  }
}

Local development:

{
  "command": "node",
  "args": ["/absolute/path/to/AI-Red-Team-MCP/dist/index.js"]
}

Environment variables

Variable

Description

—

No required env


Tools

Tool

Description

airedteam_status

Show Structured robustness/security testing configuration / health.

airedteam_scan

Scan text for policy/safety issues.

airedteam_suite

Generate a test suite outline.


License

ISC

Available Tools

3 tools
airedteam_scanC

Scan text for policy/safety issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesInput/output text

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only says 'scan' without clarifying whether the operation is read-only, what kind of output it produces, whether it modifies anything, or how results are returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence and is appropriately front-loaded. It is not bloated, though its brevity contributes to some ambiguity in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is too sparse. It does not explain what a successful scan returns, how to interpret results, or how this tool relates to its siblings. An agent would likely need additional information to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the 'text' parameter. The tool description adds no new parameter-level meaning beyond what the schema provides, but the schema is sufficient, warranting the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Scan') and resource ('text') with a specific purpose ('policy/safety issues'). It is understandable and distinct enough from the sibling tools, though it does not explicitly differentiate itself from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus airedteam_status or airedteam_suite. There is no mention of exclusions, prerequisites, or typical scenarios, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

airedteam_statusB

Show Structured robustness/security testing configuration / health.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. The verb 'Show' implies a read-only operation, but the description does not explicitly state that it has no side effects, requires no permissions, or returns a snapshot. It also does not disclose whether it reflects current state or cached data. The lack of explicit safety information is a gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the action verb is front-loaded. It is appropriately concise for a simple status/health tool, though it could be slightly more specific without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and no annotations, the description should at least clarify what the output looks like or what 'configuration / health' specifically refers to. It does not describe the return format, the scope of 'health', or any prerequisites. For a tool this simple, it is adequate but leaves the agent guessing about the exact nature of the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema has 100% coverage (empty properties). Per the baseline rule for 0 parameters, a score of 4 is appropriate because there are no parameter semantics to clarify; the description does not need to add parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Show') and a resource ('configuration / health'), and the phrase 'robustness/security testing' provides some context. It is not a tautology, and it is distinguishable from the sibling tools (scan and suite), which imply running tests. However, the phrasing is a bit ambiguous ('configuration / health' could mean either or both).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus the siblings (airedteam_scan, airedteam_suite). It does not mention typical use cases, prerequisites, or that it should be used before/after other operations. The agent must infer its role from the name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

airedteam_suiteC

Generate a test suite outline.

ParametersJSON Schema
NameRequiredDescriptionDefault
policyYesPolicy summary

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. The verb 'Generate' suggests a non-destructive action, but no side effects, permissions, or state changes are mentioned. For a tool that likely creates an outline, the description does not clarify whether it modifies any existing data or what the output structure is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is appropriately concise and free of redundancy. However, it is not front-loaded with any scoping or differentiation details, and while it earns its place, it does not provide any extra value beyond the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (one parameter, no output schema), the description is still incomplete. It does not explain what a test suite outline consists of, how the policy parameter is used, or what the return value looks like. An agent would lack sufficient context to invoke the tool correctly or interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% coverage for the single parameter 'policy' with the description 'Policy summary', so the parameter is already documented. The tool description adds no additional meaning about the policy's format, role, or how it influences the output, leaving the schema to do the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and resource ('test suite outline'), making the primary action clear. However, it does not differentiate itself from sibling tools (airedteam_status, airedteam_scan) by mentioning scope or alternatives, so it lacks the explicit contrast seen in high-quality definitions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus its siblings. The description implies usage when an outline is needed, but there is no mention of prerequisites, context, or situations where this tool is inappropriate, leaving the agent to infer the intended scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedairedteam_scan
    • First observedairedteam_status
    • First observedairedteam_suite

TDQS

B3.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: status shows health/configuration, scan performs direct text analysis, and suite generates a test plan. There is no realistic confusion between them.

Naming Consistency4/5

All tools share the consistent airedteam_ prefix and use lowercase snake_case, making the naming predictable. However, they do not follow a strict verb_noun pattern: scan is a verb while status and suite are nouns.

Tool Count5/5

Three tools is at the lower end of the well-scoped range, but each one earns its place in a focused red-team utility server. There is no redundancy or unnecessary bloat.

Completeness3/5

The set covers health/status, text scanning, and test-suite outline generation, but there is no tool to execute a generated suite or retrieve detailed scan results. This leaves notable gaps in a full red-team testing workflow.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to perform authorized security testing and penetration testing operations including SSL/TLS analysis, port scanning, vulnerability scanning, and HTTP security header audits through natural language interactions.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables scanning LLM prompts and responses for prompt injection, jailbreaks, PII leakage, secret leakage, and other malicious content using deterministic rules, returning verdicts and safe redacted text.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides deterministic scanning and detection of prompt injection, jailbreak, data-exfiltration, and system-prompt leaks, with EIP-191 signed attestations and 0G Storage anchoring for verifiable safety reports.
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables auditing an AI agent's guardrails by providing adversarial test prompts and scoring responses against known refusal/guardrail patterns.
    -