Skip to main content
Glama
Zero-Harm-AI

PromptShield MCP

by Zero-Harm-AI
README.md
# Zero Harm AI MCP

Zero Harm AI MCP is a Model Context Protocol server that lets AI agents and runtime firewalls call Zero Harm AI safety checks for text, chat messages, prompts, tool inputs, and generated outputs.

The server is a thin adapter over `zero-harm-ai-detectors`. It should not duplicate detector logic from the detector package or from the Zero Harm AI GitHub Action.

## Goals

- Expose PII, secret, and harmful-content detection through MCP tools.
- Return structured findings that agents and firewalls can enforce.
- Support local/self-hosted operation for sensitive data.
- Keep logs privacy-safe by default.
- Provide stable tool contracts that can be used by coding agents, chat agents, and firewall.

## Non-Goals

- Reimplementing `zero-harm-ai-detectors`.
- Acting as a hosted service by default.
- Making policy enforcement decisions that belong to a firewall or calling agent.
- Replacing the Zero Harm AI GitHub Action.

## Relationship To Other Projects

```text
zero-harm-ai-detectors
  Shared detector engine for PII, secrets, and harmful content.

zero-harm-ai-gh-action
  GitHub Action and CI-oriented scanner for pull requests.

zero-harm-ai-mcp
  MCP server adapter that exposes detector functionality to AI agents.

zero-harm-ai-firewall (future)
  Runtime enforcement layer. It can call zero-harm-ai-mcp or use
  zero-harm-ai-detectors directly.
```

## Installation

Install from PyPI:

```bash
pip install zero-harm-ai-mcp
```

For local development:

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
ruff check .
```

## MCP Client Configuration

```json
{
  "mcpServers": {
    "zero-harm-ai": {
      "command": "zero-harm-ai-mcp",
      "args": []
    }
  }
}
```

## MCP Tools

### `zero_harm.scan_text`

Scan one text string for PII, secrets, and harmful content.

Use this for prompt inputs, generated outputs, tool arguments, logs, and arbitrary text.

### `zero_harm.scan_messages`

Scan chat-style messages while preserving message roles and indexes.

Use this when an agent wants to inspect a conversation before sending it to a model or tool.

### `zero_harm.redact_text`

Return a redacted version of text plus findings.

Use this when the caller wants to continue safely after removing sensitive spans.

### `zero_harm.evaluate_policy`

Map detector findings to an action recommendation.

Use this when a caller wants a normalized decision such as `allow`, `warn`, `redact`, or `block`.

## Working Examples

These examples are generated from the current local server implementation.

### `zero_harm.scan_text`

Input:

```json
{
  "text": "Contact alice@example.com before sharing the token.",
  "targets": ["pii", "secret", "harmful"],
  "redact": true
}
```

Output:

```json
{
  "schema_version": "1.0.0",
  "risk_level": "medium",
  "recommended_action": "redact",
  "categories": [
    "pii"
  ],
  "summary": {
    "total_findings": 1,
    "pii": 1,
    "secret": 0,
    "harmful": 0
  },
  "findings": [
    {
      "type": "email",
      "category": "pii",
      "severity": "medium",
      "confidence": 0.99,
      "span": {
        "start": 8,
        "end": 25
      },
      "redacted": "[PII]",
      "message_index": null,
      "message_role": null,
      "evidence_available": false
    }
  ],
  "redacted_text": "Contact [PII] before sharing the token."
}
```

### `zero_harm.scan_messages`

Input:

```json
{
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant."
    },
    {
      "role": "user",
      "content": "My email is alice@example.com."
    }
  ],
  "targets": ["pii", "secret", "harmful"],
  "redact": true
}
```

Output:

```json
{
  "schema_version": "1.0.0",
  "risk_level": "medium",
  "recommended_action": "redact",
  "categories": [
    "pii"
  ],
  "summary": {
    "total_findings": 1,
    "pii": 1,
    "secret": 0,
    "harmful": 0
  },
  "findings": [
    {
      "type": "email",
      "category": "pii",
      "severity": "medium",
      "confidence": 0.99,
      "span": {
        "start": 12,
        "end": 29
      },
      "redacted": "[PII]",
      "message_index": 1,
      "message_role": "user",
      "evidence_available": false
    }
  ],
  "redacted_text": "[{\"role\": \"system\", \"content\": \"You are a helpful assistant.\"}, {\"role\": \"user\", \"content\": \"My email is [PII].\"}]"
}
```

### `zero_harm.redact_text`

Input:

```json
{
  "text": "aws_access_key_id = AKIAIOSFODNN7EXAMPLE",
  "targets": ["pii", "secret", "harmful"]
}
```

Output:

```json
{
  "schema_version": "1.0.0",
  "risk_level": "high",
  "recommended_action": "block",
  "categories": [
    "secret"
  ],
  "summary": {
    "total_findings": 1,
    "pii": 0,
    "secret": 1,
    "harmful": 0
  },
  "findings": [
    {
      "type": "api_key",
      "category": "secret",
      "severity": "high",
      "confidence": 0.95,
      "span": {
        "start": 20,
        "end": 40
      },
      "redacted": "[SECRET]",
      "message_index": null,
      "message_role": null,
      "evidence_available": false
    }
  ],
  "redacted_text": "aws_access_key_id = [SECRET]"
}
```

### `zero_harm.evaluate_policy`

Input:

```json
{
  "text": "Contact alice@example.com before sharing the token.",
  "targets": ["pii", "secret", "harmful"],
  "redact": false
}
```

Output:

```json
{
  "schema_version": "1.0.0",
  "risk_level": "medium",
  "recommended_action": "warn",
  "categories": [
    "pii"
  ],
  "summary": {
    "total_findings": 1,
    "pii": 1,
    "secret": 0,
    "harmful": 0
  }
}
```

## Privacy Requirements

- Do not log raw input text by default.
- Do not log detected secret values by default.
- Include a config option for audit logs that stores only counts, categories, severities, and request metadata.
- Avoid sending data to external services unless explicitly configured.
- Keep the default transport local-first.

## Development

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
ruff check .
```

## Release

Build and validate distribution artifacts:

```bash
python -m build
twine check dist/*
```

See [RELEASE.md](RELEASE.md) for the full PyPI release flow.

TDQS

B3.1/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a distinct operation (evaluate policy, redact, scan text, scan messages) with clear descriptions, making them easily distinguishable.

Naming Consistency5/5

All tools follow a consistent 'promptshield.<verb>_<object>' pattern with lowercase underscores, using verbs like evaluate, redact, and scan.

Tool Count5/5

Four tools is appropriate for a content safety server, covering key tasks without being too many or too few.

Completeness5/5

The set covers evaluation, scanning, and redaction for both general text and chat messages, providing a complete surface for common content security needs.

Maintenance

ActivitySlowing
ResponsivenessNo issues