Skip to main content
Glama
Jyoti429

RedactAI MCP Server

by Jyoti429

RedactAI MCP Server

A standalone, production-ready Model Context Protocol (MCP) server built in TypeScript for sensitive data detection, sanitization, and PII redaction before downstream AI agent processing.

Supports both Local STDIO MCP and Remote Streamable HTTP MCP (compatible with Render, cloud hosts, and remote agent workflows).


1. What the Project Does

RedactAI MCP Server exposes a high-performance MCP tool named sanitize_text. It inspects text for sensitive personal identifiable information (PII) including Names, Email Addresses, Phone Numbers, ID Numbers (Employee IDs, Tax IDs, National IDs), and Locations, redacting them using configurable strategies:

  • MASK (default): Replaces sensitive entities with semantic tokens (e.g. [NAME], [EMAIL], [PHONE], [ID_NUMBER], [LOCATION]).

  • HASH: Computes deterministic SHA-256 hash placeholders (e.g. [HASH:a1b2c3d4]) to preserve referential consistency across documents without exposing raw plaintext.

  • REMOVE: Strips the sensitive entity completely ("").


Related MCP server: classifinder-mcp

2. Why MCP (Model Context Protocol)?

AI agents and LLMs frequently ingest raw user prompts, support transcripts, logs, and sensitive internal records. Sending unredacted PII to external model APIs introduces critical privacy, security, and regulatory compliance risks (GDPR, HIPAA, DPDP).

By exposing the RedactAI engine as a native MCP tool, any MCP-compatible agent workflow, IDE, or client (Claude Desktop, Antigravity, Cursor, LangChain, LlamaIndex, or custom agent runtimes) can automatically call sanitize_text directly within the agent loop to scrub sensitive context before executing external prompts.


3. Architecture & Transports

Both local and remote transports share the exact same redaction and detection pipeline:

                    MCP CLIENT / AI AGENT
                   /                     \
                  /                       \
             STDIO                         HTTPS
          (Local CLI)                 (Remote Server)
               |                              |
               v                              v
          src/index.ts                   src/http.ts
               \                              /
                \                            /
                 v                          v
                  Tool: sanitize_text (/mcp)
                               |
                               v
                        Redaction Service
                               |
                               v
                       DetectionProvider
                               |
                               v
                    Teammate's Core Engine
                       (future adapter)
  • Local STDIO Transport (src/index.ts): Communicates over stdin/stdout.

  • Remote Streamable HTTP Transport (src/http.ts): Exposes standard MCP over HTTP at /mcp with full SSE/streaming and CORS support.

  • TemporaryDetectionProvider: High-accuracy deterministic regex and context-based provider for standalone demonstration.

  • Teammate Core Adapter: Drop-in replacement via the DetectionProvider interface.


4. Installation & Build

Ensure Node.js (v20+ or v22+) is installed:

# Clone or navigate into project directory
cd redactai-mcp

# Install dependencies
npm install

# Typecheck TypeScript
npm run typecheck

# Build both STDIO and HTTP entry points to dist/
npm run build

# Run all 21 unit, STDIO, and HTTP tests
npm run test

5. Running the MCP Server

A. Local STDIO MCP Server

Used for local CLI tools, Claude Desktop, Cursor, or local Antigravity setups.

# Production start
npm run start
# Or directly: node dist/index.js

# Development mode
npm run dev

B. Remote Streamable HTTP MCP Server

Used for remote web clients, cloud-hosted agents, or teammate web applications connecting over HTTPS.

# Production start (listens on PORT or 10000, binds to 0.0.0.0)
npm run start:http
# Or directly: node dist/http.js

# Development mode
npm run dev:http

Default endpoints when running locally:

  • MCP Endpoint: http://localhost:10000/mcp

  • Health Check: http://localhost:10000/health

  • Info: http://localhost:10000/


6. Testing with MCP Inspector

The official MCP Inspector provides an interactive web UI to test tools and schema definitions.

Testing Local STDIO Server:

npx @modelcontextprotocol/inspector node dist/index.js

Testing Remote Streamable HTTP Server:

  1. Start the HTTP server in one terminal:

    npm run start:http
  2. Launch MCP Inspector in another terminal connecting via HTTP transport:

    npx @modelcontextprotocol/inspector http://localhost:10000/mcp

7. Render Deployment Guide

A pre-configured render.yaml blueprint is included in the project.

Automatic Blueprint Deployment

  1. Push this repository to GitHub/GitLab.

  2. In the Render Dashboard, click New > Blueprint.

  3. Connect your repository. Render will automatically read render.yaml:

    • Environment: Node

    • Build Command: npm install && npm run build

    • Start Command: npm run start:http

    • Health Check Path: /health

    • Port: Auto-assigned by Render (process.env.PORT)

  4. Once deployed, your public MCP URL will be:

    https://YOUR-SERVICE-NAME.onrender.com/mcp

Health Check Verification

curl https://YOUR-SERVICE-NAME.onrender.com/health
# Returns: {"status":"ok"}

8. Example sanitize_text Invocation

Tool Name

sanitize_text

Description

Detect and redact sensitive personal information from text before downstream processing.

Input Schema

{
  "text": "Contact John Mehta at john.mehta@example.com or 9876543210.",
  "rules": {
    "NAME": "MASK",
    "EMAIL": "MASK",
    "PHONE": "MASK",
    "ID_NUMBER": "MASK",
    "LOCATION": "MASK"
  }
}

(Note: The rules object is optional. If omitted, all entity types default to MASK.)

Custom Rules Example (e.g. Hashing email & Removing phone)

{
  "text": "Contact John Mehta at john.mehta@example.com or 9876543210.",
  "rules": {
    "NAME": "MASK",
    "EMAIL": "HASH",
    "PHONE": "REMOVE"
  }
}

Example Structured Response

{
  "redactedText": "Contact [NAME] at [EMAIL] or [PHONE].",
  "detections": [
    {
      "type": "NAME",
      "value": "John Mehta",
      "start": 8,
      "end": 18,
      "confidence": 98,
      "redaction": "[NAME]"
    },
    {
      "type": "EMAIL",
      "value": "john.mehta@example.com",
      "start": 22,
      "end": 44,
      "confidence": 99,
      "redaction": "[EMAIL]"
    },
    {
      "type": "PHONE",
      "value": "9876543210",
      "start": 48,
      "end": 58,
      "confidence": 95,
      "redaction": "[PHONE]"
    }
  ],
  "summary": {
    "totalDetections": 3,
    "byType": {
      "NAME": 1,
      "EMAIL": 1,
      "PHONE": 1,
      "ID_NUMBER": 0,
      "LOCATION": 0
    }
  }
}

9. Connecting Teammate's Core Engine

When your teammate completes the core redaction engine or an ML/NER model, integrate it in 3 simple steps:

Step 1: Implement DetectionProvider

Create an adapter class implementing the DetectionProvider interface:

// src/teammate-adapter.ts
import { Detection, DetectionProvider } from "./types.js";
import { realTeammateEngine } from "teammate-core-engine";

export class CoreEngineAdapter implements DetectionProvider {
  async detect(text: string): Promise<Detection[]> {
    const coreResults = await realTeammateEngine.scan(text);

    return coreResults.map((item) => ({
      type: item.category as any, // "NAME" | "EMAIL" | "PHONE" | "ID_NUMBER" | "LOCATION"
      value: item.rawText,
      start: item.startIndex,
      end: item.endIndex,
      confidence: item.score,
    }));
  }
}

Step 2: Inject the Adapter

In src/index.ts and src/http.ts:

import { CoreEngineAdapter } from "./teammate-adapter.js";
import { RedactionService } from "./redaction.js";
import { runStdioServer } from "./index.js";
import { createHttpServer } from "./http.js";

// Initialize with teammate's adapter
const customProvider = new CoreEngineAdapter();
const redactionService = new RedactionService(customProvider);

// For STDIO:
// runStdioServer(redactionService);

// For HTTP:
// createHttpServer(undefined, undefined, redactionService);

Step 3: Run Tests

npm run test

10. Security & Privacy Guarantees

  • Zero Persistence: No input text or processed output is logged or written to storage.

  • Protocol Separation: All operational logs are directed strictly to stderr.

  • Deterministic Privacy: HASH mode computes one-way SHA-256 tokens preserving referential linkability without exposing raw cleartext.

  • Hackathon Note: The public HTTP endpoint is designed for hackathon demonstration and agent workflow integration.

Available Tools

1 tool
sanitize_textB

Detect and redact sensitive personal information from text before downstream processing.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text containing sensitive personal information to be sanitized.
rulesNoOptional mapping of entity types to redaction modes. Defaults to MASK for all types.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, but it only states the high-level action. It does not disclose the return shape (sanitized text only versus a structured report of detections), whether the input is mutated or a new string is returned, whether HASH is a one-way irreversible transform, or what REMOVE leaves behind in the text. For a tool that alters sensitive data, these are materially consequential gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that delivers the verb, resource, and use-phase in order of importance. There is no filler, no restating of the tool name, and no duplication of schema content. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a moderately complex nested `rules` object (5 entity types × 3 modes), no output schema, and no annotations, yet the description does not explain what the tool returns, whether it mutates the input, or the behavior of the non-default modes. An agent would have to guess whether the result is redacted text alone or redacted text plus detection metadata — a real gap given nothing else compensates for it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both `text` and `rules` have descriptive text, and each nested entity type (NAME, EMAIL, PHONE, LOCATION, ID_NUMBER) documents its enum and meaning. The tool description itself adds no parameter-level information, so per the high-coverage baseline, a 3 is appropriate. The schema's own statement that rules 'Defaults to MASK for all types' already handles the default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('detect and redact') anchored to a concrete resource ('sensitive personal information from text'), and it adds the intended use phase ('before downstream processing'). It is not a tautology of the tool name. With no sibling tools provided, there is no differentiation to demonstrate, so it stops short of the full 5 by not actively distinguishing from possible alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is 'before downstream processing,' which implies the tool should be invoked prior to other text-processing steps but never states this explicitly. No alternatives are named (there are no siblings), and no exclusion criteria are given — e.g., what to do if detection without redaction is needed. This is implied context rather than explicit when-to-use/when-not-to-use direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedsanitize_text

TDQS

A3.7/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of an agent confusing it with another operation. sanitize_text has a clear, singular purpose that is easy to identify.

Naming Consistency5/5

The single tool name follows a clean verb_noun convention and is descriptive of its function. Since there are no other tool names, there are no inconsistencies to introduce confusion.

Tool Count4/5

One tool is on the low end of typical MCP server scopes, but for a focused RedactAI service it is reasonable and not trivial. The tool directly addresses the server's intended operation without unnecessary bloat.

Completeness5/5

The tool fully covers the stated purpose of detecting and redacting sensitive personal information before downstream processing. No additional lifecycle or auxiliary operations are needed for this narrow, stateless transformation task.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers