Skip to main content
Glama
wedo911

regexguard

regexguard-mcp-server

Glama score

An MCP server that gives any AI agent a way to sanity-check a regex it just generated -- both what it actually matches and whether it's safe to run against untrusted input -- before shipping it. Fully local: no API key, no network call, no dependency beyond the MCP SDK and Zod.

Why

Regexes are a notoriously easy place to introduce a bug that looks fine in every example you happen to test. Two failure modes in particular are both common and easy to miss by eye:

  1. It doesn't match what you think it matches. explain_regex turns the pattern into a real syntax tree and describes it in plain English, so "does this actually require at least one digit?" has a fast answer that doesn't depend on trusting your own reading of nested brackets.

  2. It's a denial-of-service vector. A regex with nested quantifiers ((a+)+) or ambiguous alternation inside a repeated group ((a|a)+) can make a backtracking engine take exponential time on a crafted (or even accidental) non-matching input -- this is ReDoS, a real and repeatedly-exploited vulnerability class, and a plausible defect in any regex an agent writes without testing it against adversarial input. check_redos_risk flags the structural shape without ever executing the pattern -- it's safe to run on untrusted or deliberately malicious regex source.

Related MCP server: Regex Toolkit MCP Server

Tools

explain_regex

Parses a pattern into an AST and returns a plain-English description of what it matches.

check_redos_risk

Statically analyzes a pattern's structure for nested quantifiers and ambiguous alternation inside a repeated group -- the two classic causes of catastrophic backtracking. Returns "safe", "high", or "critical", with a specific finding for each issue found.

Both tools share one parser (src/services/parser.ts): a real recursive-descent regex parser (literals, character classes, shorthand classes, anchors, capturing/non-capturing/named groups, lookaround, alternation, quantifiers, backreferences), not a bag of string-matching heuristics against the pattern's raw source text.

This is a heuristic structural check, not a formal verifier -- check_redos_risk can tell you a pattern has the textbook exponential-blowup shape; it can't prove a pattern is fast on all inputs, and there are ReDoS patterns outside the two shapes it currently detects. Treat a "safe" result as "nothing obvious found," not a guarantee.

Install and configure

git clone https://github.com/wedo911/regexguard-mcp-server.git
cd regexguard-mcp-server
npm install
npm run build

Add it to your MCP client's config (e.g. claude_desktop_config.json, or a project's .mcp.json for Claude Code):

{
  "mcpServers": {
    "regexguard": {
      "command": "node",
      "args": ["/absolute/path/to/regexguard-mcp-server/dist/index.js"]
    }
  }
}

Run the tests

npm run build
node --test tests/parser.test.mjs tests/explain.test.mjs tests/redosCheck.test.mjs

44 tests cover the parser grammar, the explanation output, and both true positives ((a+)+, (a*)*, (a|a)+, (a|ab)+, patterns nested inside non-capturing groups) and true negatives ((cat|dog)+, a realistic username pattern, a realistic email pattern, sibling — not nested — repetitions) for the ReDoS check, so the false-positive rate on ordinary patterns is a tested property, not a hope.

Try it without a client

npx @modelcontextprotocol/inspector --cli node dist/index.js \
  --method tools/call --tool-name check_redos_risk \
  --tool-arg pattern='^(([a-zA-Z0-9])+([\.-]?([a-zA-Z0-9])+)*)$'

License

MIT — see LICENSE.

Available Tools

2 tools
check_redos_riskCheck Regex for Catastrophic Backtracking (ReDoS) RiskA
Read-onlyIdempotent

Statically analyze a regex's structure for the two classic causes of catastrophic backtracking: nested quantifiers (e.g. "(a+)+") and ambiguous alternation inside a repeated group (e.g. "(a|a)+"). Both can make a backtracking regex engine take exponential time on a crafted or even accidental non-matching input -- a real, exploitable denial-of-service vector, and a common defect in regexes generated without testing against adversarial input.

This tool NEVER executes the pattern -- it only parses and inspects the pattern's source structure, so it's safe to run on untrusted or deliberately malicious patterns without risk of hanging.

Args:

  • pattern (string, 1-1000 chars): the regex source, without delimiters.

Returns: For JSON format: { "parsed": boolean, "error": string | null, "risk": "safe" | "high" | "critical", "findings": [ { "category": "nested_quantifier" | "ambiguous_alternation", "severity": "high" | "critical", "message": string } ] }

Examples:

  • Use when: "is this regex I just generated safe to run against user input?" -> pass the pattern before using it

  • Use when: reviewing a regex from an untrusted source before adding it to a codebase

  • Don't use when: you need proof a regex is fast on all inputs -- this is a heuristic structural check (no false negatives are guaranteed to be caught), not a formal verifier

Error Handling:

  • Returns an error (not an exception) if the pattern doesn't parse.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesThe regex source, without the surrounding slashes, e.g. "^(a+)+$".

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only, idempotent, non-destructive. The description adds critical behavioral context: the tool never executes the pattern, is safe on untrusted or malicious input, returns errors rather than exceptions on parse failure, and is heuristic with possible missed risks. This is significant added transparency with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with headings for purpose, usage, arguments, return format, examples, and error handling. Every section contributes useful operational information and the purpose is front-loaded within the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since no output schema exists, the description provides the full JSON return structure, error behavior, and usage boundaries, making the tool actionable. Minor ambiguity about regex flavor is acceptable given the structural check nature, and the description is otherwise complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the pattern parameter is fully described with type, min/max length, and an example in the schema. The description repeats the 'without delimiters' point but adds no new semantic meaning beyond what the schema already provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: statically analyzing a regex's structure for the two classic causes of catastrophic backtracking (nested quantifiers and ambiguous alternation). This clearly distinguishes it from a generic regex explainer like the sibling explain_regex, even without naming it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'Use when' scenarios (generated regex before use, reviewing untrusted regex) and a clear 'Don't use when' caution that it is a heuristic structural check, not a formal verifier. This gives the agent clear selection criteria and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_regexExplain a RegexA
Read-onlyIdempotent

Parse a regular expression into a real syntax tree and describe in plain English what it matches. Use this to sanity-check a regex you (or someone else) wrote actually does what you intended, before shipping it.

Args:

  • pattern (string, 1-1000 chars): the regex source, without delimiters (no leading/trailing "/").

Returns: For JSON format: { "explanation": string, "parsed": boolean, "error": string | null } explanation is empty and error is set if the pattern could not be parsed (e.g. unbalanced parentheses, unsupported syntax).

Examples:

  • Use when: "does this regex actually require the string to start with a letter?" -> pass the pattern, read the explanation

  • Don't use when: you need to actually run the regex against text -- this tool never executes the pattern, only analyzes its source

Error Handling:

  • Returns an error (not an exception) if the pattern doesn't parse, with the parser's best guess at where the problem is.

ParametersJSON Schema
NameRequiredDescriptionDefault
patternYesThe regex source, without the surrounding slashes, e.g. "^[a-z0-9_]{3,16}$".

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses important behavioral traits: the tool never executes the pattern, returns a parsed boolean, returns an error rather than throwing an exception on unparseable input, and provides the parser's best guess at the problem. This adds real context beyond the readOnly/idempotent hints and contains no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with Args, Returns, Examples, and Error Handling sections, and it front-loads the core purpose. It is slightly redundant with the schema in the Args section, but the overall organization makes it easy for an agent to parse and apply.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only one parameter and no output schema, the description fully compensates by specifying the exact JSON return shape, the meaning of each field, error behavior, and a concrete usage example. An agent has enough information to call the tool correctly and interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description's Args section essentially repeats the schema: pattern string, 1-1000 chars, no delimiters. It adds no new semantic information beyond what the input schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Parse a regular expression into a real syntax tree and describe in plain English what it matches.' It also clearly differentiates the tool from running a regex by explicitly saying it 'never executes the pattern, only analyzes its source,' which distinguishes it from execution-oriented tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit use guidance: 'Use this to sanity-check a regex... before shipping it' and includes both a 'Use when' example and a 'Don't use when' exclusion. It does not explicitly name check_redos_risk as the alternative for security risk assessment, so it misses the full when/when-not/alternatives pattern, but the context is otherwise clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observedcheck_redos_risk
    • First observedexplain_regex

TDQS

A4.6/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: explain_regex describes what a pattern matches, while check_redos_risk analyzes vulnerability to catastrophic backtracking. They share input format but have no functional overlap, making misselection unlikely.

Naming Consistency5/5

Both tools follow a consistent verb_noun snake_case pattern (explain_regex and check_redos_risk). The naming is predictable and aligns with the domain.

Tool Count4/5

With only two tools, the server is minimally scoped, but for a dedicated regex-analysis utility this is reasonable and each tool addresses a core need. Slightly under typical range but appropriate for the narrow focus.

Completeness5/5

The server covers the essential aspects of regex analysis: comprehension (explain_regex) and security risk (check_redos_risk). Syntax validation is implicitly handled via parse errors. There are no obvious missing operations within the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    RegexForge gives AI agents a reliable way to get a regex without asking an LLM to hallucinate one. Pass in labeled examples (strings that should match, strings that shouldn't) plus an optional description; get back the regex, a proof matrix showing it handles every example, and a backtracking-risk audit flagging catastrophic-backtracking patterns. Pure symbolic synthesis over a template bank with
    -
  • F
    license
    A
    quality
    D
    maintenance
    Enables LLM agents to extract, validate, and mask personally identifiable information using deterministic regular expressions, reducing token usage and hallucination risks.
    3
    7 npm
    3
    -
  • A
    license
    A
    quality
    D
    maintenance
    Provides tools to test regex patterns for correctness, performance (ReDoS), and memory usage, and suggests safe rewrites. Enables LLMs to iterate on regex generation with verifiable feedback.
    9
    MIT