Skip to main content
Glama
vinkius-labs

doc-breach-mcp

by vinkius-labs

Using DocBreach in production? We want to hear about it β€” Join the Discord β†’ WAF horror stories, edge cases, and what you're building. The founder is in there.


πŸ›‘ The Problem: Agentic Workflows Are Blind

We're in the era of autonomous AI Agents β€” but the web was built to repel bots, not serve them.

When your Claude, Cursor, or Windsurf tries to read an obscure API's documentation, it gets annihilated by:

  1. Cloudflare WAFs throwing 403 CAPTCHAs at Node.js fetch().

  2. Empty SPA shells (Next.js, Mintlify, GitBook) that render nothing without a $300M headless browser.

  3. Legacy enterprise PDFs that crash the model's context window.

  4. Login walls that lock public API references behind OAuth gates.

  5. "AI-friendly" SaaS tools (Firecrawl, Jina, Context7) charging you $50/mo to read pages that are already public.

The LLM doesn't need a middleman. It needs raw signal.


Related MCP server: EasyPeasyMCP

βš”οΈ The Weapon: Guerrilla Architecture

DocBreach is a ruthless, 100% local MCP server. It doesn't ask for permission. It uses military-grade heuristics to extract clean, LLM-optimized Markdown from any developer portal β€” and it does it for free, forever.

Enemy Defense

DocBreach Tactical Override

πŸ›‘οΈ Cloudflare / WAF 403

Temporal Proxying β€” Hits a WAF? Silently pivots to the Wayback Machine. The docs from last week work just fine.

βš›οΈ JavaScript SPA Walls

Hydration Hijacking β€” Rips __NEXT_DATA__, __NUXT__, __GITBOOK_STATE__ straight from the DOM. Zero JS engine needed.

πŸͺŸ Hidden iFrames

Source Chasing β€” Detects embedded Swagger/Postman/Stoplight apps, destroys the wrapper, resolves the true origin URL.

πŸ“„ Legacy PDF Manuals

Native Brute-Force β€” In-memory PDF parsing. Your AI reads 2004 banking manuals like they're GitHub READMEs.

πŸ” Login Walls

Wall Detection β€” Identifies OAuth/SSO gates instantly and tells the agent to pivot to public alternatives.

πŸ•³οΈ Ghost Town Sites

Self-Healing Errors β€” No docs found? DocBreach guides the agent to search GitHub repos, SDK source code, or llms.txt files.

πŸ’Έ SaaS Scraping Taxes

Zero. Forever. Everything runs locally via Cheerio and Turndown. No API keys. No accounts. No telemetry.

"The LLM shouldn't be smart at scraping. It should be smart at coding. DocBreach handles the dirty work."


πŸš€ Quickstart

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "docbreach": {
      "command": "npx",
      "args": ["-y", "doc-breach-mcp"]
    }
  }
}

Cursor / Windsurf

Add to your MCP settings:

{
  "doc-breach": {
    "command": "npx",
    "args": ["-y", "doc-breach-mcp"]
  }
}

That's it. No API keys. No .env files. No sign-ups. It just works.


🧠 How It Thinks

DocBreach gives your AI agent 4 precision tools and lets the model drive:

You: "Integrate with the Datadog API and list all monitors"

Agent β†’ docs.discover({ query: "datadog API" })
     ← Found: docs.datadoghq.com/api/latest/ (openapi)

Agent β†’ docs.map({ domain: "docs.datadoghq.com" })
     ← πŸ—ΊοΈ Sitemap hierarchy, auto-generated Mermaid graph, and llms.txt discovery

Agent β†’ docs.read({ url: "https://docs.datadoghq.com/api/latest/" })
     ← πŸ“„ Clean Markdown + nav links + auth requirements

Agent β†’ docs.extract({ url: "https://api.datadoghq.com/api/v2/openapi.yaml", tag: "monitors" })
     ← πŸ“‹ GET /api/v1/monitor β€” List all monitors
        GET /api/v1/monitor/{id} β€” Get a monitor's details
        POST /api/v1/monitor β€” Create a monitor
        ...

Agent: "I see the API requires DD-API-KEY and DD-APPLICATION-KEY headers,
        and you need to select a DD_SITE (US1, EU, US3, US5, AP1)..."

The model reasons. DocBreach retrieves. Nobody hallucinates.

The 11-Step Reader Pipeline

Every URL passes through a battle-hardened, 11-step extraction pipeline:

 URL
  β”‚
  β”œβ”€ 1. Preflight ──────── HEAD check β†’ Content-Type, size, reject >10MB
  β”œβ”€ 2. Fetch ──────────── GET + Wayback Machine fallback on 403/503
  β”œβ”€ 3. Login Detection ── OAuth/SSO wall? β†’ abort + guide agent
  β”œβ”€ 4. Format Detection ─ OpenAPI? Postman? PDF? llms.txt? Markdown?
  β”œβ”€ 5. Binary Handling ── PDF <5MB β†’ in-memory parse
  β”œβ”€ 6. Spec Summary ───── OpenAPI/Postman β†’ structured Markdown
  β”œβ”€ 7. SPA Hydration ──── __NEXT_DATA__, __NUXT__, readme-data, GitBook
  β”œβ”€ 8. Nav Extraction ─── Sidebar links β†’ absolute URLs
  β”œβ”€ 9. iFrame Intel ───── Swagger/Postman/Stoplight embed β†’ true URL
  β”œβ”€ 10. HTML Cleaning ─── Cheerio β†’ remove headers, footers, ads, nav
  └─ 11. Markdown ──────── Turndown + boundary-aware truncation
  β”‚
  β–Ό
 Clean, LLM-ready Markdown

πŸͺ– The Uncomfortable Truth

The developer tooling market has a parasite problem.

Companies like Firecrawl, Jina Reader, and Context7 take public documentation β€” pages that are freely accessible to any browser β€” wrap them in a proprietary API, and charge you a monthly subscription to access what was already yours.

They aren't adding value. They're adding a toll booth to the public internet.

DocBreach exists because:

  • Documentation is public. If a human can read it, an agent should too.

  • Scraping is a solved problem. Cheerio + Turndown have existed for a decade. You don't need a $20M startup to parse HTML.

  • Your AI runs locally. Why should it phone home to a SaaS to read a README?

This is not a product. This is a crowbar.


πŸ“Š DocBreach vs. The Toll Booths

DocBreach

Firecrawl

Jina Reader

Context7

Cost

$0

$50+/mo

$30+/mo

Free (limited)

Runs locally

βœ…

❌ Cloud

❌ Cloud

❌ Cloud

No API keys

βœ…

❌

❌

❌

No telemetry

βœ…

❌

❌

❌

WAF bypass

βœ… Wayback

βœ… Paid proxy

❌

❌

SPA extraction

βœ… Hydration

βœ… Headless

❌

❌

PDF parsing

βœ… Native

βœ…

❌

❌

OpenAPI extraction

βœ…

❌

❌

❌

HATEOAS navigation

βœ…

❌

❌

❌

Cognitive rules

βœ…

❌

❌

❌

Open source

βœ… MIT

Partial

❌

βœ…


πŸ”§ Tools Reference

docs.discover

Find documentation sources for any service, library, or API.

docs.discover({ query: "stripe webhooks API" })
// β†’ [ { url, title, type: "openapi", source: "probe" }, ... ]

docs.map

Map the complete documentation structure of any domain. Extracts sitemaps, robots.txt, and llms.txt, returning an architectural blueprint.

docs.map({ domain: "docs.stripe.com" })
// β†’ { total: 1200, sections: { "Root": [...], "API": [...] }, ... }

docs.read

Read any documentation URL and return clean, LLM-ready Markdown.

docs.read({ url: "https://docs.stripe.com/webhooks" })
// β†’ { content: "# Webhooks\n\n...", nav_links: [...], format: "html" }

docs.search

Search for specific topics within a documentation site.

docs.search({ query: "authentication", site: "docs.stripe.com" })
// β†’ [ { url: ".../authentication", title: "Authentication", ... } ]

docs.extract

Extract structured endpoint information from OpenAPI/Swagger/Postman specs.

docs.extract({ url: "https://api.stripe.com/openapi/spec.json", tag: "charges" })
// β†’ [ { method: "POST", path: "/v1/charges", summary: "Create a charge" }, ... ]

πŸ† Beyond the MCP Specification

Google and Anthropic's official MCP best practices ask for "Single Responsibility," "Clear Descriptions," and "Structured Error Handling." That is the bare minimum.

Thanks to Vurb.ts, DocBreach elevates these concepts to the tenth power, operating years ahead of the standard protocol:

  • MVA Architecture (Model β†’ View β†’ Agent): Standard MCP returns raw JSON strings. We route everything through Fluent Presenters acting as smart egress firewalls, stripping noise before the LLM ever sees it.

  • HATEOAS Navigation: Instead of the agent guessing what to do next, every DocBreach response includes a .suggestActions() payload telling the model exactly which tool to call next.

  • JIT System Rules: Dynamic instructions injected mid-flight based on payload context (e.g., "The content was truncated, use search").

  • Self-Healing Errors: Standard MCP throws an error. DocBreach returns an error and the exact prompt/tool required to recover from it.

  • Server-Side Mermaid UI: Sends native ui.mermaid() visual graphs to the MCP Inspector to help humans see the architecture the agent sees.

  • State Sync & Cache Control: Emits .cached() directives at the protocol level to eliminate duplicate requests and save LLM token context.


πŸ“„ License

MIT β€” because documentation should be free, and so should the tools that read it.


Available Tools

5 tools
docs_discoverA
Read-only

[INSTRUCTIONS] Use this as the FIRST step when you need to find documentation. Use specific, descriptive queries. Example: "stripe API webhooks" instead of just "stripe". Combine your search intents into a single query. Do NOT call this tool in rapid loops β€” refine your query instead.

Find documentation sources for any service, library, or API [READ-ONLY]

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesWhat to search for (e.g., "datadog API monitoring endpoints")

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds behavioral context: it is read-only and should not be called in rapid loops, providing valuable guidance beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise and front-loaded with the key instruction. Each sentence adds value, though it could be slightly streamlined without losing important guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the tool's input well but does not describe the output format (e.g., what a documentation source looks like). Given the simple parameter set and annotations, the missing return description leaves some ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage with a description for the query parameter. The description adds rich semantic guidance: use specific, descriptive queries, combine intents, and provides an example, significantly enhancing understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Find documentation sources for any service, library, or API'. It specifies it should be the first step, distinguishing it from sibling tools like docs_search which might be for more specific searches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage instructions: 'Use this as the FIRST step', examples of good queries, and a warning against rapid loops. It lacks explicit alternatives but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_extractA
Read-only

[INSTRUCTIONS] Use this ONLY when docs.read has identified an OpenAPI/Swagger spec URL or Postman Collection. Provide a tag to filter endpoints by API group. If no tag is provided, returns a summary of all tags with endpoint counts.

Extract structured endpoint information from an OpenAPI, Swagger, or Postman spec [READ-ONLY]

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter by API tag (e.g., "monitors", "users")
urlYesURL of the OpenAPI/Swagger spec or Postman Collection
methodNoFilter by HTTP method
searchNoSearch endpoint descriptions

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds context by noting it is 'READ-ONLY' and explains the dual behavior (summary without tag, filtered endpoints with tag). This goes beyond the safety profile to clarify the tool's operational characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences) and well-structured. It front-loads the usage condition and then explains behavior. The '[INSTRUCTIONS]' prefix is slightly redundant but not detrimental. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description covers the main use cases (with/without tag, filtering by method/search). It could mention the output format but is adequate for an agent to understand the tool's purpose and behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so parameters are well-documented. The description adds value by explaining that omitting 'tag' returns a summary of all tags with endpoint counts, which is not in the schema. This clarifies default behavior and enhances interpretability.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool extracts structured endpoint information from OpenAPI, Swagger, or Postman specs. It specifies the action (extract), resource (endpoint info from specs), and scope (filtering by tag). This distinguishes it from sibling tools like docs_read (which likely reads raw content) and docs_search (which searches).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to use 'ONLY when docs.read has identified an OpenAPI/Swagger spec URL or Postman Collection.' It also explains behavior when tag is provided vs. not, giving clear guidance on when to use and what to expect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_mapA
Read-only

[INSTRUCTIONS] Use this as the FIRST step when exploring a new API or documentation site. Provide a domain (e.g., "stripe.com") and receive a complete table of contents with every documentation page organized by section. Then use docs.read on specific pages from the map.

Map the complete documentation structure of any domain [READ-ONLY]

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to map (e.g., "stripe.com", "docs.github.com")

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description includes '[READ-ONLY]' which aligns with annotations (readOnlyHint: true). It adds that it returns a 'complete table of contents with every documentation page organized by section', providing behavioral context beyond annotations. However, it does not detail potential limitations or error conditions, which would be needed for a higher score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the key instruction and purpose. It uses two clear sentences without unnecessary words. The structure is logical: instructional note, purpose, example, and guidance. It could be slightly more streamlined but is very efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema), the description explains the output type (table of contents organized by section) and provides a use case. It is complete enough for an agent to understand what to expect. Minor gaps like output format specifics are acceptable without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with a clear description of the 'domain' parameter including examples. The description adds no additional semantic value beyond what is already in the schema, so the score is at the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool maps the complete documentation structure of any domain, and explicitly distinguishes its role as the first step when exploring a new API. The phrase 'Map the complete documentation structure' combined with the example usage provides a specific verb and resource, setting it apart from siblings like docs_read and docs_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs to use this tool as the FIRST step when exploring a new API, and advises subsequent use of docs.read on specific pages. This provides clear when-to-use guidance and mentions the follow-up tool. It could be improved by explicitly stating when not to use it (e.g., if the user already knows the structure), but it is still strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docs_readA
Read-only

[INSTRUCTIONS] Reads a documentation page and returns clean Markdown. Handles HTML, JSON, YAML, OpenAPI specs, Postman Collections, PDFs (<5MB), and llms.txt. The response includes a "Related Documentation Links" section extracted from page navigation. ALWAYS check these links for authentication and getting-started pages before generating code.

Read any documentation URL and return clean, LLM-ready Markdown [READ-ONLY]

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFull URL of the documentation page to read
max_lengthNoMaximum output length in characters (default: 20000)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, and the description adds specific behaviors: handles multiple formats (HTML, JSON, YAML, etc.), PDFs under 5MB, and returns related documentation links. Fully consistent with annotations and provides extra detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two paragraphs: one explaining functionality and one giving instructions. It is mostly concise but could be slightly more streamlined. The key information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema, the description fully explains the output format (Markdown) and includes handling of multiple input types and extraction of related links. Also provides actionable instructions for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% coverage with descriptions for both parameters. The description adds value by explaining the tool returns Markdown and handles various input formats, which is not in the schema. Slight redundancy but acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads documentation pages and returns clean Markdown, listing supported formats. It distinguishes from sibling tools like docs_discover and docs_search by focusing on reading a specific page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to 'ALWAYS check these links for authentication and getting-started pages before generating code,' providing clear guidance on when and how to use the tool's output. Implicitly suggests using this tool for reading docs rather than discovering or extracting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.3/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: discovering sources, exploring structure, reading content, extracting specs, and searching within a known site. No overlap despite some similarity between discover and search.

Naming Consistency5/5

All tools follow the consistent 'docs_verb' pattern with imperative verbs, making the set predictable and easy to navigate.

Tool Count5/5

Five tools is an ideal size for a documentation assistant, covering core needs without excess.

Completeness5/5

The tool set covers the full workflow of finding, reading, mapping, and extracting documentation, with no obvious gaps for typical use cases.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    A locally-hosted MCP server that provides AI assistants with advanced web crawling capabilities, including structured data extraction, deep site crawling, and page screenshots. It enables users to convert single or multiple URLs into clean Markdown content for processing by LLMs without requiring external API keys for basic features.
  • A
    license
    Not graded
    quality
    C
    maintenance
    A lightweight, zero-config MCP server that makes documentation and API specifications instantly accessible to AI models using the llms.txt standard. It enables searching and retrieving full documentation, OpenAPI, and AsyncAPI specs without requiring a complex RAG infrastructure or vector database.
    16
    1
    Apache 2.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/vinkius-labs/doc-breach-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server