Skip to main content
Glama
Orchardxyz

Mammoth MCP Server

by Orchardxyz

mammoth-mcp

NPM version NPM downloads

A Model Context Protocol (MCP) server for converting DOCX files to HTML using mammoth.js.

Features

  • convert_docx_to_html: Convert DOCX files to clean HTML with advanced styling options

  • convert_docx_to_html_with_images: Convert DOCX files to HTML with embedded base64 images

  • extract_raw_text: Extract plain text content from DOCX files

  • convert_docx_to_markdown: Convert DOCX files to Markdown format

Related MCP server: MCP Document Converter

Installation

$ pnpm install

Development

$ npm run dev
$ npm run build

Usage

Configure MCP Client

Add the server to your MCP client configuration (e.g., Claude Desktop).

The configuration file is located at:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

This will automatically install and run the latest version:

{
  "mcpServers": {
    "mammoth": {
      "command": "npx",
      "args": ["-y", "mammoth-mcp"]
    }
  }
}

Option 2: Using local installation

If you've cloned the repository locally:

{
  "mcpServers": {
    "mammoth": {
      "command": "node",
      "args": ["/absolute/path/to/mammoth-mcp/dist/esm/index.js"]
    }
  }
}

After updating the configuration, restart Claude Desktop for the changes to take effect.

Available Tools

convert_docx_to_html

Convert a DOCX file to HTML.

Parameters:

  • filePath (string, required): Absolute path to the DOCX file

  • styleMap (string, optional): Custom style map to control Word styles to HTML conversion

  • ignoreEmptyParagraphs (boolean, optional): Whether to ignore empty paragraphs (default: true)

  • idPrefix (string, optional): Prefix for generated IDs (bookmarks, footnotes, endnotes)

  • includeDefaultStyleMap (boolean, optional): Whether to include default style map (default: true)

  • includeEmbeddedStyleMap (boolean, optional): Whether to include embedded style map from document (default: true)

Example:

{
  "filePath": "/path/to/document.docx",
  "styleMap": "p[style-name='Section Title'] => h1:fresh\nb[style-name='Emphasis'] => em",
  "ignoreEmptyParagraphs": false,
  "idPrefix": "doc-"
}

convert_docx_to_html_with_images

Convert a DOCX file to HTML with images embedded as base64 data URIs.

Parameters:

  • filePath (string, required): Absolute path to the DOCX file

  • styleMap (string, optional): Custom style map to control Word styles to HTML conversion

  • ignoreEmptyParagraphs (boolean, optional): Whether to ignore empty paragraphs (default: true)

  • idPrefix (string, optional): Prefix for generated IDs (bookmarks, footnotes, endnotes)

  • includeDefaultStyleMap (boolean, optional): Whether to include default style map (default: true)

  • includeEmbeddedStyleMap (boolean, optional): Whether to include embedded style map from document (default: true)

Example:

{
  "filePath": "/path/to/document-with-images.docx",
  "styleMap": "p[style-name='Code'] => pre:separator('\\n')"
}

extract_raw_text

Extract raw text from a DOCX file, ignoring all formatting. Useful for indexing, search, or text analysis.

Parameters:

  • filePath (string, required): Absolute path to the DOCX file

Example:

{
  "filePath": "/path/to/document.docx"
}

convert_docx_to_markdown

Convert a DOCX file to Markdown format. Note: This feature is deprecated by mammoth.js but remains functional.

Parameters:

  • filePath (string, required): Absolute path to the DOCX file

  • styleMap (string, optional): Custom style map to control Word styles to Markdown conversion

  • ignoreEmptyParagraphs (boolean, optional): Whether to ignore empty paragraphs (default: true)

  • idPrefix (string, optional): Prefix for generated IDs (bookmarks, footnotes, endnotes)

  • includeDefaultStyleMap (boolean, optional): Whether to include default style map (default: true)

  • includeEmbeddedStyleMap (boolean, optional): Whether to include embedded style map from document (default: true)

Example:

{
  "filePath": "/path/to/document.docx",
  "styleMap": "p[style-name='Heading 1'] => # "
}

How It Works

This MCP server uses mammoth.js to convert DOCX documents to clean, semantic HTML or Markdown. The conversion preserves:

  • Headings

  • Paragraphs

  • Lists

  • Tables

  • Bold/italic/underline formatting

  • Images (when using the with_images variant)

  • Custom style mappings (via styleMap parameter)

Style Maps

Style maps allow you to control how Word styles are converted to HTML/Markdown. Each line in a style map represents a mapping from a Word style to an HTML/Markdown element.

Example style map:

p[style-name='Section Title'] => h1:fresh
p[style-name='Subsection Title'] => h2:fresh
p[style-name='Code'] => pre:separator('\n')

For more information on style maps, see the mammoth.js documentation.

LICENSE

MIT

Available Tools

2 tools
convert_docx_to_htmlC

Convert a DOCX file to HTML using mammoth. Supports reading from a file path and returns the HTML content.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesAbsolute path to the DOCX file to convert

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool 'returns the HTML content', which gives basic output information, but lacks details on error handling, performance traits (e.g., speed, memory usage), or constraints like file size limits. For a conversion tool, this leaves gaps in understanding its operational behavior and reliability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, stating the core functionality in the first sentence. Both sentences earn their place by covering conversion purpose and input method. However, it could be slightly more structured by explicitly separating input and output details, but overall, it avoids unnecessary verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (file conversion), no annotations, and no output schema, the description is adequate but incomplete. It covers the basic what and how but misses key contextual elements like error cases, output format details beyond 'HTML content', or integration with the sibling tool. For a standalone tool, it meets minimum viability but lacks depth for robust agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with 'filePath' clearly documented as an absolute path. The description adds minimal value beyond this, only reiterating 'reading from a file path' without providing additional context like supported file formats beyond DOCX or path validation rules. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: converting DOCX files to HTML using mammoth. It specifies the verb ('convert'), resource ('DOCX file'), and technology ('using mammoth'), making it specific and actionable. However, it doesn't explicitly differentiate from its sibling tool 'convert_docx_to_html_with_images', which likely handles images differently, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions 'supports reading from a file path', which hints at input method, but doesn't clarify scenarios where this tool is preferred over its sibling or other conversion methods. There's no mention of prerequisites, limitations, or comparative contexts, leaving usage decisions ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

convert_docx_to_html_with_imagesA

Convert a DOCX file to HTML with embedded images as base64 data URIs

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesAbsolute path to the DOCX file to convert

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. While it describes the core transformation behavior, it lacks critical details like error handling, performance characteristics, file size limitations, or what happens with unsupported DOCX features. For a file conversion tool with zero annotation coverage, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that conveys the complete purpose without any wasted words. It's appropriately sized for a single-parameter tool and front-loads the key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with 100% schema coverage but no annotations and no output schema, the description adequately covers the basic transformation purpose. However, it lacks information about the output format details, error conditions, or limitations that would be important for practical usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents the single parameter 'filePath'. The description doesn't add any parameter-specific information beyond what the schema provides, such as file format requirements or path validation rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Convert'), resource ('DOCX file'), and output format ('HTML with embedded images as base64 data URIs'). It distinguishes from the sibling tool 'convert_docx_to_html' by explicitly mentioning the image embedding feature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by specifying the input format (DOCX) and output format (HTML with images). However, it doesn't explicitly state when to use this tool versus the sibling 'convert_docx_to_html' or provide any exclusion criteria or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updates
    • First observedconvert_docx_to_html
    • First observedconvert_docx_to_html_with_images

TDQS

B3/5.0
Disambiguation2/5

The two tools have highly overlapping purposes—both convert DOCX to HTML—with the only difference being image handling. An agent could easily misselect between them, as the distinction is subtle and not clearly differentiated in the names alone. This creates ambiguity about which tool to use for a given conversion task.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern with 'convert_docx_to_html' as the base, and the second tool adds a descriptive suffix '_with_images'. The naming is predictable and clear, using snake_case uniformly without any deviations or mixed conventions.

Tool Count2/5

With only 2 tools, the server feels thin and under-scoped for a 'Mammoth MCP Server' that implies broader DOCX processing capabilities. This minimal set lacks basic operations like converting to other formats (e.g., Markdown, plain text) or handling DOCX metadata, making it inadequate for comprehensive document conversion tasks.

Completeness2/5

The tool surface is severely incomplete for a DOCX conversion server. It only covers HTML output with two variants, missing essential operations such as conversion to other formats (e.g., PDF, Markdown), extraction of text or images separately, or handling of DOCX properties. This will cause agent failures when broader document processing is needed.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Orchardxyz/mammoth-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server