Skip to main content
Glama

Nougat-MCP

PyPI version Python versions License: GPL v3 MCP Protocol

nougat-mcp is a Model Context Protocol (MCP) server for high-fidelity OCR of scientific PDFs using Meta's Nougat.

It is designed for agent workflows where you need equations, tables, and structure preserved better than traditional OCR.

Why This Server

  • Scientific OCR quality tailored for papers, formulas, and dense layouts.

  • MCP-native interface for Codex, Claude, Cursor, Antigravity, and other clients.

  • Output-format control:

    • mmd: raw Nougat/Mathpix-style output.

    • md: renderer-friendly conversion (math delimiter and KaTeX compatibility fixes).

  • Settings file support so agents can read a shared default format policy.

Related MCP server: MCP-MinerU

Installation

Install from PyPI:

uv pip install nougat-mcp

This package installs nougat-ocr and pins known-sensitive dependencies for stability.

Tools

parse_research_paper

Arguments:

  • file_path (string): Absolute path to a local PDF.

  • output_format (string, optional):

    • default (default): uses server settings.

    • mmd: raw Nougat output.

    • md: converted markdown-friendly output.

Returns:

  • OCR result as a single text string in the requested format.

get_output_settings

Returns resolved server output settings, including where settings were loaded from.

Output Conversion (mmd -> md)

When output_format="md", the server applies compatibility conversions:

  • \[ ... \] -> $$ ... $$

  • \( ... \) -> $ ... $

  • \tag{...} -> visible equation label \qquad\text{(...)}

  • KaTeX delimiter normalization, for example:

    • \bigl{\|} ... \bigr{\|} -> \bigl\| ... \bigr\|

This avoids common renderer parse errors in markdown environments that are not fully MathJax-compatible.

Server Settings

Settings are read in this order:

  1. NOUGAT_MCP_SETTINGS (if set)

  2. ./settings.json (current working directory)

Example settings.json:

{
  "nougat_mcp": {
    "default_output_format": "md",
    "md_rewrite_tags": true,
    "md_fix_sized_delimiters": true
  }
}

Agent Configuration

Codex CLI

Add to ~/.codex/config.toml:

[mcp_servers.nougat]
command = "uvx"
args = ["nougat-mcp"]
enabled = true

[mcp_servers.nougat.env]
NOUGAT_MCP_SETTINGS = "/absolute/path/to/settings.json"

Claude Desktop

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "nougat": {
      "command": "uvx",
      "args": ["nougat-mcp"],
      "env": {
        "NOUGAT_MCP_SETTINGS": "/absolute/path/to/settings.json"
      }
    }
  }
}

Antigravity / Gemini Desktop

Add to ~/.gemini/settings.json:

{
  "mcpServers": {
    "nougat": {
      "type": "stdio",
      "command": "uvx",
      "args": ["nougat-mcp"],
      "env": {
        "NOUGAT_MCP_SETTINGS": "/absolute/path/to/settings.json"
      }
    }
  }
}

Cursor

In Cursor MCP settings, add:

{
  "mcpServers": {
    "nougat": {
      "command": "uvx",
      "args": ["nougat-mcp"],
      "env": {
        "NOUGAT_MCP_SETTINGS": "/absolute/path/to/settings.json"
      }
    }
  }
}

Note: Cursor MCP config location can vary by version/platform; use the MCP settings UI or your current JSON settings file.

Showcase (Real Page Example)

A real extraction from page 5 of src/2405.08770v1.pdf is included:

Quick comparison:

# mmd
\[DV=V_{x}. \tag{3.2}\]

# md
$$
DV=V_{x}. \qquad\text{(3.2)}
$$

Performance Notes

  • First run may download model weights (~1.4 GB).

  • CPU inference is significantly slower than GPU inference.

  • Use page subsets whenever possible to reduce runtime.

Compatibility Pins

To keep Nougat stable across environments, the package pins sensitive dependency ranges:

  • transformers>=4.35,<4.38

  • albumentations>=1.3,<1.4

  • pypdfium2<5.0

  • huggingface-hub<1.0

  • fsspec<=2025.10.0

Credits

License

GNU General Public License v3.0 (LICENSE).

Available Tools

2 tools
get_output_settingsB

Return resolved output settings so agents can adapt behavior. Reads NOUGAT_MCP_SETTINGS or ./settings.json.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals that it reads from environment variable or a file, which is helpful. However, it does not mention failure modes (e.g., missing file), whether it is read-only, or any other side effects. With no annotations, a bit more detail would improve transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise with two short sentences, no filler, and front-loaded with the core purpose. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (no params, no output schema), the description covers the basic functionality. However, it does not specify the format or structure of the returned settings, which might be needed for agents to use the output effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so schema coverage is 100% by default. The description adds meaning by specifying data sources (NOUGAT_MCP_SETTINGS or ./settings.json), which is useful beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Return') and the resource ('resolved output settings'), with a clear purpose for agents to adapt behavior. While the sibling tool is unrelated, the purpose is specific enough to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool or when to avoid it. The description lacks context about prerequisites or alternatives, such as if settings might be unavailable or how often to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parse_research_paperA

Highly accurate OCR for academic papers and scientific PDFs using Meta's Nougat model. Converts visual structures like tables, formulas, and multi-column layouts into clean Markdown.

Args: file_path (str): The absolute path to the PDF file on the local system. output_format (str): "default" uses settings.json preferences. "mmd" returns raw Nougat output. "md" converts math delimiters for broader Markdown renderer compatibility.

Returns: str: The extracted text in the requested markup format.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes
output_formatNodefault

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses use of Meta's Nougat model and conversion of visual structures, but lacks details on error handling, performance, or file size limits. The Returns section provides some clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a clear purpose sentence followed by argument and return descriptions. It is appropriately sized but could be slightly more concise by removing redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (not shown but noted), the description adequately covers input parameters and return format. It does not discuss errors or edge cases, but is generally complete for a straightforward parsing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It clearly describes 'file_path' as absolute path and elaborates on 'output_format' enum values: 'default', 'mmd', and 'md' with their behaviors. This adds meaningful context beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: high-accuracy OCR for academic papers and scientific PDFs, converting visual structures to Markdown. It distinguishes itself from the sibling tool 'get_output_settings' by being a parsing tool rather than a settings retrieval tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides guidance on when to use each output format option and mentions the 'default' utilizes settings.json. However, it does not specify prerequisites (e.g., file existence) or when to avoid using this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updatesv0.1.0
    • First observedget_output_settings
    • First observedparse_research_paper

TDQS

A3.9/5.0
Disambiguation5/5

The two tools have entirely distinct purposes: one retrieves configuration settings, the other parses PDFs.

Naming Consistency5/5

Both tools use the consistent verb_noun pattern with underscores, e.g., get_output_settings and parse_research_paper.

Tool Count4/5

With only two tools, the server is minimal but appropriate for its focused scope on PDF parsing with a settings helper.

Completeness4/5

The server provides the core functionality (parse and settings) for its domain, though it lacks auxiliary features like model listing.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables document and image parsing to extract text, tables, and formulas from PDFs, screenshots, and scanned documents. Features OCR capabilities, table recognition, LaTeX formula conversion, and MLX acceleration optimized for Apple Silicon.
    2
    16
    Apache 2.0
  • F
    license
    A
    quality
    C
    maintenance
    Enables AI agents to efficiently process large local and online PDFs through selective extraction of text, images, and metadata. It provides tools for content search and document outline navigation to optimize context window usage.
    7
    14
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables LLMs to read and extract content from PDF files with high-fidelity LaTeX recognition and layout awareness using a Python-based extraction engine. It includes a robust Node.js fallback and supports page range filtering for efficient processing of large documents.
    1
    60
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/svretina/nougat-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server