pdftomarkdown
by HEMANT16
README.md
# ð PDF to Markdown (PDF to MD) â Fast, Private, Open-Source Converter
> **The ultimate open-source PDF to Markdown converter for Humans, Developers, and AI Agents.**
> Effortlessly convert **PDFs, Word (`.docx`), Excel & CSVs (`.xlsx`, `.csv`), and YouTube URLs** into clean, structured, token-efficient Markdown (`.md`).
[](https://pdftomarkdown.fyi)
[](https://www.npmjs.com)
[](https://github.com/HEMANT16/pdftomarkdown)
[](https://opensource.org/licenses/MIT)
[](https://pdftomarkdown.fyi)
[](https://modelcontextprotocol.io)
---
## ð Table of Contents
- [Why PDF to Markdown?](#-why-pdf-to-markdown)
- [Architecture & Conversion Flow](#-architecture--conversion-flow)
- [Key Features](#-key-features)
- [How to Use (3 Ways)](#-how-to-use-3-ways)
- [1. Live Web Application](#1-live-web-application-zero-install)
- [2. Zero-Install CLI (`npx`)](#2-zero-install-cli-npx)
- [3. AI Agent & MCP Server](#3-ai-agent--mcp-server-claude-cursor)
- [Benchmark & Comparison Matrix](#-benchmark--comparison-matrix)
- [Supported Formats](#-supported-formats)
- [CLI Reference & Options](#-cli-reference--options)
- [Programmatic API (Node / TypeScript)](#-programmatic-api-node--typescript)
- [Technical SEO & Performance](#-technical-seo--performance)
- [Frequently Asked Questions (FAQ)](#-frequently-asked-questions-faq)
- [License](#-license)
---
## ðŊ Why PDF to Markdown?
PDFs are designed for **visual printing**, not editing or AI computing. They trap content inside fixed-coordinate binary blobs. **Markdown (`.md`)** is the universal standard for:
1. ðĪ **AI & LLM Workflows**: Maximizes prompt token density for Claude, ChatGPT, Cursor, and Gemini by stripping layout bloat.
2. ðïļ **Knowledge Management**: Effortless importing into **Obsidian, Notion, Logseq, Roam, and Apple Notes**.
3. ð **Technical Documentation**: Direct ingestion into static site generators like **Docusaurus, VitePress, MkDocs, and Astro**.
4. ðŋ **Git Version Control**: Clean line-by-line diffs instead of opaque binary changes.
---
## ðïļ Architecture & Conversion Flow
```
INPUT SOURCE
ââââââââââââââââââââââââââââââââŽâââââââââââââââââââââââââââââââ
âž âž âž
ð PDF Document ð Word / Excel ðĨ YouTube Link
(Local / Remote URL) (.docx, .xlsx, .csv) (Video URL / Shorts)
â â â
ââââââââââââââââââââââââââââââââžâââââââââââââââââââââââââââââââ
âž
ââââââââââââââââââââââââââââââââââââââââ
â Unified Geometric & NLP Engine â
ââââââââââââââââââââââââââââââââââââââââĪ
â âĒ Visual Reading Order Sorter â
â âĒ H1âH4 Typography Hierarchy â
â âĒ GFM Table Reconstructor â
â âĒ Code Block & List Parser â
â âĒ Link & Metadata Extractor â
ââââââââââââââââââââââââââââââââââââââââ
â
âž
CLEAN MARKDOWN
ââââââââââââââââââââââââââââââââžâââââââââââââââââââââââââââââââ
âž âž âž
ð Web UI Workspace ðŧ CLI Output Stream ðĪ AI Agent Context
(Split Editor + Preview) (stdout / .md file export) (Claude / Cursor MCP)
```
---
## âĻ Key Features
- ð **100% In-Browser Privacy**: In the web app, documents are converted locally in memory via WebAssembly / JavaScript. Zero files are uploaded to any external server.
- ⥠**Zero-Install CLI (`npx`)**: Convert any PDF, Word doc, or CSV straight from your command line with no setup.
- ð **Smart Table Detection**: Automatically aligns multi-column rows into clean GitHub Flavored Markdown (GFM) pipe tables (`| Header | ... |`).
- ðïļ **Multi-Format Extensibility**: Converts PDFs, Word (`.docx`), Excel/CSVs (`.csv`), and YouTube videos into Markdown.
- ðĪ **Native Model Context Protocol (MCP) Server**: Connect directly to **Claude Desktop**, **Cursor**, **Windsurf**, and **Antigravity** with zero coding.
- ðĻ **Minimalist Dual-Panel Workspace**: Live side-by-side raw Markdown editor and rendered HTML preview with line, word, and character counters.
- ð **Blazing Fast**: Converts standard multi-page documents in milliseconds without requiring heavy GPUs or multi-gigabyte models.
---
## ð How to Use (3 Ways)
### 1. Live Web Application (Zero Install)
Visit **[pdftomarkdown.fyi](https://pdftomarkdown.fyi)**:
1. Drag and drop any PDF file.
2. Click **Convert to Markdown**.
3. Preview the rendered document side-by-side and click **Copy Markdown** or **Download .md**.
---
### 2. Zero-Install CLI (`npx`)
No installation required. Run directly in your terminal:
```bash
# Convert a local PDF file
npx pdftomarkdown research-paper.pdf -o paper.md
# Convert a Word document (.docx)
npx pdftomarkdown document.docx -o doc.md
# Convert a CSV table to GFM pipe tables
npx pdftomarkdown financials.csv -o financials.md
# Extract timestamped transcript from a YouTube video
npx pdftomarkdown "https://www.youtube.com/watch?v=dQw4w9WgXcQ" > video.md
# Convert with an automatic Table of Contents
npx pdftomarkdown manual.pdf --toc -o manual.md
# Pipe directly into an AI tool / LLM
cat quarterly-report.pdf | npx pdftomarkdown | llm "Summarize the key takeaways"
```
---
### 3. AI Agent & MCP Server (Claude / Cursor)
Equip your AI assistants with PDF to Markdown capabilities with **zero coding**.
#### Claude Desktop Configuration:
Add this entry to your `claude_desktop_config.json`:
```json
{
"mcpServers": {
"pdftomarkdown": {
"command": "npx",
"args": ["-y", "pdftomarkdown", "--mcp"]
}
}
}
```
Now you can prompt Claude:
> *"Read `whitepaper.pdf` and convert it to clean markdown with table of contents."*
---
## ð Benchmark & Comparison Matrix
| Capability | **pdftomarkdown** | **PSPDFKit** | **Marker** | **Microsoft MarkItDown** |
| :--- | :---: | :---: | :---: | :---: |
| **Instant Live Web App** | â
**Yes (pdftomarkdown.fyi)** | â No | â No | â No |
| **100% In-Browser Privacy** | â
**Yes (0 uploads)** | â No | â No | â No |
| **Zero-Install CLI (`npx`)** | â
**Yes** | â
Yes | â (Heavy Python) | â (Python pip) |
| **Live Rendered Preview** | â
**Yes (Side-by-side)** | â No | â No | â No |
| **Word (`.docx`) to MD** | â
**Yes** | â No | â No | â
Yes |
| **Excel / CSV to GFM Table** | â
**Yes** | â No | â No | â
Yes |
| **YouTube Video to Transcript** | â
**Yes** | â No | â No | â
Yes |
| **Model Context Protocol (MCP)** | â
**Yes (Built-in)** | â ïļ Partial | â No | â ïļ Python only |
| **Execution Speed** | ⥠**< 200 ms / page** | ⥠< 200 ms / page | ðĒ 5â20 s (GPU) | ⥠< 500 ms / page |
| **Hardware Overhead** | ðŠķ **Featherweight (JS)** | ðŠķ Featherweight | ð Giant (10GB+ PyTorch) | ð Python Env |
---
## ðïļ Supported Formats
| Format | Extension | Output Style |
| :--- | :--- | :--- |
| **PDF Documents** | `.pdf` | Headings (H1âH4), lists, GFM tables, links, code blocks |
| **Word Documents** | `.docx` | Headings, bold/italic, bullet lists, clean links |
| **Excel & CSV** | `.csv`, `.tsv`, `.xlsx` | GitHub Flavored Markdown (GFM) pipe tables |
| **YouTube URLs** | `https://youtube.com/...` | Timestamped transcript (`00:00`), description, and metadata |
| **Plain Text & Code** | `.txt`, `.json`, `.xml` | Formatted code fences & readable Markdown blocks |
---
## ð ïļ CLI Reference & Options
```bash
pdftomarkdown â The universal document to Markdown converter.
Usage:
npx pdftomarkdown <file-or-url> [options]
Options:
-o, --output <file> Write output directly to a file (.md)
--toc Automatically generate a Table of Contents
--no-tables Disable automatic GFM table detection
--mcp Start Model Context Protocol (MCP) server
-v, --version Show current version
-h, --help Show help documentation
```
---
## ðŧ Programmatic API (Node / TypeScript)
You can import and use the converter directly inside your Node.js or TypeScript projects:
```typescript
import { convertPdfToMarkdown, convertAnyToMarkdown } from 'pdftomarkdown';
// Convert a PDF buffer or file
const markdown = await convertPdfToMarkdown(pdfBuffer, {
includeToc: true,
detectTables: true,
preservePageBreaks: true,
});
console.log(markdown);
```
---
## ð Technical SEO & Performance
- **Primary Keywords**: PDF to Markdown, PDF to MD, PDF to Markdown converter, convert PDF to Markdown.
- **Secondary Keywords**: PDF to Markdown online, free PDF to Markdown converter, convert PDF file to Markdown, extract Markdown from PDF, Word to Markdown, YouTube to Markdown.
- **Technical Standards**: Semantic HTML5, Schema.org `WebSite`, `WebApplication`, and `FAQPage` JSON-LD schemas, sitemap.xml, robots.txt.
- **Zero Client Overhead**: No heavy UI frameworks on the parser engine, optimized WebAssembly worker threads.
---
## â Frequently Asked Questions (FAQ)
### What is PDF to Markdown conversion?
PDF to Markdown conversion parses the geometry, fonts, and text positions inside fixed-layout PDF files and translates them into plain-text Markdown (`.md`) format, preserving headings, lists, tables, and links.
### How do I convert PDF to Markdown for free?
You can use the live web app at [pdftomarkdown.fyi](https://pdftomarkdown.fyi) or run `npx pdftomarkdown input.pdf -o output.md` in your terminal. Both are 100% free with no registration or subscriptions.
### Is my document kept private?
Yes. When using the web app, all parsing occurs **locally in your browser's memory**. No files or text streams are ever sent to our servers.
### Can I convert tables from PDF to Markdown?
Yes. Our engine detects horizontally and vertically aligned text columns and formats them into standard GitHub Flavored Markdown (GFM) pipe tables.
### Can I use the output with ChatGPT, Claude, and LLMs?
Yes. Clean Markdown is the optimal input format for AI prompts and RAG vector databases, minimizing token usage while maintaining document hierarchy.
---
## ð License & Disclaimer
Distributed under the **MIT License**. See `LICENSE` for more information.
> **Disclaimer:** All product names, logos, and brands (including Microsoft, PSPDFKit/Nutrient, Marker, Claude, and ChatGPT) are property of their respective owners and are used here solely for identification, compatibility, and comparative purposes.
Copyright ÂĐ 2026 [PDF to Markdown](https://pdftomarkdown.fyi)
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues