Skip to main content
Glama

shuck-file

PyPI MCP Registry License: MIT

Любой файл на входе — Markdown на выходе: читайте только то, что важно.

shuck-file преобразует документы в чистый Markdown для AI-агентов и LLM. Небольшие файлы выводятся напрямую; для больших файлов возвращается карта документа с краткими описаниями разделов, количеством токенов и практическими следующими шагами — чтобы агенты извлекали только то, что им нужно.

Зачем нужен shuck-file?

AI-агентам нужен мост, который учитывает контекст:

  • Небольшой файлshuck report.docx → полный Markdown в stdout

  • Большой файлshuck report.docx → карта документа с разделами и вариантами извлечения

  • Точечное извлечениеshuck report.docx --sections s1,s3 → только то, что нужно

  • Поискshuck report.docx --grep "revenue" → найти, не читая всё подряд

Related MCP server: mcp-document-converter

Поддерживаемые форматы

Формат

Расширение

Библиотека

Что сохраняется

Word

.docx

python-docx

Заголовки, жирный/курсив, списки, таблицы

PDF

.pdf

pdfplumber

Текстовое содержимое, разрывы страниц

Excel

.xlsx

openpyxl

Все листы в виде таблиц Markdown

PowerPoint

.pptx

python-pptx

Заголовки, текст, таблицы, заметки докладчика

CSV

.csv

stdlib

Все строки/столбцы в виде таблицы

Установка

Через pip (рекомендуется)

pip install shuck-file

Это устанавливает команду CLI shuck и сервер MCP.

Из исходного кода

git clone https://github.com/Shan-Zhu/shuck-file.git
cd shuck-file
pip install -e .

Быстрый старт

# Convert a document
shuck report.docx

# Force full output (bypass map mode)
shuck large-report.pdf --all

# Search within a document
shuck report.pdf --grep "revenue"

Использование

Автоматическая маршрутизация (по умолчанию)

Небольшие файлы выводятся напрямую, для больших файлов возвращается карта документа.

# Small file → direct Markdown output
shuck document.pdf

# Large file → document map with sections table + next steps
shuck large-report.pdf

Параметры извлечения

# Force full output (bypass map mode)
shuck report.pdf --all

# Extract specific sections
shuck report.pdf --sections s1,s3

# Tables only
shuck report.pdf --tables-only

# Search within document
shuck report.pdf --grep "revenue"

# Token budget (smart compression)
shuck report.pdf --budget 4000

# Combinations work
shuck report.pdf --sections s2,s3 --budget 2000

Специфика Excel/CSV

# Column headers and types
shuck data.xlsx --schema-only

# Headers + first N rows
shuck data.xlsx --sample 5

Дополнительные подкоманды для опытных пользователей

# Force map mode (even on small files)
shuck probe document.docx

# Force full extraction (alias for --all)
shuck pull document.docx

Управление выводом

# Write to file
shuck document.pdf -o output.md

# Write to directory (auto-named)
shuck document.pdf -d ./converted/

# Skip YAML frontmatter
shuck document.pdf --no-frontmatter

# List supported formats
shuck --formats

Режим карты документа

Когда файл большой, shuck возвращает карту документа:

# Document Map: quarterly-report.pdf

**6 pages | ~12,400 tokens | 6 sections**

## Sections

| # | Title | Type | Tokens | Density |
|---|-------|------|--------|---------|
| s1 | Executive Summary | narrative | 450 | high |
| s2 | Q3 Financial Results | mixed | 2,800 | high |
| s3 | Revenue Breakdown | tabular | 3,200 | high |
| ...

## Next Steps

- `shuck quarterly-report.pdf --all` -- full document (~12,400 tokens)
- `shuck quarterly-report.pdf --sections s1,s2` -- high-density (~3,250 tokens)
- `shuck quarterly-report.pdf --grep "..."` -- search for keywords

MCP-сервер

shuck-file включает сервер MCP (Model Context Protocol), что делает его доступным для любого AI-инструмента, совместимого с MCP.

Claude Code

claude mcp add shuck-file -- shuck-file

Или добавьте в .mcp.json вашего проекта:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Cursor

Добавьте в ~/.cursor/mcp.json:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Windsurf

Добавьте в вашу конфигурацию MCP:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Любой MCP-клиент

shuck-file регистрируется как MCP-сервер через точку входа mcp.servers. Доступные инструменты:

  • shuck — Преобразование документа в Markdown со всеми параметрами (режим, разделы, grep, бюджет и т. д.)

  • list_formats — Список поддерживаемых форматов документов

Плагин для Claude Code

Установите как плагин для Claude Code, чтобы получить навык /shuck:

claude plugin add /path/to/shuck-file

Архитектура

src/shuck_file/
├── cli.py                # CLI entrypoint
├── server.py             # MCP Server (FastMCP)
├── core/
│   ├── router.py          # Auto-routing logic
│   ├── segmenter.py       # Document segmentation
│   ├── mapper.py          # Map mode renderer
│   ├── budget.py          # Smart compression
│   ├── grep.py            # In-document search
│   ├── frontmatter.py     # YAML frontmatter
│   └── models.py          # Data models
├── extractors/
│   ├── base.py            # Base extractor ABC
│   ├── docx_ext.py        # Word extractor
│   ├── pdf_ext.py         # PDF extractor
│   ├── xlsx_ext.py        # Excel extractor
│   ├── pptx_ext.py        # PowerPoint extractor
│   └── csv_ext.py         # CSV extractor
plugin/                    # Claude Code plugin wrapper
tests/
├── test_extractors.py
├── test_router.py
├── test_segmenter.py
├── test_budget.py
└── test_grep.py

Лицензия

MIT

A
license - permissive license
Not graded
quality - not tested
D
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to search, deep-read, and build knowledge bases from Markdown, PDF, DOCX, and PPTX documents via MCP tools for retrieval, document navigation, and ingestion.
    70
    616
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides AI agents with comprehensive document parsing capabilities including PDF text extraction, OCR, HTML-to-markdown conversion, table extraction, and summarization, optimized for agent workflows.
    101
    MIT

View all related MCP servers

Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.

  • Persistent docs and memory for AI agents — read, write, organize & search a shared workspace.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Shan-Zhu/shuck-file'

If you have feedback or need assistance with the MCP directory API, please join our Discord server