Skip to main content
Glama

shuck-file

PyPI MCP Registry License: MIT

Cualquier archivo entra, Markdown sale: lee solo lo que importa.

shuck-file convierte documentos a Markdown limpio para agentes de IA y LLMs. Los archivos pequeños se emiten directamente; los archivos grandes devuelven un mapa del documento con resúmenes de secciones, recuentos de tokens y próximos pasos accionables, para que los agentes solo extraigan lo que necesitan.

¿Por qué shuck-file?

Los agentes de IA necesitan un puente que sea consciente del contexto:

  • Archivo pequeñoshuck report.docx → Markdown completo en stdout

  • Archivo grandeshuck report.docx → mapa del documento con secciones y opciones de extracción

  • Extracción específicashuck report.docx --sections s1,s3 → solo lo que necesitas

  • Búsquedashuck report.docx --grep "revenue" → encuentra sin leerlo todo

Related MCP server: mcp-document-converter

Formatos admitidos

Formato

Extensión

Biblioteca

Lo que se conserva

Word

.docx

python-docx

Encabezados, negrita/cursiva, listas, tablas

PDF

.pdf

pdfplumber

Contenido de texto, saltos de página

Excel

.xlsx

openpyxl

Todas las hojas como tablas Markdown

PowerPoint

.pptx

python-pptx

Títulos, texto, tablas, notas del orador

CSV

.csv

stdlib

Todas las filas/columnas como tabla

Instalación

Mediante pip (recomendado)

pip install shuck-file

Esto instala el comando CLI shuck y el servidor MCP.

Desde el código fuente

git clone https://github.com/Shan-Zhu/shuck-file.git
cd shuck-file
pip install -e .

Inicio rápido

# Convert a document
shuck report.docx

# Force full output (bypass map mode)
shuck large-report.pdf --all

# Search within a document
shuck report.pdf --grep "revenue"

Uso

Enrutamiento automático (predeterminado)

Los archivos pequeños se emiten directamente; los archivos grandes devuelven un mapa del documento.

# Small file → direct Markdown output
shuck document.pdf

# Large file → document map with sections table + next steps
shuck large-report.pdf

Opciones de extracción

# Force full output (bypass map mode)
shuck report.pdf --all

# Extract specific sections
shuck report.pdf --sections s1,s3

# Tables only
shuck report.pdf --tables-only

# Search within document
shuck report.pdf --grep "revenue"

# Token budget (smart compression)
shuck report.pdf --budget 4000

# Combinations work
shuck report.pdf --sections s2,s3 --budget 2000

Específico para Excel/CSV

# Column headers and types
shuck data.xlsx --schema-only

# Headers + first N rows
shuck data.xlsx --sample 5

Subcomandos para usuarios avanzados

# Force map mode (even on small files)
shuck probe document.docx

# Force full extraction (alias for --all)
shuck pull document.docx

Control de salida

# Write to file
shuck document.pdf -o output.md

# Write to directory (auto-named)
shuck document.pdf -d ./converted/

# Skip YAML frontmatter
shuck document.pdf --no-frontmatter

# List supported formats
shuck --formats

Salida del modo mapa

Cuando un archivo es grande, shuck devuelve un mapa del documento:

# Document Map: quarterly-report.pdf

**6 pages | ~12,400 tokens | 6 sections**

## Sections

| # | Title | Type | Tokens | Density |
|---|-------|------|--------|---------|
| s1 | Executive Summary | narrative | 450 | high |
| s2 | Q3 Financial Results | mixed | 2,800 | high |
| s3 | Revenue Breakdown | tabular | 3,200 | high |
| ...

## Next Steps

- `shuck quarterly-report.pdf --all` -- full document (~12,400 tokens)
- `shuck quarterly-report.pdf --sections s1,s2` -- high-density (~3,250 tokens)
- `shuck quarterly-report.pdf --grep "..."` -- search for keywords

Servidor MCP

shuck-file incluye un servidor MCP (Model Context Protocol), lo que lo hace disponible para cualquier herramienta de IA compatible con MCP.

Claude Code

claude mcp add shuck-file -- shuck-file

O añádelo al .mcp.json de tu proyecto:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Cursor

Añádelo a ~/.cursor/mcp.json:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Windsurf

Añádelo a tu configuración de MCP:

{
  "mcpServers": {
    "shuck-file": {
      "command": "shuck-file",
      "args": []
    }
  }
}

Cualquier cliente MCP

shuck-file se registra como servidor MCP mediante el punto de entrada mcp.servers. Herramientas expuestas:

  • shuck — Convierte un documento a Markdown con todas las opciones (mode, sections, grep, budget, etc.)

  • list_formats — Lista los formatos de documento admitidos.

Plugin de Claude Code

Instálalo como plugin de Claude Code para la habilidad /shuck:

claude plugin add /path/to/shuck-file

Arquitectura

src/shuck_file/
├── cli.py                # CLI entrypoint
├── server.py             # MCP Server (FastMCP)
├── core/
│   ├── router.py          # Auto-routing logic
│   ├── segmenter.py       # Document segmentation
│   ├── mapper.py          # Map mode renderer
│   ├── budget.py          # Smart compression
│   ├── grep.py            # In-document search
│   ├── frontmatter.py     # YAML frontmatter
│   └── models.py          # Data models
├── extractors/
│   ├── base.py            # Base extractor ABC
│   ├── docx_ext.py        # Word extractor
│   ├── pdf_ext.py         # PDF extractor
│   ├── xlsx_ext.py        # Excel extractor
│   ├── pptx_ext.py        # PowerPoint extractor
│   └── csv_ext.py         # CSV extractor
plugin/                    # Claude Code plugin wrapper
tests/
├── test_extractors.py
├── test_router.py
├── test_segmenter.py
├── test_budget.py
└── test_grep.py

Licencia

MIT

A
license - permissive license
Not graded
quality - not tested
D
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to search, deep-read, and build knowledge bases from Markdown, PDF, DOCX, and PPTX documents via MCP tools for retrieval, document navigation, and ingestion.
    70
    616
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides AI agents with comprehensive document parsing capabilities including PDF text extraction, OCR, HTML-to-markdown conversion, table extraction, and summarization, optimized for agent workflows.
    101
    MIT

View all related MCP servers

Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.

  • Persistent docs and memory for AI agents — read, write, organize & search a shared workspace.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Shan-Zhu/shuck-file'

If you have feedback or need assistance with the MCP directory API, please join our Discord server