Skip to main content
Glama

pptx-mcp-server

Servidor MCP (stdio) que converte um arquivo HTML — escrito inteiramente pela LLM chamadora — em um PowerPoint (.pptx) nativo e editável. O servidor nunca gera nem embute nenhuma imagem/screenshot do slide inteiro; ele só usa um Chromium headless para medir o layout real (posições, fontes, cores) do HTML e converte o que foi marcado em objetos de verdade do PPTX (caixas de texto, formas, imagens).

Como funciona

  1. A LLM escreve tudo. Todo o design, conteúdo e HTML/CSS (e imagens locais, se houver) são escritos pela LLM chamadora diretamente no workspace do usuário — o servidor não desenha, não escolhe layout, não tem opinião de design. Ele só converte.

  2. Contrato de marcação. Cada slide é um elemento HTML com data-pptx-slide; dentro dele, qualquer elemento com data-pptx="text", data-pptx="shape" ou data-pptx="image" vira um objeto real no .pptx. Tudo que não tem data-pptx é só andaime de layout (divs de flexbox, grid, etc.) e é ignorado na conversão. Veja a tool get_pptx_authoring_guide para o contrato completo, com exemplo.

  3. Conversão. O servidor abre o HTML num Chromium headless (só para ler getBoundingClientRect()/getComputedStyle() de cada elemento marcado — nenhum screenshot é tirado) e usa esses dados para montar o .pptx via pptxgenjs: texto vira caixa de texto editável, formas viram retângulos/retângulos arredondados com fill/borda reais, imagens viram objetos de imagem nativos.

Related MCP server: PPTX Generator MCP Server

Tools expostas

  • get_pptx_authoring_guide() — devolve o guia completo de como escrever o HTML (o contrato data-pptx-slide/data-pptx, o que cada tipo lê de CSS, o que não é suportado, e orientações de design não-restritivas). Chame antes de escrever qualquer HTML — a conversão depende desse contrato; HTML sem essas marcações vira um .pptx vazio.

  • convert_html_to_pptx({ htmlPath, outputPath? })a tool que entrega o arquivo. Recebe o caminho absoluto de um único arquivo HTML já escrito pela LLM, extrai os elementos marcados e grava o .pptx. Devolve o caminho salvo e, se houver, avisos sobre elementos que não converteram bem (tamanho zero, imagem não encontrada, slide com tamanho diferente do primeiro).

Resolução de outputPath

  1. Caminho absoluto informado → usado diretamente.

  2. Caminho relativo ou omitido → tenta a primeira root declarada pelo cliente MCP (se ele suportar o recurso roots do protocolo).

  3. Se o cliente não suportar roots → usa a variável de ambiente PPT_MCP_OUTPUT_DIR, se definida, ou ~/Documents/PPT-MCP como último fallback.

Uso

npm install   # também baixa o Chromium do Playwright (postinstall)
npm run build
npm start     # inicia o servidor MCP via stdio

Configuração nos principais harnesses/clientes MCP

O pacote está publicado no npm como pptx-mcp-server, então na maioria dos clientes basta apontar command: npx, args: ["-y", "pptx-mcp-server"]. A env var PPT_MCP_OUTPUT_DIR é opcional (ver Resolução de outputPath); troque C:\caminho\padrao\de\saida pela pasta que preferir.

Claude Code

claude mcp add pptx-mcp-server -e PPT_MCP_OUTPUT_DIR=C:\caminho\padrao\de\saida -- npx -y pptx-mcp-server

Ou editando .mcp.json (projeto) / config de usuário diretamente — mesmo formato JSON do bloco "Claude Desktop / Cursor" abaixo.

Claude Desktop

Edite claude_desktop_config.json:

{
  "mcpServers": {
    "pptx-mcp-server": {
      "command": "npx",
      "args": ["-y", "pptx-mcp-server"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

Cursor

Mesmo formato do Claude Desktop, em ~/.cursor/mcp.json (global) ou .cursor/mcp.json (projeto):

{
  "mcpServers": {
    "pptx-mcp-server": {
      "command": "npx",
      "args": ["-y", "pptx-mcp-server"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

OpenAI Codex CLI

Adicione em ~/.codex/config.toml (ou .codex/config.toml do projeto):

[mcp_servers.pptx-mcp-server]
command = "npx"
args = ["-y", "pptx-mcp-server"]
env = { PPT_MCP_OUTPUT_DIR = "C:\\caminho\\padrao\\de\\saida" }

Google Antigravity (CLI agy / IDE)

Antigravity 2.0, a IDE e a CLI compartilham a mesma config, em ~/.gemini/config/mcp_config.json (ou .agents/mcp_config.json no workspace). Mesmo formato mcpServers/command/args de cima. Reinicie o Antigravity depois de editar.

VS Code (GitHub Copilot Chat, modo agente)

Crie/edite .vscode/mcp.json no projeto — repare que o VS Code usa a chave servers (não mcpServers) e exige "type": "stdio":

{
  "servers": {
    "pptx-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "pptx-mcp-server"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

GitHub Copilot CLI

Via wizard interativo (/mcp add dentro do copilot) ou editando ~/.copilot/mcp-config.json:

{
  "mcpServers": {
    "pptx-mcp-server": {
      "type": "local",
      "command": "npx",
      "args": ["-y", "pptx-mcp-server"],
      "tools": ["*"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

Windsurf

Mesmo formato mcpServers/command/args, em ~/.codeium/windsurf/mcp_config.json.

Zed

Em settings.json (~/.config/zed/settings.json no macOS/Linux, %APPDATA%\Zed\settings.json no Windows), a chave é context_servers e servidores customizados precisam de "source": "custom":

{
  "context_servers": {
    "pptx-mcp-server": {
      "source": "custom",
      "command": "npx",
      "args": ["-y", "pptx-mcp-server"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

Outros clientes

Quase todo cliente MCP stdio segue o mesmo formato básico — uma entrada com command: "npx" e args: ["-y", "pptx-mcp-server"], variando só a chave de agrupamento (mcpServers, servers, context_servers) e o caminho do arquivo de config. Consulte a documentação do cliente específico se ele não estiver nesta lista.

Rodando localmente (sem publicar/instalar via npm)

Durante desenvolvimento, ou se preferir não depender do registro npm, aponte command direto pro build local em vez de npx:

{
  "mcpServers": {
    "pptx-mcp-server": {
      "command": "node",
      "args": ["C:\\Projetos\\mcps\\ppt\\dist\\server.js"],
      "env": {
        "PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
      }
    }
  }
}

Available Tools

2 tools
build_pptxBuild PPTX (creates the actual presentation file)A

THE DELIVERABLE TOOL. Renders a list of complete HTML documents (one per slide) and assembles them into a single PowerPoint (.pptx) file written to disk, one image-backed slide per HTML input, in order. Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck; the tool result gives you the saved file path to report back to the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoSlide width in px (default 1280)
heightNoSlide height in px (default 720)
slidesYesArray of complete, self-contained HTML documents, one per slide, in order
outputPathNoWhere to save the .pptx. Absolute path recommended. If a bare filename or nothing is given, the file is saved in the client's workspace root (if the client shares one) or in a configured default folder otherwise.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. It clearly states that the tool writes a file to disk, produces one image-backed slide per HTML input, preserves input order, and returns the saved file path. It could go further by noting overwrite behavior or whether the output directory must exist, but the core side effects and output behavior are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loads the essence with 'THE DELIVERABLE TOOL,' though that opening phrase is somewhat redundant given the title already says 'creates the actual presentation file.' The rest is efficient: it explains the mechanism, the ordering, the distinction from preview_slide, and the result to report.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a file-generating tool with no output schema, the description covers the key contextual needs: what input is required, how it is processed, that a file is saved to disk, and that the serialized output is the file path. The schema handles parameter details, and the sibling is addressed explicitly. It could mention edge cases or constraints around HTML rendering, but the definition is largely complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents width, height, slides, and outputPath adequately. The description adds some nuance by emphasizing that each HTML document becomes 'one image-backed slide' in the given order, but it does little to clarify parameter formats beyond what the schema provides. This is the baseline 3 expected when the schema is the primary source of parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Renders a list of complete HTML documents... and assembles them into a single PowerPoint (.pptx) file written to disk.' It also distinguishes itself from its sibling by saying 'Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck.' This leaves no ambiguity about what the tool produces or how it differs from preview_slide.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool versus the alternative: 'Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck.' It also explains what to do with the result: 'the tool result gives you the saved file path to report back to the user.' This gives the agent both a trigger and a follow-up action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_slidePreview slide (design check only — does NOT save a presentation)A

DESIGN-ITERATION TOOL ONLY. Renders a single complete HTML document to a PNG image and saves it to disk, returning the file path (not the image data). Open/view that PNG file with whatever image-reading capability your agent has to check the design — this tool does NOT produce a .pptx and is NOT how a presentation gets delivered. Use it only to check how one slide looks while you're still tweaking its HTML/CSS. Once the user's slides are finalized, you MUST call build_pptx to actually produce the .pptx file — that is the only tool that creates a deliverable presentation file.

ParametersJSON Schema
NameRequiredDescriptionDefault
htmlYesComplete, self-contained HTML document (with inline CSS in a style block) for the slide
widthNoSlide width in px (default 1280)
heightNoSlide height in px (default 720)
outputPathNoWhere to save the preview PNG. Absolute path recommended. If a bare filename or nothing is given, it's saved in the client's workspace root (if the client shares one) or in a configured default folder otherwise.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does so thoroughly: it discloses the side effect ('saves it to disk'), the return contract ('file path (not the image data)'), the required follow-up action ('Open/view that PNG file with whatever image-reading capability your agent has'), and the hard boundary ('is NOT how a presentation gets delivered'). No behavioral trait an agent needs is left to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the scoping tag 'DESIGN-ITERATION TOOL ONLY' and the core mechanism in the first sentence, with each subsequent sentence adding operational guidance. However, the non-deliverable warning is asserted three times ('does NOT produce a .pptx', 'NOT how a presentation gets delivered', 'the only tool that creates a deliverable'), so a few words are redundant without adding new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema and no annotations, the description covers every operational fact an agent needs: input form (complete self-contained HTML), output contract (file path, not image data), side effect (PNG written to disk), the agent's next step (view the image), and the workflow handoff to build_pptx. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's 'single complete HTML document (with inline CSS)' loosely echoes the schema's documented html parameter but adds no syntax or format details beyond it, and the other three parameters are left entirely to the schema. The description does not compensate further, but with full schema coverage, it doesn't need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Renders a single complete HTML document to a PNG image and saves it to disk, returning the file path' — and immediately distinguishes itself from the sibling tool: 'does NOT produce a .pptx and is NOT how a presentation gets delivered.' The title's 'design check only' reinforcement makes the scope unmistakable. An agent cannot confuse this with build_pptx without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes when to use it: 'Use it only to check how one slide looks while you're still tweaking its HTML/CSS.' It also gives the when-not condition and the alternative: 'Once the user's slides are finalized, you MUST call build_pptx to actually produce the .pptx file.' This is direct routing guidance with an explicit exclusion, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.5/5.0
Disambiguation5/5

preview_slide is explicitly scoped to single-slide PNG design checks, while build_pptx is the only tool that produces the deliverable PPTX. The descriptions draw a clear boundary and even warn against using preview_slide for final delivery.

Naming Consistency5/5

Both tools follow the same verb_noun pattern: preview_slide and build_pptx. The naming is consistent, predictable, and each verb clearly indicates the action.

Tool Count4/5

Two tools is slightly lean relative to the typical 3-15 range, but the server's scope is deliberately narrow: preview a slide and build a deck. Each tool earns its place and there is no redundancy.

Completeness5/5

The workflow is complete for the stated purpose: preview_slide covers design iteration and build_pptx covers final deliverable creation from a list of HTML slides. There are no obvious missing operations in this focused domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/RAFAEL-SILVASOUZA/pptx-mcp-server'

If you have feedback or need assistance with the MCP directory API, please join our Discord server