pptx-mcp-server
This server lets you create PowerPoint presentations from HTML: preview slide designs and build .pptx files from HTML slide documents.
preview_slide: Renders a single complete HTML document to a PNG file on disk for design checking only; it does not save a presentation.
build_pptx: Takes an ordered list of complete HTML documents (one per slide) and assembles them into a single .pptx file — this is the tool to use for actual PPT/PPTX delivery.
Native, editable PowerPoint objects (per README): Elements marked with
data-pptx="text|shape|image"are converted into real PPTX text boxes, shapes, and images; headless Chromium is used only to measure layout, not to embed screenshots.Authoring guide: The server exposes a tool that returns the full HTML markup contract needed for successful conversion (avoiding empty decks).
Configurable output: You can set slide width/height and choose where the output
.pptxor preview PNG is saved, with fallback to workspace root or a default directory.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pptx-mcp-serverBuild a 3-slide presentation from this HTML and save it as pitch.pptx"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
pptx-mcp-server
Servidor MCP (stdio) que converte um arquivo HTML — escrito inteiramente pela LLM chamadora —
em um PowerPoint (.pptx) nativo e editável. O servidor nunca gera nem embute nenhuma
imagem/screenshot do slide inteiro; ele só usa um Chromium headless para medir o layout real
(posições, fontes, cores) do HTML e converte o que foi marcado em objetos de verdade do PPTX
(caixas de texto, formas, imagens).
Como funciona
A LLM escreve tudo. Todo o design, conteúdo e HTML/CSS (e imagens locais, se houver) são escritos pela LLM chamadora diretamente no workspace do usuário — o servidor não desenha, não escolhe layout, não tem opinião de design. Ele só converte.
Contrato de marcação. Cada slide é um elemento HTML com
data-pptx-slide; dentro dele, qualquer elemento comdata-pptx="text",data-pptx="shape"oudata-pptx="image"vira um objeto real no.pptx. Tudo que não temdata-pptxé só andaime de layout (divs de flexbox, grid, etc.) e é ignorado na conversão. Veja a toolget_pptx_authoring_guidepara o contrato completo, com exemplo.Conversão. O servidor abre o HTML num Chromium headless (só para ler
getBoundingClientRect()/getComputedStyle()de cada elemento marcado — nenhum screenshot é tirado) e usa esses dados para montar o.pptxviapptxgenjs: texto vira caixa de texto editável, formas viram retângulos/retângulos arredondados com fill/borda reais, imagens viram objetos de imagem nativos.
Related MCP server: PPTX Generator MCP Server
Tools expostas
get_pptx_authoring_guide()— devolve o guia completo de como escrever o HTML (o contratodata-pptx-slide/data-pptx, o que cada tipo lê de CSS, o que não é suportado, e orientações de design não-restritivas). Chame antes de escrever qualquer HTML — a conversão depende desse contrato; HTML sem essas marcações vira um.pptxvazio.convert_html_to_pptx({ htmlPath, outputPath? })— a tool que entrega o arquivo. Recebe o caminho absoluto de um único arquivo HTML já escrito pela LLM, extrai os elementos marcados e grava o.pptx. Devolve o caminho salvo e, se houver, avisos sobre elementos que não converteram bem (tamanho zero, imagem não encontrada, slide com tamanho diferente do primeiro).
Resolução de outputPath
Caminho absoluto informado → usado diretamente.
Caminho relativo ou omitido → tenta a primeira
rootdeclarada pelo cliente MCP (se ele suportar o recursorootsdo protocolo).Se o cliente não suportar
roots→ usa a variável de ambientePPT_MCP_OUTPUT_DIR, se definida, ou~/Documents/PPT-MCPcomo último fallback.
Uso
npm install # também baixa o Chromium do Playwright (postinstall)
npm run build
npm start # inicia o servidor MCP via stdioConfiguração nos principais harnesses/clientes MCP
O pacote está publicado no npm como pptx-mcp-server,
então na maioria dos clientes basta apontar command: npx, args: ["-y", "pptx-mcp-server"]. A
env var PPT_MCP_OUTPUT_DIR é opcional (ver Resolução de outputPath);
troque C:\caminho\padrao\de\saida pela pasta que preferir.
Claude Code
claude mcp add pptx-mcp-server -e PPT_MCP_OUTPUT_DIR=C:\caminho\padrao\de\saida -- npx -y pptx-mcp-serverOu editando .mcp.json (projeto) / config de usuário diretamente — mesmo formato JSON do bloco
"Claude Desktop / Cursor" abaixo.
Claude Desktop
Edite claude_desktop_config.json:
{
"mcpServers": {
"pptx-mcp-server": {
"command": "npx",
"args": ["-y", "pptx-mcp-server"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}Cursor
Mesmo formato do Claude Desktop, em ~/.cursor/mcp.json (global) ou .cursor/mcp.json (projeto):
{
"mcpServers": {
"pptx-mcp-server": {
"command": "npx",
"args": ["-y", "pptx-mcp-server"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}OpenAI Codex CLI
Adicione em ~/.codex/config.toml (ou .codex/config.toml do projeto):
[mcp_servers.pptx-mcp-server]
command = "npx"
args = ["-y", "pptx-mcp-server"]
env = { PPT_MCP_OUTPUT_DIR = "C:\\caminho\\padrao\\de\\saida" }Google Antigravity (CLI agy / IDE)
Antigravity 2.0, a IDE e a CLI compartilham a mesma config, em
~/.gemini/config/mcp_config.json (ou .agents/mcp_config.json no workspace). Mesmo formato
mcpServers/command/args de cima. Reinicie o Antigravity depois de editar.
VS Code (GitHub Copilot Chat, modo agente)
Crie/edite .vscode/mcp.json no projeto — repare que o VS Code usa a chave servers (não
mcpServers) e exige "type": "stdio":
{
"servers": {
"pptx-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "pptx-mcp-server"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}GitHub Copilot CLI
Via wizard interativo (/mcp add dentro do copilot) ou editando ~/.copilot/mcp-config.json:
{
"mcpServers": {
"pptx-mcp-server": {
"type": "local",
"command": "npx",
"args": ["-y", "pptx-mcp-server"],
"tools": ["*"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}Windsurf
Mesmo formato mcpServers/command/args, em ~/.codeium/windsurf/mcp_config.json.
Zed
Em settings.json (~/.config/zed/settings.json no macOS/Linux, %APPDATA%\Zed\settings.json
no Windows), a chave é context_servers e servidores customizados precisam de "source": "custom":
{
"context_servers": {
"pptx-mcp-server": {
"source": "custom",
"command": "npx",
"args": ["-y", "pptx-mcp-server"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}Outros clientes
Quase todo cliente MCP stdio segue o mesmo formato básico — uma entrada com command: "npx" e
args: ["-y", "pptx-mcp-server"], variando só a chave de agrupamento (mcpServers, servers,
context_servers) e o caminho do arquivo de config. Consulte a documentação do cliente específico
se ele não estiver nesta lista.
Rodando localmente (sem publicar/instalar via npm)
Durante desenvolvimento, ou se preferir não depender do registro npm, aponte command direto pro
build local em vez de npx:
{
"mcpServers": {
"pptx-mcp-server": {
"command": "node",
"args": ["C:\\Projetos\\mcps\\ppt\\dist\\server.js"],
"env": {
"PPT_MCP_OUTPUT_DIR": "C:\\caminho\\padrao\\de\\saida"
}
}
}
}Available Tools
2 toolsbuild_pptxBuild PPTX (creates the actual presentation file)A
THE DELIVERABLE TOOL. Renders a list of complete HTML documents (one per slide) and assembles them into a single PowerPoint (.pptx) file written to disk, one image-backed slide per HTML input, in order. Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck; the tool result gives you the saved file path to report back to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Slide width in px (default 1280) | |
| height | No | Slide height in px (default 720) | |
| slides | Yes | Array of complete, self-contained HTML documents, one per slide, in order | |
| outputPath | No | Where to save the .pptx. Absolute path recommended. If a bare filename or nothing is given, the file is saved in the client's workspace root (if the client shares one) or in a configured default folder otherwise. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral disclosure. It clearly states that the tool writes a file to disk, produces one image-backed slide per HTML input, preserves input order, and returns the saved file path. It could go further by noting overwrite behavior or whether the output directory must exist, but the core side effects and output behavior are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the essence with 'THE DELIVERABLE TOOL,' though that opening phrase is somewhat redundant given the title already says 'creates the actual presentation file.' The rest is efficient: it explains the mechanism, the ordering, the distinction from preview_slide, and the result to report.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a file-generating tool with no output schema, the description covers the key contextual needs: what input is required, how it is processed, that a file is saved to disk, and that the serialized output is the file path. The schema handles parameter details, and the sibling is addressed explicitly. It could mention edge cases or constraints around HTML rendering, but the definition is largely complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents width, height, slides, and outputPath adequately. The description adds some nuance by emphasizing that each HTML document becomes 'one image-backed slide' in the given order, but it does little to clarify parameter formats beyond what the schema provides. This is the baseline 3 expected when the schema is the primary source of parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Renders a list of complete HTML documents... and assembles them into a single PowerPoint (.pptx) file written to disk.' It also distinguishes itself from its sibling by saying 'Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck.' This leaves no ambiguity about what the tool produces or how it differs from preview_slide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool versus the alternative: 'Call this — not preview_slide — whenever the user asks for a PPT/PowerPoint/presentation/apresentação/slide deck.' It also explains what to do with the result: 'the tool result gives you the saved file path to report back to the user.' This gives the agent both a trigger and a follow-up action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_slidePreview slide (design check only — does NOT save a presentation)A
DESIGN-ITERATION TOOL ONLY. Renders a single complete HTML document to a PNG image and saves it to disk, returning the file path (not the image data). Open/view that PNG file with whatever image-reading capability your agent has to check the design — this tool does NOT produce a .pptx and is NOT how a presentation gets delivered. Use it only to check how one slide looks while you're still tweaking its HTML/CSS. Once the user's slides are finalized, you MUST call build_pptx to actually produce the .pptx file — that is the only tool that creates a deliverable presentation file.
| Name | Required | Description | Default |
|---|---|---|---|
| html | Yes | Complete, self-contained HTML document (with inline CSS in a style block) for the slide | |
| width | No | Slide width in px (default 1280) | |
| height | No | Slide height in px (default 720) | |
| outputPath | No | Where to save the preview PNG. Absolute path recommended. If a bare filename or nothing is given, it's saved in the client's workspace root (if the client shares one) or in a configured default folder otherwise. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does so thoroughly: it discloses the side effect ('saves it to disk'), the return contract ('file path (not the image data)'), the required follow-up action ('Open/view that PNG file with whatever image-reading capability your agent has'), and the hard boundary ('is NOT how a presentation gets delivered'). No behavioral trait an agent needs is left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the scoping tag 'DESIGN-ITERATION TOOL ONLY' and the core mechanism in the first sentence, with each subsequent sentence adding operational guidance. However, the non-deliverable warning is asserted three times ('does NOT produce a .pptx', 'NOT how a presentation gets delivered', 'the only tool that creates a deliverable'), so a few words are redundant without adding new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema and no annotations, the description covers every operational fact an agent needs: input form (complete self-contained HTML), output contract (file path, not image data), side effect (PNG written to disk), the agent's next step (view the image), and the workflow handoff to build_pptx. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's 'single complete HTML document (with inline CSS)' loosely echoes the schema's documented html parameter but adds no syntax or format details beyond it, and the other three parameters are left entirely to the schema. The description does not compensate further, but with full schema coverage, it doesn't need to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Renders a single complete HTML document to a PNG image and saves it to disk, returning the file path' — and immediately distinguishes itself from the sibling tool: 'does NOT produce a .pptx and is NOT how a presentation gets delivered.' The title's 'design check only' reinforcement makes the scope unmistakable. An agent cannot confuse this with build_pptx without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes when to use it: 'Use it only to check how one slide looks while you're still tweaking its HTML/CSS.' It also gives the when-not condition and the alternative: 'Once the user's slides are finalized, you MUST call build_pptx to actually produce the .pptx file.' This is direct routing guidance with an explicit exclusion, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
preview_slide is explicitly scoped to single-slide PNG design checks, while build_pptx is the only tool that produces the deliverable PPTX. The descriptions draw a clear boundary and even warn against using preview_slide for final delivery.
Both tools follow the same verb_noun pattern: preview_slide and build_pptx. The naming is consistent, predictable, and each verb clearly indicates the action.
Two tools is slightly lean relative to the typical 3-15 range, but the server's scope is deliberately narrow: preview a slide and build a deck. Each tool earns its place and there is no redundancy.
The workflow is complete for the stated purpose: preview_slide covers design iteration and build_pptx covers final deliverable creation from a list of HTML slides. There are no obvious missing operations in this focused domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate, edit, merge, translate and PDF-convert PowerPoint (.pptx) over MCP. 8 tools.
Generate professional PowerPoint presentations from text, YouTube videos, or structured JSON data.…
Generate, edit, and export AI presentations to PDF, PPTX, or a shareable link.
AI presentation and report generation: slides, diagrams, PPTX export, live preview MCP App.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceCreates professional PowerPoint presentations from Markdown or JSON with intelligent layout recommendations, rich content support including tables and images, and automatic template selection based on content analysis.7Apache 2.0
- FlicenseAqualityDmaintenanceGenerates professional PowerPoint presentations from Markdown with support for code blocks, tables, custom branding, and mixed formatting. Transforms lesson plans and documentation into styled PPTX files with syntax highlighting and customizable themes.59
- AlicenseAqualityDmaintenanceConvert HTML to PDF/PNG/WebP/PPTX slide carousels with 11 themes. For LinkedIn carousels, decks, Instagram posts, and infographics — Puppeteer-based pixel-perfect rendering.62MIT
- FlicenseNot gradedqualityCmaintenanceConverts slide-oriented HTML into editable PowerPoint files via an MCP server, using Chromium for layout measurement and supporting editable text, shapes, images, and raster fallbacks.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/RAFAEL-SILVASOUZA/pptx-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server