Images to PDF
images_to_pdfCombine PNG/JPEG images (base64) into a PDF, one image per page. Returns a file_id and a ~1h URL.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| images | Yes | Base64-encoded PNG/JPEG images, one per page (max 20). |
images_to_pdfCombine PNG/JPEG images (base64) into a PDF, one image per page. Returns a file_id and a ~1h URL.
| Name | Required | Description | Default |
|---|---|---|---|
| images | Yes | Base64-encoded PNG/JPEG images, one per page (max 20). |
Changes observed during successful MCP inspections.
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions the return value (file_id and ~1h URL) but does not disclose whether input images or the generated PDF are stored, how long the file persists, or any side effects beyond creating a PDF. The annotations are not contradicted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two short sentences that directly convey the action and the return value. No unnecessary words or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity and single parameter, the description is largely complete: it explains the input, output, and page arrangement. It could mention whether the file_id can be used with other tools, but that is not essential for invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description for the single parameter is detailed: it specifies base64-encoded PNG/JPEG images, one per page, with a max of 20. This gives an agent clear guidance on the accepted input format and limits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool combines PNG/JPEG images into a PDF with one image per page, which is a specific verb and resource. It also distinguishes itself from sibling PDF tools by focusing on image input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied by the description, but it does not explicitly state when to choose this tool over alternatives like markdown_to_pdf or url_to_pdf, nor does it mention exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Add one secure layer between your agents and this server.