Skip to main content
Glama

🎨 pruna-mcp-server

CI PyPI Python License: MIT

Pruna AI용 MCP 서버 — AI 어시스턴트에서 직접 초고속 이미지 생성, 편집, 업스케일링 및 비디오 생성을 수행합니다.

Pruna AI는 이미지 및 비디오 생성에 특화된 추론 API입니다. 이미지당 $0.005부터 시작하는 2초 미만의 이미지 생성 기능을 제공하며, 텍스트-이미지 변환, 이미지 편집, 업스케일링 및 비디오 생성 모델을 지원합니다. 이 MCP 서버는 해당 API를 래핑하여 MCP 호환 클라이언트(Claude Desktop, Kiro, Cursor)가 기본적으로 시각적 콘텐츠를 생성할 수 있도록 합니다.

MCP 사양 2025-11-25를 준수합니다.

주요 기능

  • 6개의 MCP 도구: generate_image, edit_image, upscale_image, generate_video, list_models, upload_file

  • 7개의 MCP 프롬프트: 제품 사진, 가상 스테이징, 소셜 미디어 비주얼, 게임 컨셉 아트, 광고 크리에이티브, 비디오 광고, 이미지 향상

  • 2개의 MCP 리소스: 도구 호출 없이 모델을 탐색할 수 있는 pruna://models 카탈로그

  • 18개의 모델: 텍스트-이미지 10개, 편집 3개, 업스케일 1개, 비디오 4개

  • 스마트 동기/비동기: 빠른 이미지 모델은 동기식, 비디오는 폴링을 통한 비동기식 처리

  • 투명한 파일 처리: 로컬 경로 또는 URL 전달 시 자동 업로드 처리

  • 기본 MCP 이미지 반환: 인라인 표시를 지원하는 클라이언트를 위한 ImageContent 블록

  • 완벽한 MCP 준수: 도구 주석, 구조화된 콘텐츠, 진행 상황 알림

Related MCP server: jgkme/kilo-image-gen-mcp

빠른 시작

# With uvx (zero install)
uvx pruna-mcp-server

# Or with pip
pip install pruna-mcp-server
pruna-mcp

API 키를 설정하세요. pruna.ai에서 키를 발급받을 수 있습니다(개발자 포털로 이동하거나 Pruna에 문의하여 액세스 권한을 요청하세요).

# macOS Keychain (recommended)
security add-generic-password -a $USER -s PRUNA_API_KEY -w "your-api-key"

# Or environment variable
export PRUNA_API_KEY="your-api-key"

MCP 클라이언트 구성

Kiro CLI

에이전트 구성(예: ~/.kiro/agents/default.json)에 추가하세요:

mcpServers 내:

"pruna": {
  "command": "sh",
  "args": ["-c", "PRUNA_API_KEY=$(security find-generic-password -a $USER -s PRUNA_API_KEY -w) uv run --directory /path/to/pruna-mcp-server pruna-mcp"],
  "autoApprove": ["generate_image", "edit_image", "upscale_image", "generate_video", "list_models", "upload_file"]
}

tools에 @pruna/*를 추가하세요.

allowedTools에 "generate_image", "edit_image", "upscale_image", "generate_video", "list_models", "upload_file"을 추가하세요.

참고: Kiro 에이전트는 @server-name/* 구문을 사용하는 tools 화이트리스트와 allowedTools 목록을 사용합니다. Pruna 도구를 사용하려면 두 곳 모두에 포함되어야 합니다.

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json(macOS)에 추가하세요:

{
  "mcpServers": {
    "pruna": {
      "command": "sh",
      "args": ["-c", "PRUNA_API_KEY=$(security find-generic-password -a $USER -s PRUNA_API_KEY -w) /path/to/uv run --directory /path/to/pruna-mcp-server pruna-mcp"]
    }
  }
}

중요: uv의 전체 경로(예: /Users/you/.local/bin/uv)를 사용하세요. Claude Desktop은 ~/.local/bin을 포함하지 않는 최소한의 PATH로 프로세스를 실행합니다.

참고: Claude Desktop은 채팅 내에서 ImageContent를 인라인으로 렌더링하지 않습니다. 이미지는 생성되어 로컬에 저장되며, Claude는 응답에 파일 경로를 참조합니다.

Cursor

.cursor/mcp.json에 추가하세요:

{
  "mcpServers": {
    "pruna": {
      "command": "uvx",
      "args": ["pruna-mcp-server"],
      "env": { "PRUNA_API_KEY": "your-api-key" }
    }
  }
}

도구

도구

설명

가격

generate_image

10개 모델을 사용한 텍스트-이미지 생성

이미지당 $0.0001부터

edit_image

텍스트 지침으로 1-5개의 이미지 편집

이미지당 $0.010부터

upscale_image

1-8 메가픽셀로 AI 업스케일링

이미지당 $0.005부터

generate_video

텍스트/이미지/오디오를 비디오로 변환

초당 $0.005부터

list_models

가격을 포함한 모든 사용 가능한 모델 탐색

무료

upload_file

편집/비디오 워크플로우를 위한 파일 업로드

무료

이미지 도구는 JSON 메타데이터 블록과 기본 MCP ImageContent 블록(5MB 미만 이미지의 경우 base64)을 모두 반환합니다.

프롬프트

일반적인 사용 사례를 위한 내장 워크플로우 템플릿:

프롬프트

사용 사례

예시

product-photo

이커머스 제품 사진

"깔끔한 배경의 흰색 가죽 스니커즈"

virtual-staging

부동산 방 스테이징

가구가 없는 방을 가구로 채우기

social-media-visual

플랫폼 최적화 비주얼

플랫폼별 자동 화면 비율

game-concept-art

게임 에셋 및 환경

캐릭터, 무기, 풍경

ad-creative

텍스트 오버레이가 포함된 디지털 광고

이미지 내에 렌더링된 헤드라인

video-ad

짧은 비디오 광고

토킹 헤드, 제품 데모

image-enhance

업스케일 + 향상 워크플로우

AI 생성 이미지 정교화

구성

환경 변수

필수

기본값

설명

PRUNA_API_KEY

✅

—

Pruna AI API 키

PRUNA_OUTPUT_DIR

—

./pruna-output

다운로드된 파일 저장 디렉토리

PRUNA_POLL_INTERVAL

—

2

비동기 폴링 간격(초)

PRUNA_TIMEOUT

—

120

HTTP 타임아웃(초)

PRUNA_MAX_RETRIES

—

3

일시적 오류 발생 시 최대 재시도 횟수

클라이언트 호환성

클라이언트

전송 방식

상태

참고

Kiro CLI

STDIO

✅ 테스트 완료

tools + allowedTools 구성 필요

Claude Desktop

STDIO

✅ 테스트 완료

uv 전체 경로 사용; 인라인 이미지 표시 불가

Cursor

STDIO

🔲 계획 중

—

Claude Code

STDIO

🔲 계획 중

—

개발

git clone https://github.com/charlesrapp/pruna-mcp-server.git
cd pruna-mcp-server
uv sync --extra dev

# Run tests (100 tests, 94% coverage)
uv run pytest --cov

# Lint & type check
uv run ruff check src/ tests/
uv run mypy src/

지침은 CONTRIBUTING.md를 참조하세요.

라이선스

MIT — LICENSE 참조.

Available Tools

8 tools
edit_imageA

Edit one or more images with text instructions using Pruna AI.

Args: prompt: Edit instruction describing the desired changes images: 1-5 image URLs or local file paths model: Model to use (default: p-image-edit) aspect_ratio: Output aspect ratio seed: Random seed for reproducible generation

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
modelNop-image-edit
imagesYes
promptYes
aspect_ratioNomatch_input_image

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate the tool is modifiable (readOnlyHint=false) and not destructive. The description adds that it uses 'Pruna AI' (external dependency) but does not elaborate on side effects, rate limits, or required permissions. Some value added beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, starting with a clear purpose sentence followed by a parameter list. Every sentence adds value, and the structure is easy to scan. No redundant or unnecessary text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core functionality and parameters but omits details like valid model options, aspect ratio formats, error handling, output format, and prerequisites. Given the complexity and lack of output schema, more completeness would help an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description lists each parameter with a brief explanation (e.g., 'prompt: Edit instruction describing the desired changes'). This adds meaning beyond the schema's bare titles, though some params like aspect_ratio receive minimal clarification.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool edits one or more images using text instructions, specifying the verb (edit) and resource (images) distinctly. This differentiates it from sibling tools like generate_image or generate_video, which create new content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when an existing image needs modification via a text prompt, and provides parameter docs like prompt and images. However, it does not explicitly state when not to use it or mention alternative tools for similar tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_imageA

Generate an image from a text prompt using Pruna AI.

Args: prompt: Text description of the image to generate model: Model to use (default: p-image) aspect_ratio: Output aspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, custom) width: Custom width 256-1440, multiple of 16. Only when aspect_ratio=custom height: Custom height 256-1440, multiple of 16. Only when aspect_ratio=custom seed: Random seed for reproducible generation

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
modelNop-image
widthNo
heightNo
promptYes
aspect_ratioNo16:9

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, so the description's addition of constraints (e.g., dimension multiples) and seed for reproducibility adds some value. However, it does not disclose other behavioral traits like rate limits, authentication, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using a docstring format with a clear one-line summary followed by an organized Args list. Every sentence adds value—no redundancy or irrelevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains input parameters well but does not describe the output format (e.g., image URL, base64). Given no output schema, this omission could leave an agent uncertain about what is returned. Additionally, no usage context or error handling is mentioned.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully explains all six parameters. It defines defaults (model: p-image, aspect_ratio: 16:9), lists allowed aspect ratios, and specifies valid ranges and conditions for width/height. This provides essential meaning beyond the schema's type and title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action 'Generate an image from a text prompt' and specifies the tool (Pruna AI). The verb and resource are specific, and it naturally distinguishes from sibling tools like 'edit_image' and 'generate_video'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus alternatives (e.g., edit_image for modifications). It is implied that this is for generating new images from text, but no when-not-to-use or exclusion criteria are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_videoB

Generate a video from text, image, or audio using Pruna AI.

Args: prompt: Text prompt for video generation model: Model to use (p-video, wan-t2v, wan-i2v, vace) image: Input image URL/path for image-to-video audio: Input audio URL/path for audio-conditioned video duration: Duration in seconds (1-20) resolution: Video resolution (720p or 1080p) aspect_ratio: Aspect ratio (ignored when image is provided) fps: Frames per second (24 or 48) seed: Random seed for reproducible generation

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
seedNo
audioNo
imageNo
modelNop-video
promptYes
durationNo
resolutionNo720p
aspect_ratioNo16:9

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate non-read-only, open-world, non-idempotent, non-destructive behavior, but the description adds no behavioral context beyond listing parameters. There is no mention of costs, generation time, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the purpose, followed by a clear parameter list. However, it could be better structured with explicit defaults or grouped parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks mention of prerequisites (e.g., uploading image/audio files) and does not clarify behavior with multiple inputs. Since output schema exists, return values are not expected, but the tool's complexity warrants more completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description adds meaning by listing parameters and brief explanations (e.g., 'prompt: Text prompt for video generation'), but it does not provide constraints or allowed values for parameters like model or resolution.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Generate a video from text, image, or audio using Pruna AI,' specifying the verb and resource, and implicitly distinguishes from sibling tools like generate_image and transform_video.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives such as generate_image or transform_video. The description lists inputs but does not provide decision criteria or mention prerequisites like file uploads.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA
Read-onlyIdempotent

List available Pruna AI models with capabilities and pricing.

Args: category: Filter by category: image, editing, try-on, upscale, video, video-edit

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnlyHint, idempotentHint, destructiveHint), the description adds that the tool returns 'capabilities and pricing' and supports filtering. It does not describe pagination or rate limits, but for a simple list operation this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with a clear main sentence and a docstring-style args section. It could be slightly more structured, but it efficiently conveys the purpose and filter option.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional parameter, output schema exists), the description covers its purpose and filter. It does not mention pagination or output structure, but the output schema likely handles that. It is complete enough for a listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning to the 'category' parameter by listing the allowed values: image, editing, try-on, upscale, video, video-edit. The input schema only specifies type string/null with no enums, so the description compensates for low schema coverage (0%).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and resource 'Pruna AI models', and specifies it returns 'capabilities and pricing'. It distinguishes from sibling tools that perform actions (edit, generate, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. The description only mentions filtering by category, but does not discuss when listing is appropriate or when to prefer other tools. Sibling tools are all action-oriented, so the distinction is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transform_videoA

Transform a source video using reference images (video-to-video).

Two models are available:

  • p-video-animate: animate a single subject reference image using the motion from the source video (provide exactly 1 reference).

  • p-video-replace: replace the character(s) in the source video using 1-3 identity reference images.

Motion, timing, camera movement, and scene structure are preserved.

Args: video: Source video URL or local file path (.mp4) references: Reference images (URLs or local file paths). Exactly 1 for p-video-animate, 1-3 for p-video-replace. model: Model to use (p-video-animate or p-video-replace) resolution: Output resolution (720p or 1080p) target_fps: Working FPS (original, 24, or 48) instruction_prompt: Optional guidance on how to apply the transform turbo: Faster generation for slightly lower quality save_audio: Save the output video with audio ignore_audio: Ignore source audio during generation seed: Random seed for reproducible generation

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
modelNop-video-animate
turboNo
videoYes
referencesYes
resolutionNo720p
save_audioNo
target_fpsNooriginal
ignore_audioNo
instruction_promptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide minimal info (readOnlyHint=false, destructiveHint=false), so the description carries the burden. It reveals that motion, timing, etc. are preserved and that turbo mode offers faster generation at lower quality. However, it lacks disclosure of limitations (e.g., maximum video length, supported input formats) or side effects beyond the stated transformations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with a lead sentence, model breakdown, and a bullet list of parameters. It is slightly verbose but every sentence adds value. Could be condensed slightly, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters, no nested objects, and an existing output schema, the description covers all parameters and model usage. It lacks mention of prerequisites (e.g., pre-uploaded files), potential error conditions, or integration with sibling tools like upload_file. Nearly complete, with minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, leaving all parameter meaning to the description. The description provides thorough inline explanations for all 10 parameters, including details on reference count constraints per model and optional instruction prompt. This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies it is a video-to-video transform using reference images. It distinguishes two models (p-video-animate and p-video-replace) with different use cases, which differentiates it from sibling tools like generate_video or edit_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use each model based on number of references and what is preserved (motion, timing, etc.). However, it does not explicitly state when not to use this tool or mention alternatives for other scenarios, which would strengthen guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

try_on_imageA

Virtually fit one or more garments onto a person's photo using Pruna AI.

Args: person_image: Image URL or local file path of the person garment_images: 1-11 garment reference images (URLs or local file paths). Up to 6 recommended for best quality. model: Model to use (default: p-image-try-on) prompt: Experimental guidance for non-flatlay garment images (e.g. which garment from which image to use) turbo: Faster generation. Not recommended for more than 4 garments reference_pose: Experimental. Image URL/path to repose the person before try-on seed: Random seed for reproducible generation output_format: Output format (webp, jpg, png) output_quality: Quality for jpg/webp outputs (0-100) preserve_input_size: Resize the result back to the person image size

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
modelNop-image-try-on
turboNo
promptNo
person_imageYes
output_formatNojpg
garment_imagesYes
output_qualityNo
reference_poseNo
preserve_input_sizeNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide no behavioral hints (readOnlyHint=false, destructiveHint=false), so the description carries full burden. It discloses recommendations (up to 6 garments, turbo not for >4), experimental flags (prompt, reference_pose), and defaults, but omits details on failure modes, rate limits, or exact output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief and well-structured: a one-sentence summary followed by a bullet-like list of parameters. Every sentence adds value, with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite 10 parameters and no output schema, the description explains each parameter well and offers usage tips. However, it lacks an explicit description of what the tool returns (e.g., an image URL) and does not specify input format requirements (e.g., valid file types).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must fully explain parameters. It does so with clear explanations for all 10 parameters, including defaults, recommendations, and experimental notes—adding significant meaning beyond raw schema property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it virtually fits garments onto a person's photo using Pruna AI. The verb 'fit' and resource 'garments onto person's photo' are specific, and the tool is distinct from siblings like edit_image or generate_image due to its focused try-on functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly defines usage for virtual try-on but does not explicitly state when to use this tool over alternatives or when not to use it. Sibling tools like edit_image or generate_image are not mentioned, so an agent lacks guidance on trade-offs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileA
Idempotent

Upload a local file to Pruna AI for use in editing/video workflows.

Args: file_path: Local file path to upload (max 20MB)

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a key behavioral constraint (max 20MB) beyond what annotations provide. It aligns with annotations (idempotentHint, non-destructive). However, it could disclose whether duplicate file paths overwrite or create new versions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise: two sentences plus an args section. It is front-loaded with the primary purpose and the args section is clear and efficient. No extraneous text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with an output schema, the description covers the essential behavior. It could be improved by indicating what the tool returns (e.g., file ID) or prerequisites (file existence), but it is largely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds the 20MB size limit and clarifies that file_path is a local path. Since the schema has 0% description coverage, this is a valuable addition. It could specify allowed file types or formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (upload), the target (Pruna AI), and the context (editing/video workflows). It effectively distinguishes from sibling tools that perform other operations like editing or generating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on when to use this tool (for editing/video workflows) but does not explicitly mention when not to use it or suggest alternatives. The sibling tools list implies related use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upscale_imageA
Idempotent

Upscale an image using Pruna AI.

Args: image: Image URL or local file path to upscale target: Target resolution in megapixels (1-128, capped at 128 MP) output_format: Output format (webp, jpg, png) enhance_details: Enhance fine textures enhance_realism: Improve realism (recommended for AI-generated images)

ParametersJSON Schema
NameRequiredDescriptionDefault
imageYes
targetNo
output_formatNojpg
enhance_detailsNo
enhance_realismNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=false (mutation), destructiveHint=false, and idempotentHint=true. The description adds that it upsamples and enhances details/realism but does not reveal additional behavioral traits like file size limits or processing time. It meets the baseline but adds minimal value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise: one-line purpose followed by a clear bullet list of parameters. Every sentence is necessary and front-loaded, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all input parameters but omits details about the output (e.g., returned image URL). Given the tool's moderate complexity and lack of output schema, describing the return value would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description fully explains each parameter: image source, target megapixels with range, output format options, and boolean enhancements. This adds crucial meaning beyond the schema's type/defaults, enabling correct usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Upscale an image using Pruna AI.' This directly conveys the purpose, distinguishing it from sibling tools like edit_image and generate_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. It does not mention any prerequisites or situations where the tool is inappropriate, leaving the agent to infer usage from the parameter descriptions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.2.0
    • First observededit_image
    • First observedgenerate_image
    • First observedgenerate_video
    • First observedlist_models
    • First observedtransform_video
    • First observedtry_on_image
    • First observedupload_file
    • First observedupscale_image

TDQS

A4.1/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: image generation, image editing, video generation, video transformation, try-on, upscaling, model listing, and file upload. No two tools overlap in function.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (e.g., generate_image, edit_image, list_models) using snake_case, making them predictable and easy to understand.

Tool Count5/5

With 8 tools, the server is well-scoped for its purpose—covering core media generation and editing tasks without being overly sparse or cluttered.

Completeness4/5

The set covers major operations: generation, editing, transformation, upscaling, try-on, model listing, and file upload. A minor gap is the lack of a text-based video editing tool analogous to edit_image, but overall it's quite complete.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers