gimp-mcp
gimp-mcp
GIMP 3을 구동하는 MCP 서버로, 스크립트 기반 이미지 편집을 위한 도구입니다: 자르기, 크기 조정, 종횡비 맞추기, 가벼운 색상 보정, 치수 사양 검증, 폴더 전체에 대한 일괄 처리를 제공합니다.
Windows의 GIMP 3.2.4에서 빌드 및 검증되었으며, 기존 2.x Script-Fu 인터페이스 대신
GIMP 3의 GObject Introspection Python API(gi.repository.Gimp)를 사용합니다.
용도
동일한 결정적(deterministic) 처리를 반복적으로 적용해야 하고, 클릭으로 처리하기보다 설명으로 처리하고 싶은 모든 워크플로에 적합합니다:
사진을 대상 종횡비로 자르거나, 가장 큰 중앙 정사각형으로 자르기
이미지 폴더의 가장 긴 변이 최대 2000px가 되도록 크기 조정
게시 전에 이미지가 크기/방향 요구 사항을 충족하는지 확인
한 번의 패스로 전체 촬영본에 동일한 자르기-크기조정 파이프라인 적용
Related MCP server: gimp-mcp
가장 주의해야 할 문제: EXIF 방향
휴대폰과 많은 카메라의 사진은 EXIF 방향 태그와 함께 가로 방향으로 저장되는 경우가 흔하며, 이 태그는 뷰어에게 회전하라고 지시합니다. 모두가 3000x4000 세로 사진으로 보는 사진이 4000x3000으로 저장되어 있을 수 있습니다.
GIMP의 비대화형 로더는 해당 태그를 적용하지 않습니다. "중앙 정사각형 자르기"를 단순하게 적용하면 잘못된 축을 자르게 되어 옆으로 누운 이미지가 생성됩니다. 그런데도 그럴듯한 치수가 보고되므로, 출력물을 열기 전까지는 명백히 잘못된 것이 눈에 띄지 않습니다.
이 프로젝트의 모든 로드는 load_image()를 거치며, 이 함수는 먼저
Gimp.Image.policy_rotate()를 호출하므로 모든 지오메트리 — 그리고 이 서버가 보고하는
모든 치수 — 는 표시 방향(displayed orientation), 즉 뷰어가 실제로 보는 방향을
기준으로 합니다. 이는 테스트로 검증됩니다.
아키텍처
두 가지 실행 백엔드, 하나의 공유 작업 런타임:
┌───────────────────────────────┐
MCP client ──────►│ gimp_mcp/server.py (stdio) │
└───────────┬───────────────────┘
│
┌─────────────────┴──────────────────┐
▼ ▼
HeadlessBackend BridgeBackend
spawns gimp-console-3.exe TCP 127.0.0.1:50472
(no running GIMP needed) (into a running GIMP)
│ │
▼ ▼
bootstrap.py plug-ins/gimp-mcp-bridge/
│ │
└──────────────┬─────────────────────┘
▼
gimp_mcp/gimp_runtime.py
THE single source of truth for every
image operation. Both paths share it,
so batch and live cannot drift apart.install_plugin.py는 gimp_runtime.py를 복사하는 대신 설치된 플러그인 옆에
runtime_path.txt 포인터를 작성하므로, 작업 코드의 단일 사본만 디스크에 존재합니다.
백엔드 선택. headless가 기본값이며 모든 일괄 및 결정적 작업에 사용됩니다 —
열려 있는 GIMP가 필요 없고 신뢰할 수 있는 경로입니다. bridge는 이미 열려 있는
문서에 대해 실시간 작업을 할 때 사용합니다. 두 백엔드 모두 픽셀 단위로 동일한
출력을 생성하는 것으로 검증되었습니다.
TCP를 사용하는 이유, D-Bus가 아닌 이유
기존의 실시간 GIMP 제어 프로젝트는 D-Bus를 사용하지만, D-Bus는 Windows에
존재하지 않습니다. 루프백 TCP 소켓은 동일한 기능을 달성하면서 크로스 플랫폼입니다.
127.0.0.1에만 바인딩되며 네트워크에 절대 노출되지 않습니다.
설치
GIMP 3.x(3.2.4 기준으로 개발됨)와 mcp Python 패키지가 필요합니다.
mcp의존성에 대한 참고. 이 프로젝트는mcp1.x SDK를 대상으로 하며mcp>=1.0,<2로 고정되어 있습니다. 2.0 버전은mcp.server.fastmcp를 제거하고FastMCP를MCPServer로 이름을 바꾸었습니다. 2.0으로의 포팅은 아직 완료되지 않았으며, 고정하지 않고 설치하면 2.x가 설치되어 import 시 실패합니다.
pip install -r requirements.txt
python install_plugin.py # install the bridge plug-in (optional)
python install_plugin.py --list # show detected GIMP config dirs브리지 플러그인은 실시간 제어 도구에만 필요합니다. 일괄 및 단일 이미지 도구는 GIMP에 아무것도 설치하지 않고도 작동합니다.
플러그인 위치
install_plugin.py는 버전을 하드코딩하는 대신 실제로 존재하는 GIMP 3.x 설정
디렉터리를 탐색합니다. Windows에서는 다음과 같습니다:
%APPDATA%\GIMP\3.2\plug-ins\gimp-mcp-bridge\gimp-mcp-bridge.py버전이 포함된 디렉터리(GIMP 3.2의 경우 3.2, 3.0이 아님)이며, GIMP 3에서는
각 플러그인이 .py 파일과 이름이 일치하는 폴더에 있어야 합니다. Linux와 macOS에서
설치 프로그램은 각각 ~/.config/GIMP/3.x/와 ~/Library/Application Support/GIMP/3.x/를
찾습니다.
MCP 서버 등록
패키지를 설치하면 gimp-mcp 콘솔 스크립트가 제공되며, 이는 작업 디렉터리에 의존하지
않으므로 등록하기에 가장 깔끔한 방법입니다:
python -m venv .venv
.venv/Scripts/python -m pip install -e . # .venv/bin/python on Unix{
"mcpServers": {
"gimp": {
"type": "stdio",
"command": "/path/to/gimp-mcp/.venv/Scripts/gimp-mcp.exe",
"args": []
}
}
}Claude Code에서는 다음과 같은 한 줄 명령으로 동일하게 처리할 수 있습니다:
claude mcp add gimp --scope user -- /path/to/gimp-mcp/.venv/Scripts/gimp-mcp.exe해당 인터프리터에서 mcp를 import할 수 있다면 모듈을 직접 실행해도 됩니다:
{
"mcpServers": {
"gimp": {
"command": "python",
"args": ["-m", "gimp_mcp"],
"cwd": "/path/to/gimp-mcp"
}
}
}선택적 환경 변수:
변수 | 용도 |
| 자동 감지되지 않는 경우 |
|
|
| 브리지 포트, 기본값 |
도구
검사
도구 | 용도 |
| GIMP에 연결할 수 있는지 확인; 두 백엔드를 모두 보고합니다. 문제가 있을 때 여기서 시작하세요. |
| 치수, 레이어, 방향. 치수는 표시된 대로입니다. |
| 치수 사양에 대한 검증; 측정된 치수와 평이한 언어의 이유와 함께 통과/실패를 반환합니다. |
단일 이미지
도구 | 용도 |
| 정확한 픽셀 사각형. 범위를 벗어나면 조용히 클램프하지 않고 거부합니다. |
| 가장 큰 정사각형; |
| 대상 비율(1.0 정사각형, 4:3은 1.3333, 16:9는 1.7778), 최대 면적. |
| 너비, 높이 또는 |
| 밝기/대비, -0.5..0.5로 제한됩니다. |
| 한 번에 처리: 자르기로 방향 수정, 최소 크기로 업스케일, 최대 크기로 다운스케일, 선택적 보정. |
| 한 번의 패스로 사용자 정의 작업 파이프라인(JPEG 재인코딩 1회). |
일괄 처리
도구 | 용도 |
| 폴더 전체에 대한 임의의 파이프라인. |
| 전체 폴더를 하나의 치수 사양에 맞춥니다. |
| 읽기 전용 감사; 편집 전 분류에 사용합니다. |
전체 일괄 처리는 하나의 GIMP 호출 안에서 실행됩니다. GIMP 콘솔은 시작하는 데
수 초가 걸리므로 파일별로 프로세스를 생성하면 느릴 것입니다 — 작은 폴더의 경우
파일당 약 ~2.4배 저렴한 것으로 측정되었으며, 폴더가 클수록 절감 효과는 커집니다.
실패한 파일이 실행을 중단하지 않으며, errors에 기록되고 나머지는 계속 진행됩니다.
실시간 제어(브리지 플러그인 필요)
도구 | 용도 |
| 실행 중인 GIMP에서 열려 있는 이미지 목록. |
| 캔버스의 병합된 스냅샷으로, 보고 반복할 수 있습니다. |
| 실시간 컨텍스트에서 임의의 Python 실행; |
| 브리지를 중지하고 GIMP는 열어 둡니다. |
GIMP에서 브리지 시작: 필터 > 개발 > MCP 브리지 시작.
이미지 사양
check_image_spec, fit_to_spec 및 해당 일괄 버전은 하나의 사양 모델을 공유합니다.
모든 제약 조건은 선택 사항입니다 — 0은 제한 없음을 의미하고, 방향 any는 방향
요구 사항이 없음을 의미합니다.
필드 | 값 |
| 픽셀, 최소값 없음은 |
| 픽셀, 최대값 없음은 |
|
|
fit_to_spec는 세 단계의 순서로 사양을 충족합니다: 방향을 수정하기 위한 자르기,
최소 크기에 도달하기 위한 업스케일, 최대 크기를 존중하기 위한 다운스케일. 이미 충족된
제약 조건은 프레이밍을 변경하지 않습니다.
// A square image at least 1000x1000, capped at 2000x2000
{ "orientation": "square", "min_width": 1000, "min_height": 1000,
"max_width": 2000, "max_height": 2000 }색상 조정은 의도적으로 제한적입니다
adjust_image는 밝기/대비를 -0.5..0.5로 제한하며, 범위를 벗어나는 값은 클램프하지
않고 거부합니다. 약 ±0.15를 넘는 값은 사진의 특성을 눈에 띄게 바꾸며, 이는
이미지가 실제 대상을 충실히 표현해야 할 때 중요합니다. 채도 부스트나 "자동 향상"은
의도적으로 없습니다.
검증
테스트 스위트 실행:
python -m pytest tests/ -v실제 이미지가 필요한 테스트는 이미지를 지정하지 않으면 건너뜁니다:
export GIMP_MCP_TEST_IMAGE=/path/to/photo.jpg # ideally EXIF-rotated
export GIMP_MCP_TEST_REFERENCE=/path/to/photo-square.jpgGIMP_MCP_TEST_REFERENCE는 GIMP_MCP_TEST_IMAGE의 독립적으로 생성된 중앙 정사각형
자르기여야 합니다 — 예를 들어 GIMP에서 수동으로 자른 것입니다. 핵심 테스트는
crop_square가 오류 없이 실행되는 것에 그치지 않고 해당 참조를 재현한다는 것을
검증합니다.
개발 중 사용된 참조 사진(EXIF 방향 6의 4000x3000 JPEG, 3000x4000으로 표시됨)에서:
crop_square vs hand-made reference : mean abs diff 0.236, max 18, outliers 0.0014%
same crop via the bridge backend : mean abs diff 0.236, max 18, outliers 0.0014%그 잔차는 JPEG 재인코딩 노이즈입니다 — 재인코딩만으로도 평균 ~0.5가 발생하며 — 지오메트리 차이가 아니며, 두 백엔드는 정확히 일치합니다.
테스트 스위트는 또한 표시 방향 보고, 방향 및 최소 크기 사양, 범위를 벗어난 자르기 거부, 범위를 벗어난 조정 거부, 밝기가 픽셀을 올바른 방향으로 이동하는지, 체인 파이프라인, 종횡비 자르기, 폴더 전체 일괄 처리, 읽기 전용 감사, 누락된 파일에 대한 명확한 오류, 실제 MCP stdio 프로토콜에 대한 전체 통과를 다룹니다.
문제 해결
gimp-console not found — GIMP_CONSOLE에 gimp-console-3.exe의 전체 경로를
설정하세요.
브리지 도구가 "Could not reach the GIMP bridge"로 실패 — GIMP가 열려 있지 않거나
브리지가 시작되지 않았습니다. 필터 > 개발 > MCP 브리지 시작을 실행하세요.
gimp_status는 두 백엔드를 동시에 표시합니다.
설치 후 메뉴 항목이 보이지 않음 — GIMP를 다시 시작하세요. 플러그인은 시작 시에만
스캔합니다. 레이아웃이 plug-ins/gimp-mcp-bridge/gimp-mcp-bridge.py인지 확인하세요
(폴더 이름은 파일 이름과 일치해야 합니다).
플러그인 진단 — GIMP 플러그인은 별도 프로세스이며, Windows에서 GIMP가 GUI 앱으로
실행될 때 stderr가 보이지 않습니다. 브리지는 설치된 플러그인 옆의 bridge.log에
기록합니다.
색상 프로필 대화상자가 GIMP 시작을 차단 — GUI 모드에서 임베디드 프로필이 있는 이미지를 열 때 발생합니다. 헤드리스 모드에서는 나타나지 않으며, 이것이 일괄 작업이 헤드리스 백엔드를 사용하는 또 다른 이유입니다.
일괄 처리 시간 초과 — 기본값은 전체 실행에 600초입니다. 매우 큰 폴더는 더 많은 시간이 필요할 수 있습니다.
알려진 제한 사항
라이브 제어는 가볍게만 검증되었다. 작동이 확인되었지만(이미지 열기, 목록, 스크린샷, 라이브 편집, 브리지를 통한 크롭이 헤드리스와 동일한 출력으로 수행됨), 헤드리스 경로보다 사용량이 훨씬 적다. 헤드리스를 신뢰할 수 있는 경로로 취급하라.
브리지는 설계상 임의의 Python을 실행한다. 루프백 전용이며 자동이 아닌 수동으로 시작되지만, 머신의 localhost에 도달할 수 있는 모든 것은 GIMP가 실행되는 동안 GIMP를 구동할 수 있다. 사용하지 않을 때는 중지하라.
브리지 시작은 자체 플러그인 프로세스를 차단한다 — 이것이 브리지를 살아 있게 유지하는 방식이다. GIMP의 UI를 멈추지는 않지만, GIMP는 플러그인이 실행 중인 것으로 표시한다.
GUI 메뉴 항목 자체는 자동화 테스트로 검증되지 않았다. 해당 항목이 호출하는 절차는 검증되었지만, 클릭 경로는 검증되지 않았다.
Windows만 검증되었다. 코드 경로는 크로스 플랫폼이며 설치 프로그램이 Linux/macOS 구성 디렉터리를 처리하지만, 둘 다 테스트되지 않았다.
mcp2.x SDK는 아직 지원되지 않는다 — Install 아래의 참고 사항을 보라.AI 배경 제거나 스타일 전환은 없다. 일부 유사 프로젝트는 작동하는 구현 없이 이를 광고하지만, 여기서는 의도적으로 주장하지 않는다.
선행 기술에 대한 참고 사항
브리지를 노출하는 GIMP 측 플러그인과 클라이언트로 연결되는 독립형 MCP 서버 프로세스 간의 분리는 이 문제에 자연스러운 형태이며 다른 GIMP MCP 프로젝트에서도 사용된다. 배치 처리와 프리셋 스타일 파이프라인은 여러 프로젝트에 공통적이다. 라이브 캔버스 제어는 다른 곳에서 D-Bus를 통해 존재하지만, 여기서는 Windows 지원을 위해 루프백 TCP로 대체되었다. 이들 중 어느 것에서도 코드를 복사하지 않았다. Windows 특정 사항 — 실제 플러그인 경로, 플러그인 프로세스 수명, 실행 콜백 시그니처, EXIF 동작 — 은 GIMP 3.2.4에 대해 직접 확인되었다.
라이선스
MIT — LICENSE 참조.
Available Tools
18 toolsadjust_imageA
Brightness/contrast touch-up.
Both values run -1.0..1.0 -- GIMP 3's real range for this operation -- and
are rejected outside it rather than clamped. enhance_image uses the same
range for contrast; they wrap the same GIMP call.
Useful values are far smaller than the limits. GIMP 3 applies this in linear light, so it bites harder than the same number did in GIMP 2.x, and anything much past +/-0.1 visibly changes the character of a photo. Keep adjustments small when the image needs to represent a real subject faithfully.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| quality | No | ||
| contrast | No | ||
| brightness | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well. It discloses that out-of-range values are rejected rather than clamped, that the operation runs in linear light and therefore has a stronger effect than in GIMP 2.x, and that values past +/-0.1 visibly change the image. This goes well beyond a generic 'adjusts brightness and contrast.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then adds only high-value behavioral and practical details. Every sentence earns its place: the range/rejection rule, the relationship to enhance_image, the linear-light warning, and the guidance to keep adjustments small.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no annotations and 0% schema coverage, the description is nearly complete: it explains the main operation, the key parameters, edge-case behavior, and how values behave differently in GIMP 3. It misses only the optional backend and quality semantics, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides rich semantics for brightness and contrast (range, rejection behavior, practical limits), but says nothing about backend or quality, leaving the quality default of 0.92 unexplained. The required input_path and output_path are self-evident from their names, so the gap is mainly in the optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Brightness/contrast touch-up,' a specific verb and resource that immediately identifies the tool's function. It also references enhance_image, providing a sibling distinction even though the exact selection criteria are not spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful operational guidance about value ranges and suggests keeping adjustments small, but it does not explicitly state when to choose adjust_image over enhance_image or other siblings. The mention that both wrap the same GIMP call is informative but stops short of routeing the agent to the right tool in a specific scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_check_image_specA
Audit a folder: which images already satisfy a specification.
Read-only; writes nothing. Use it to triage a folder before editing.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| pattern | No | ||
| input_dir | Yes | ||
| max_width | No | ||
| min_width | No | ||
| recursive | No | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and explicitly states 'Read-only; writes nothing.' This is a valuable, non-obvious disclosure about side effects. It does not go further into permissions or edge cases, but it covers the primary behavioral concern.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and followed by a concise safety/usage note. Every word earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, safety, and a clear use case, and an output schema exists so return values need not be described. However, with 9 parameters and zero schema descriptions, the description leaves parameter semantics and default behavior underspecified. It is adequate for a first correct invocation but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the tool has 9 parameters. The description only mentions 'specification' in general terms and does not explain how parameters like pattern, recursive, orientation, or min/max dimensions map to that specification. Parameter names are suggestive, but the description itself provides almost no semantic value beyond the schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Audit') and resource ('a folder') plus the precise condition ('which images already satisfy a specification'). The batch/folder framing clearly distinguishes it from singular tools like check_image_spec.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit use context: 'Use it to triage a folder before editing.' This tells an agent when it is appropriate, though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_fit_to_specC
Make every image in a folder satisfy one dimension specification.
The common bulk case: point it at a folder of photos and get conforming copies, with each file's final dimensions and pass/fail reported.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| suffix | No | _out | |
| backend | No | ||
| pattern | No | ||
| quality | No | ||
| input_dir | Yes | ||
| max_width | No | ||
| min_width | No | ||
| recursive | No | ||
| max_height | No | ||
| min_height | No | ||
| output_dir | Yes | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose key behavior: it produces conforming copies, operates on every image in a folder, and reports final dimensions and pass/fail. However, it does not clarify how the dimension constraints interact, whether originals are left untouched, naming/overwrite behavior, or recursive handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core function and the second gives the primary use case and expected reporting. It is not padded, though it sacrifices useful detail to achieve that brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter batch tool with 0% schema coverage and no annotations, the description is too thin. It gives a scenario and outcome but leaves the agent without enough information about parameter semantics, alternatives, and behavior constraints; the presence of an output schema only partially offsets the missing return-value detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 13 parameters, and the description adds no parameter-level meaning. 'One dimension specification' hints at constraints but does not explain input_dir, output_dir, min/max width/height, orientation, pattern, quality, backend, suffix, recursive, or anchor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete operation: make every image in a folder satisfy a dimension specification, and adds bulk/copy/reporting context. It does not explicitly name a sibling such as fit_to_spec or batch_check_image_spec, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'common bulk case' implies this is for folder-wide operations and contrasts with single-image tools, but no explicit when-to-use/when-not-to-use guidance or alternative tool names are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_processA
Apply the same operations to every image in a folder.
All files are handled inside a single GIMP session, so a large folder
costs one GIMP startup rather than one per file. A file that fails does
not abort the run: it is reported in errors and the rest continue.
operations is a JSON list, same format as process_image. Set suffix to
e.g. "_out" to avoid overwriting, and output_format to "jpg"/"png" to
convert.
| Name | Required | Description | Default |
|---|---|---|---|
| suffix | No | ||
| backend | No | ||
| pattern | No | ||
| quality | No | ||
| input_dir | Yes | ||
| recursive | No | ||
| operations | Yes | ||
| output_dir | Yes | ||
| output_format | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses single-session processing, failure isolation with reported errors, and overwrite avoidance through suffix, all of which are valuable beyond the schema. It does not cover every behavioral nuance, but the core runtime traits are well described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: one sentence for purpose, one short paragraph for runtime behavior, and one for parameter guidance. Every sentence contributes useful information without repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Core behavior, error reporting, and key parameter usage are covered, and an output schema is indicated to exist. However, with 9 parameters and zero schema descriptions, several parameters such as backend, pattern, recursive, and quality remain under-specified for an agent to invoke the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain operations as a JSON list in process_image format, suffix for avoiding overwrites, and output_format for conversion, but it leaves backend, pattern, recursive, quality, input_dir, and output_dir semantically unexplained despite having 9 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: apply the same operations to every image in a folder. This clearly distinguishes the tool from single-image tools like process_image and from specialized batch tools like batch_fit_to_spec or batch_check_image_spec.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the batch usage context clear by explaining that a large folder costs one GIMP startup rather than one per file, which strongly implies the batch counterpart to process_image. It does not explicitly name alternatives or say when not to use this tool, but the context is not misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_image_specA
Check an image against a dimension specification.
Every constraint is optional: 0 means "no limit", and orientation "any" means no orientation requirement. Valid orientation values are: any, square, landscape, portrait, square_or_landscape, square_or_portrait.
Returns pass/fail with the actual measured dimensions and a plain-language reason for each failure. Useful for validating images against a publishing platform's requirements, a print size, or an asset pipeline's conventions before spending time editing.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| backend | No | ||
| max_width | No | ||
| min_width | No | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that every constraint is optional, that 0 means no limit, that 'any' means no orientation requirement, and that the tool returns pass/fail with measured dimensions and failure reasons. This is useful behavioral detail, though it does not cover error cases or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a clear one-line purpose, then a compact explanation of constraint semantics and orientation values, then return behavior and use cases. Every sentence earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 7 parameters, no annotations, and 0% schema coverage, the description covers the key semantics well: optional constraints, 0 meaning, orientation values, and return payload. It does not explain the backend parameter or explicitly compare itself with sibling validation tools, but an output schema exists to cover return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add real meaning by explaining the 0-as-no-limit convention for numeric constraints and enumerating valid orientation values. However, the optional 'backend' parameter is never explained, which is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check an image against a dimension specification.' It also states the return contract (pass/fail with dimensions and reasons), which clearly distinguishes it from siblings like crop_image, resize_image, or inspect_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: validating images against publishing requirements, print sizes, or asset pipeline conventions before editing. It does not explicitly exclude alternatives or name sibling tools such as fit_to_spec or inspect_image, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_imageB
Crop to an exact pixel rectangle.
x/y are the top-left offset in DISPLAYED orientation. Fails clearly if the rectangle falls outside the image rather than silently clamping.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| width | Yes | ||
| height | Yes | ||
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers important details: x/y are interpreted in DISPLAYED orientation and out-of-bounds rectangles fail clearly instead of silently clamping. It could disclose overwrite behavior or backend semantics, but the core failure and orientation behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the most decision-relevant information is front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and 0% schema coverage, the description is too thin. It omits backend and quality meaning and does not state whether output_path overwrites existing files, leaving an agent to guess on non-obvious options even though required paths are given in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% parameter descriptions, so the description must compensate. It adds meaning for x/y as top-left offsets in displayed orientation and implies width/height are pixel-based, but backend and quality are left unexplained, and input_path/output_path semantics are still only inferable from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Crop to an exact pixel rectangle', which clearly identifies the operation and resource. It also distinguishes this tool from siblings like crop_to_aspect and crop_square by emphasizing exact pixel dimensions, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus crop_to_aspect, crop_square, or other siblings. The phrase 'exact pixel rectangle' implies a use case, but no when-to-use or when-not-to-use context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_squareB
Crop to the largest possible square.
anchor picks which part of the frame to keep: center (default), top, bottom, left, right, or a corner such as topleft.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the behavioral burden. It does explain that the crop keeps the largest possible square and that the anchor determines which part of the frame survives, including the center default. However, it does not disclose what backend or quality do, whether files are overwritten, or what side effects occur beyond writing output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. The main action appears in the first sentence, and the second sentence earns its place by clarifying the anchor parameter. There is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimally adequate for a simple crop operation: it states the core behavior and the key anchor option. But with five parameters and no annotations, important details like backend and quality semantics are missing, and no usage context versus sibling tools is provided. The presence of an output schema lowers the burden for return-value documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's silence. It adds helpful semantics for anchor by enumerating valid values and the default. It leaves backend and quality completely unexplained, and input_path/output_path relationships are only implicit from the tool's name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action and resource: 'Crop to the largest possible square.' It also adds meaningful detail about the anchor parameter. It does not explicitly name or differentiate from sibling tools like crop_to_aspect or crop_image, but the square-only behavior is reasonably distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use crop_square versus the many sibling cropping/resizing tools. The description does not state prerequisites, exclusions, or conditions that would help an agent select this tool over crop_to_aspect, crop_image, or adjust_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_to_aspectB
Crop to a target aspect ratio (width/height), keeping maximum area.
Use 1.0 for square, 1.3333 for 4:3, 1.5 for 3:2, 1.7778 for 16:9.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | Yes | ||
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It discloses the 'keeping maximum area' behavior and explains the ratio meaning, which is useful. However, it does not mention important behaviors like default anchoring, quality handling, or whether the operation modifies the input image.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core behavior, and every sentence adds value. The ratio examples are practical and directly aid correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, 6 parameters, and zero schema description coverage, the description is too sparse. It fails to clarify optional parameters or provide enough context to confidently invoke the tool beyond the basic ratio and paths.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It meaningfully explains the ratio parameter with examples, but provides no additional semantics for anchor, backend, quality, input_path, or output_path. This leaves several parameters underdocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool crops to a target aspect ratio while keeping maximum area, with concrete ratio examples. It is specific enough to be understood, though it does not explicitly differentiate itself from sibling tools like crop_image or crop_square.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides common ratio values but offers no guidance on when to choose this tool over alternatives such as crop_image, crop_square, or fit_to_spec. There are no exclusions or explicit usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enhance_imageA
Tone and detail enhancement in one pass.
gamma lifts midtones and shadows via levels, leaving the black and white points alone so nothing clips. 1.0 = off. contrast GIMP 3 native -1..1. 0 = off. saturation -100..100. 0 = off. sharpen high-pass sharpen blended back at this percent opacity, 0 = off. Preferred over unsharp mask, which haloes. sharpen_radius blur radius in pixels for the high pass (default 8).
Contrast bites harder than the same nominal value did in GIMP 2.x, because GIMP 3 runs the operation in linear light: 2.x's "+12" is roughly 0.020 here, not 0.094. Calibrate against output rather than remapping an old number.
| Name | Required | Description | Default |
|---|---|---|---|
| gamma | No | ||
| backend | No | ||
| quality | No | ||
| sharpen | No | ||
| contrast | No | ||
| input_path | Yes | ||
| saturation | No | ||
| output_path | Yes | ||
| sharpen_radius | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it explains that gamma avoids clipping by preserving black/white points, sharpen is high-pass blended back at a given opacity, and contrast runs in GIMP 3 linear light, with explicit calibration caveats. This gives agents a strong model of the tool's actual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded, each parameter gets a compact line, and the GIMP 3 calibration note earns its place. There is no filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations, the description is thorough on the core enhancement behavior and parameter semantics. It is not fully complete because backend and quality are unexplained, and no sibling-tool comparison is given, leaving some ambiguity in tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates well by explaining gamma, contrast, saturation, sharpen, and sharpen_radius with scales and defaults. However, backend and quality are left undocumented in the description, so their semantics remain ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation ('Tone and detail enhancement in one pass') and elaborates on what each control does, so an agent understands the tool's role. It does not explicitly differentiate it from sibling tools like adjust_image or process_image, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to choose enhance_image over the many sibling image tools, nor exclusions for when not to use it. The parameter-level note about preferring high-pass sharpen over unsharp mask is useful but not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fit_to_specA
Transform an image until it satisfies a dimension specification.
Crops to fix the orientation if required, upscales to reach a minimum size, downscales to respect a maximum, and optionally applies a light touch-up -- all in one pass, so the JPEG is re-encoded only once. Images already satisfying a constraint keep their framing.
Example: to produce a square image at least 1000x1000, pass orientation="square" with min_width=1000 and min_height=1000.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| upscale | No | ||
| contrast | No | ||
| max_width | No | ||
| min_width | No | ||
| brightness | No | ||
| input_path | Yes | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any | |
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does well: it discloses cropping, upscaling, downscaling, optional touch-up, single re-encode of JPEG, and preservation of already-compliant framing. It leaves some specifics unstated (e.g., overwrite behavior, failure conditions, handling of non-JPEG inputs), which prevents a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, all informative: the first sentence states the purpose, the second explains the mechanism, and the example grounds the parameters. No filler or repetition of schema field names, and the key operations are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no annotations and no schema descriptions, the description covers the core workflow adequately but omits enough optional-parameter semantics to be fully self-contained. The presence of an output schema means return values need not be described, but the parameter gaps and lack of sibling differentiation leave clear holes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters; it gives meaning to orientation, min_width, min_height, and the max constraints via the example and operation summary. However, many of the 13 parameters (anchor, backend, quality, upscale, contrast, brightness, max_width/max_height defaults) are not explained in the description or schema, leaving significant inference required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Transform an image until it satisfies a dimension specification,' then enumerates the exact operations (crop for orientation, upscale, downscale, optional touch-up) and gives a concrete square-image example. This clearly separates the tool from generic resize/crop siblings by emphasizing the all-in-one constraint-satisfaction behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied clearly: call this tool when an image must meet dimension constraints (min/max width/height, orientation) in one pass. However, it never explicitly tells an agent when to prefer fit_to_spec over sibling tools like crop_to_aspect, resize_image, or process_image, nor states when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gimp_statusA
Check that GIMP is reachable and report its version.
Use this first if anything seems wrong. Reports both the headless backend and whether the live bridge plug-in is running inside an open GIMP.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what the tool reports (headless backend status and live bridge plug-in state) beyond a simple reachability check, which adds useful behavioral context. It doesn't explicitly state non-destructive behavior, but 'check' and 'report' strongly imply a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the core purpose front-loaded and a clear usage directive. There is no filler or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status tool with one optional parameter, an output schema, and no required inputs, the description covers purpose, usage, and report contents. The only notable gap is the undocumented 'backend' parameter, but the tool can be correctly invoked with no arguments, so the description is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions the single optional 'backend' parameter or its meaning. An agent cannot determine what value to pass or why the parameter exists. The description mentions 'headless backend' as reported output, but that doesn't clarify the parameter's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Check that GIMP is reachable and report its version.' It also distinguishes itself from the many sibling processing tools by being a diagnostic/status tool, not an image manipulation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Use this first if anything seems wrong.' This tells the agent when to invoke it, though it doesn't name specific alternatives or exclusions. The first-step diagnostic role is clear enough for practical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_imageB
Report an image's dimensions, layers, and orientation.
Dimensions are reported as DISPLAYED (EXIF orientation applied), which is what a viewer sees -- not necessarily how the pixels are stored.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| backend | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that dimensions are reported in displayed form after EXIF orientation is applied, not as raw stored pixels. This is a meaningful behavioral nuance, though it could also mention that this is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief, front-loads the main purpose, and includes only the essential EXIF detail. Every sentence contributes value and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple inspection tool with an output schema, the description covers the core result and a key display nuance. However, it lacks an explanation of the backend parameter and gives no guidance on when this tool should be preferred over related tools, leaving some gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it adds no information about the path or backend parameters. The backend parameter is completely unexplained, and the description only implies the image is referenced by path without discussing either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Report') and a specific resource ('an image's dimensions, layers, and orientation'). This distinguishes it from sibling tools that crop, resize, or process images rather than inspect metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided, and it does not mention alternatives such as check_image_spec or other inspection-like siblings. The intended context must be inferred from the tool name and description, so there is no explicit usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_list_imagesA
List the images currently open in the running GIMP.
Requires the bridge plug-in (Filters > Development > Start MCP Bridge).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It usefully reveals that the tool depends on the bridge being started and reads current live GIMP state. 'List' implies a non-mutating operation, and the output schema covers return structure, so this is transparent enough for a simple read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the core purpose front-loaded and the required setup immediately after. Every sentence earns its place and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter live listing tool, the description covers what it does, the environment prerequisite, and leaves return details to the output schema. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing for the description to clarify about arguments. Per the baseline for no-parameter tools, this is appropriately handled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a specific resource ('images currently open in the running GIMP'), making the tool's function immediately clear. It is naturally distinguishable from sibling tools like crop_image, gimp_status, and inspect_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys that this tool is for querying the current live set of open images in GIMP, and it flags the prerequisite of the bridge plug-in. It does not explicitly compare to alternatives, but the zero-parameter live-listing purpose is self-evident enough for an agent to know when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_run_pythonA
Execute Python inside the running GIMP and return result.
The Gimp module and every operation helper are already in scope. Assign to
a variable named result to return a value. Escape hatch for anything the
typed tools above do not cover; requires the bridge plug-in.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It usefully explains that GIMP modules and helpers are in scope and that a `result` variable must be assigned to return a value. However, it does not disclose that arbitrary Python execution can mutate or destroy GIMP state, crash the session, or have irreversible side effects, which is a significant transparency gap for an unbounded execution tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the main purpose, and every sentence earns its place: execution semantics, scope context, return-value convention, use-case, and prerequisite. No filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema, the description does not need to explain return structure. It covers the parameter, the execution environment, the return mechanism, and the prerequisite. The main missing piece is a warning about the destructive or uncontrolled nature of raw Python execution, which matters for a tool of this complexity and power.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does meaningfully: it tells the agent that `code` is Python to execute inside GIMP, that modules are already in scope, and that assigning to `result` controls the return value. This gives the single parameter real semantic grounding beyond the bare name 'code'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb+resource: 'Execute Python inside the running GIMP and return `result`'. It also differentiates itself from the typed sibling tools by describing itself as an 'escape hatch for anything the typed tools above do not cover', so an agent can distinguish it at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool: when the typed tools do not cover the needed operation. It also establishes the prerequisite that the bridge plug-in must be present. This provides clear when-to-use context and points to the sibling tools as the preferred alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_screenshotA
Save a flattened snapshot of an image open in the running GIMP.
Lets you see the current state of a document you are editing live, then iterate on it. Requires the bridge plug-in.
| Name | Required | Description | Default |
|---|---|---|---|
| max_edge | No | ||
| image_index | No | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does disclose key behavioral traits: it saves a flattened/live snapshot and depends on the bridge plug-in. It does not discuss whether the original image is modified or whether an existing output file is overwritten, but 'snapshot' implies a non-destructive capture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the action, and the bridge plug-in requirement is worth stating. The sentence about seeing current state and iterating is somewhat redundant with 'snapshot' and 'live,' but it adds useful intent context without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters, no annotations, and 0% schema description coverage, so the description must compensate. It gives a clear purpose and prerequisite but leaves max_edge and image_index undefined, and does not explain output_path semantics or file format. An agent could invoke it with defaults, but not understand the full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain any of the three parameters. Output_path is only implied by 'Save', while max_edge and image_index—especially which open image is captured—are not described at all, so the description adds no parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Save a flattened snapshot of an image open in the running GIMP' names a specific verb, a concrete resource, and a key behavior (flattening). This clearly separates it from siblings like inspect_image and live_list_images, so an agent can distinguish the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it when you need to see the current state of a live document and iterate on it. It also states a prerequisite ('Requires the bridge plug-in'), but it does not explicitly mention alternative tools or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_stop_bridgeA
Stop the bridge inside the running GIMP, leaving GIMP itself open.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that the bridge is stopped and that GIMP remains open, which are the two key behavioral consequences. It does not mention edge cases like stopping an already-stopped bridge, but for a zero-parameter control action this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the core action and adds the important clarifying detail about GIMP remaining open. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with a simple, well-defined action, the description is complete. It tells the agent exactly what happens and what does not happen. The output schema is present, so return-value details are not required in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description correctly includes no parameter-specific details because none are needed. There is no schema information to supplement or contradict.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Stop the bridge inside the running GIMP'. It also explicitly clarifies that GIMP itself remains open, which disambiguates this from closing GIMP. This is a clear, distinct purpose among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the context: the bridge is running inside GIMP, and the tool stops only the bridge. It does not explicitly name alternatives or when-not-to-use, but no sibling tool appears to perform a similar stop action, so the context is sufficient for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_imageA
Apply a custom sequence of operations in one pass.
operations is a JSON list, e.g.
[{"op":"crop_square","anchor":"center"},
{"op":"resize","max_edge":2000},
{"op":"adjust","brightness":0.05}]
Available ops: crop, crop_square, crop_aspect, resize, adjust, autocrop, flatten, fit_spec, enhance. Running them as one pipeline re-encodes the JPEG only once, which avoids stacking compression artefacts.
The optional spec arguments are checked against the FINAL result and
reported under spec, so a pipeline that both reshapes and edits an image
can be validated without a second pass over it.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| quality | No | ||
| max_width | No | ||
| min_width | No | ||
| input_path | Yes | ||
| max_height | No | ||
| min_height | No | ||
| operations | Yes | ||
| orientation | No | any | |
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It adds meaningful details: single-pass processing, one-time JPEG re-encoding, and that optional spec arguments are validated against the final result and reported under `spec`. It does not cover failure modes or input format restrictions, but the core execution behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: a one-sentence purpose, a concrete operations example, a concise list of available ops, and two short paragraphs explaining the pipeline benefit and spec-checking behavior. Every sentence earns its place, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no annotations and zero schema descriptions, the description is not fully complete. It covers the central `operations` parameter and the pipeline concept well, but it leaves the optional spec arguments and other tuning parameters under-specified, which an agent would need to invoke the tool correctly for advanced use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does a good job on `operations`, showing a full JSON example and listing valid op names, and it alludes to 'spec arguments'. However, it does not explain other parameters such as `quality`, `backend`, `max_width`, `min_width`, `max_height`, `min_height`, or `orientation`, leaving significant semantic gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Apply a custom sequence of operations in one pass.' It explicitly lists the available operations, which map directly to the sibling individual-operation tools, so an agent can tell that this is the composite/pipeline counterpart to crop_image, resize_image, adjust_image, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete reason to choose this tool over chaining siblings: running operations as one pipeline re-encodes the JPEG only once and avoids stacking compression artifacts. It does not explicitly state when to prefer a single-operation sibling, but the 'custom sequence' framing and op list make the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resize_imageA
Resize an image.
Give width, height, both, or max_edge (longest side, aspect preserved). With preserve_aspect and both dimensions, the image is fitted inside the box rather than distorted.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | ||
| height | No | ||
| backend | No | ||
| quality | No | ||
| max_edge | No | ||
| input_path | Yes | ||
| output_path | Yes | ||
| preserve_aspect | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It usefully explains aspect-ratio preservation and that both dimensions with preserve_aspect fits the image inside the box rather than distorting it. However, it says nothing about defaults, behavior when no sizing parameter is provided, backend handling, quality interpretation, or whether output files are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loaded, with 'Resize an image' first and only the most decision-relevant parameter guidance following. Every sentence earns its place and there is minimal fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and zero schema descriptions, the description covers the core sizing behavior well but leaves gaps around backend, quality, validation, and edge cases. The presence of an output schema reduces the need to explain return values, but the description is still only moderately complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It adds meaning for width, height, max_edge, and preserve_aspect, and the required input/output paths are reasonably self-explanatory from their names. But backend and quality receive no explanation, leaving two parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line 'Resize an image' states a specific verb and resource, and the parameter combinations (width, height, max_edge, preserve_aspect) make the tool's purpose clear. It does not explicitly distinguish itself from sibling tools like crop_image or adjust_image, but 'resize' is distinct enough on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to use the tool's sizing modes ('Give width, height, both, or max_edge') and the behavior of preserve_aspect, which is operational guidance. However, it does not say when to prefer this tool over alternatives such as crop_to_aspect, fit_to_spec, or process_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.1.0- First observed
adjust_image - First observed
batch_check_image_spec - First observed
batch_fit_to_spec - First observed
batch_process - First observed
check_image_spec - First observed
crop_image - First observed
crop_square - First observed
crop_to_aspect - First observed
enhance_image - First observed
fit_to_spec - First observed
gimp_status - First observed
inspect_image - First observed
live_list_images - First observed
live_run_python - First observed
live_screenshot - First observed
live_stop_bridge - First observed
process_image - First observed
resize_image
TDQS
Scored across 18 tools
Several tools overlap in purpose: adjust_image and enhance_image share the same contrast call, and process_image/batch_process can reproduce the effects of most individual editing tools. The descriptions generally clarify scope, but an agent could easily hesitate between a dedicated single-op tool and its pipeline equivalent.
Most tools follow a clear verb_object snake_case pattern such as crop_image, resize_image, inspect_image, and batch_check_image_spec. The batch_ and live_ prefixes are applied consistently, with only gimp_status and fit_to_spec deviating slightly from the otherwise predictable pattern.
At 18 tools the server is slightly above the ideal 3-15 range, but the count is justified by the distinct clusters: single-image operations, spec checking/fitting, batch variants, and live GIMP bridge tools. The single/batch pairs add surface area but each serves a real workload.
The toolset covers the core image pipeline well: inspect, validate, crop, resize, adjust, enhance, process in one pass, and batch over folders. Obvious gaps like rotation or flipping are absent, but live_run_python and process_image provide workarounds for most missing operations.
Maintenance
Related MCP Connectors
Image toolkit: resize, compress, crop, watermark, convert, rotate, EXIF read/strip.
- MochifyOAuthapp.mochify
Image and PDF toolkit: convert to AVIF/WebP/JXL, resize, crop, remove backgrounds, optimize PDFs.
Image toolkit: resize, convert, compress, crop, metadata, hashes, favicons, OCR, QR codes, barcodes.
181LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform GIMP-style image operations such as open, resize, crop, flip, rotate, blur, desaturate, text overlay, export, and batch processing via MCP tools, supporting both mock (Pillow) and live GIMP backends.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to control GIMP for image editing tasks such as opening, resizing, filtering, exporting, and batch processing images through Python-Fu scripting.36 npmMIT
- AlicenseBqualityBmaintenanceEnables controlling GIMP 3 locally through natural language, providing tools for image editing, layer management, selections, and PDB procedure invocation. Keeps all images and files on the user's machine with a local-first, secure design.401GPL 3.0
- AlicenseBqualityAmaintenanceEnables AI agents to operate GIMP 3 end-to-end: open and inspect images, call every PDB procedure, apply GEGL filters destructively or as layer effects, measure pixels, render before/after/diff comparisons, cut out subjects with AI segmentation, and run multi-step recipes across folders.3265 PyPI4Apache 2.0