gimp-mcp
gimp-mcp
Ein MCP-Server, der GIMP 3 für skriptgesteuerte Bildbearbeitung steuert: Zuschneiden, Größenänderung, Seitenverhältnis-Anpassung, leichte Farbkorrekturen, Validierung von Dimensionsspezifikationen und Stapelverarbeitung über einen Ordner.
Erstellt und verifiziert auf Windows mit GIMP 3.2.4, unter Verwendung von GIMP 3s
GObject-Introspection-Python-API (gi.repository.Gimp) anstelle der alten
2.x-Script-Fu-Schnittstelle.
Wofür es gedacht ist
Jeder Workflow, bei dem Bilder wiederholt dieselbe deterministische Behandlung benötigen und man sie lieber beschreiben möchte, als durchzuklicken:
Ein Foto auf ein Ziel-Seitenverhältnis zuschneiden oder auf das größte zentrierte Quadrat
Einen Ordner mit Bildern so verkleinern, dass die längste Kante höchstens 2000px beträgt
Prüfen, ob Bilder vor der Veröffentlichung eine Größen-/Orientierungsanforderung erfüllen
Eine Crop-and-Resize-Pipeline in einem Durchgang auf einen ganzen Shoot anwenden
Related MCP server: gimp-mcp
Die eine Sache, die dich beißen wird: EXIF-Orientierung
Fotos von Handys und vielen Kameras werden häufig im Querformat mit einem EXIF-Orientierungs-Tag gespeichert, das dem Betrachter sagt, sie zu drehen. Ein Foto, das jeder als 3000x4000 Hochformat sieht, kann als 4000x3000 gespeichert sein.
GIMPs nicht-interaktiver Lader wendet dieses Tag nicht an. Ein naives "Zuschneiden auf Quadrat, zentriert" schneidet daher die falsche Achse und erzeugt ein seitenverkehrtes Bild – während es dennoch plausibel aussehende Abmessungen meldet, sodass nichts offensichtlich kaputt aussieht, bis man die Ausgabe öffnet.
Jeder Ladevorgang in diesem Projekt läuft über load_image(), das zuerst
Gimp.Image.policy_rotate() aufruft, sodass alle Geometrie – und jede Dimension, die dieser
Server meldet – in angezeigter Orientierung ist, d.h. was ein Betrachter tatsächlich sieht. Dies ist durch einen Test abgedeckt.
Architektur
Zwei Ausführungs-Backends, eine gemeinsame Operations-Laufzeit:
┌───────────────────────────────┐
MCP client ──────►│ gimp_mcp/server.py (stdio) │
└───────────┬───────────────────┘
│
┌─────────────────┴──────────────────┐
▼ ▼
HeadlessBackend BridgeBackend
spawns gimp-console-3.exe TCP 127.0.0.1:50472
(no running GIMP needed) (into a running GIMP)
│ │
▼ ▼
bootstrap.py plug-ins/gimp-mcp-bridge/
│ │
└──────────────┬─────────────────────┘
▼
gimp_mcp/gimp_runtime.py
THE single source of truth for every
image operation. Both paths share it,
so batch and live cannot drift apart.install_plugin.py schreibt einen runtime_path.txt-Zeiger neben das installierte
Plug-in, anstatt gimp_runtime.py zu kopieren, sodass genau eine Kopie des
Operationscodes auf der Festplatte existiert.
Backend-Wahl. headless ist die Standardeinstellung und wird für alle Stapel- und
deterministischen Arbeiten verwendet – es benötigt kein geöffnetes GIMP und ist der zuverlässige Pfad.
bridge ist für Live-Arbeit an einem Dokument, das bereits geöffnet ist. Beide sind
verifiziert, pixelidentische Ausgaben zu erzeugen.
Warum TCP und nicht D-Bus
Bestehende Live-GIMP-Steuerungsprojekte verwenden D-Bus, das es auf
Windows nicht gibt. Ein Loopback-TCP-Socket erreicht dasselbe und ist
plattformübergreifend. Es bindet nur an 127.0.0.1 und ist niemals dem
Netzwerk ausgesetzt.
Installation
Erfordert GIMP 3.x (entwickelt gegen 3.2.4) und das mcp-Python-Paket.
Hinweis zur
mcp-Abhängigkeit. Dies zielt auf dasmcp1.x-SDK und ist aufmcp>=1.0,<2festgelegt. Version 2.0 entferntemcp.server.fastmcpund benannteFastMCPinMCPServerum; die Portierung darauf ist noch nicht abgeschlossen, und eine nicht festgelegte Installation zieht 2.x und schlägt beim Import fehl.
pip install -r requirements.txt
python install_plugin.py # install the bridge plug-in (optional)
python install_plugin.py --list # show detected GIMP config dirsDas Bridge-Plug-in wird nur für die Live-Steuerungs-Werkzeuge benötigt. Die Stapel- und Einzelbild-Werkzeuge funktionieren ohne Installation von etwas in GIMP.
Plug-in-Speicherort
install_plugin.py erkennt, welche GIMP-3.x-Konfigurationsverzeichnisse tatsächlich
existieren, anstatt eine Version hart zu codieren. Auf Windows ist das:
%APPDATA%\GIMP\3.2\plug-ins\gimp-mcp-bridge\gimp-mcp-bridge.pyBeachte, dass es das versionierte Verzeichnis ist (3.2 für GIMP 3.2, nicht 3.0), und
GIMP 3 erfordert, dass jedes Plug-in in einem Ordner sitzt, dessen Name mit der .py-Datei übereinstimmt.
Auf Linux und macOS sucht der Installer in ~/.config/GIMP/3.x/ bzw.
~/Library/Application Support/GIMP/3.x/.
MCP-Server registrieren
Die Installation des Pakets stellt ein gimp-mcp-Konsolenskript bereit, das die
sauberste Möglichkeit zur Registrierung ist, da es nicht von einem Arbeitsverzeichnis abhängt:
python -m venv .venv
.venv/Scripts/python -m pip install -e . # .venv/bin/python on Unix{
"mcpServers": {
"gimp": {
"type": "stdio",
"command": "/path/to/gimp-mcp/.venv/Scripts/gimp-mcp.exe",
"args": []
}
}
}Mit Claude Code ist das Äquivalent in einer Zeile:
claude mcp add gimp --scope user -- /path/to/gimp-mcp/.venv/Scripts/gimp-mcp.exeDas direkte Ausführen des Moduls funktioniert ebenfalls, wenn mcp in diesem
Interpreter importierbar ist:
{
"mcpServers": {
"gimp": {
"command": "python",
"args": ["-m", "gimp_mcp"],
"cwd": "/path/to/gimp-mcp"
}
}
}Optionale Umgebungsvariablen:
Variable | Zweck |
| Vollständiger Pfad zu |
|
|
| Bridge-Port, Standard |
Werkzeuge
Inspektion
Werkzeug | Zweck |
| Prüft, ob GIMP erreichbar ist; meldet beide Backends. Beginne hier, wenn etwas nicht stimmt. |
| Abmessungen, Ebenen, Orientierung. Abmessungen wie angezeigt. |
| Validierung gegen eine Dimensionsspezifikation; bestanden/nicht bestanden mit gemessenen Abmessungen und einer verständlichen Begründung. |
Einzelbild
Werkzeug | Zweck |
| Exaktes Pixelrechteck. Lehnt außerhalb des Bereichs liegende Werte ab, anstatt stillschweigend zu klemmen. |
| Größtes Quadrat; |
| Ziel-Seitenverhältnis (1.0 Quadrat, 1.3333 für 4:3, 1.7778 für 16:9), maximale Fläche. |
| Nach Breite, Höhe oder |
| Helligkeit/Kontrast, beschränkt auf -0.5..0.5. |
| In einem Durchgang: Orientierung durch Zuschneiden korrigieren, auf ein Minimum hochskalieren, auf ein Maximum herunterskalieren, optionale Korrektur. |
| Benutzerdefinierte Operations-Pipeline in einem Durchgang (eine JPEG-Neucodierung). |
Stapel
Werkzeug | Zweck |
| Beliebige Pipeline über einen Ordner. |
| Einen ganzen Ordner an eine Dimensionsspezifikation anpassen. |
| Schreibgeschützte Prüfung; Triage vor der Bearbeitung. |
Ein ganzer Stapel läuft innerhalb eines GIMP-Aufrufs. GIMPs Konsole benötigt mehrere Sekunden
zum Starten, daher wäre das Spawnen pro Datei langsam – gemessen bei ~2,4x
günstiger pro Datei für einen kleinen Ordner, und die Ersparnis wächst mit der Ordnergröße. Eine
Datei, die fehlschlägt, bricht den Lauf nicht ab; sie landet in errors und der Rest
wird fortgesetzt.
Live-Steuerung (benötigt das Bridge-Plug-in)
Werkzeug | Zweck |
| Was im laufenden GIMP geöffnet ist. |
| Flache Momentaufnahme der Leinwand, damit du sehen und iterieren kannst. |
| Beliebiger Python-Code im Live-Kontext; weise |
| Stoppt die Bridge, lässt GIMP geöffnet. |
Starte die Bridge in GIMP: Filter > Entwicklung > MCP-Bridge starten.
Bildspezifikationen
check_image_spec, fit_to_spec und ihre Stapel-Äquivalente teilen sich ein
Spezifikationsmodell. Jede Einschränkung ist optional – 0 bedeutet keine Begrenzung, und Orientierung
any bedeutet keine Orientierungsanforderung.
Feld | Werte |
| Pixel, |
| Pixel, |
|
|
fit_to_spec erfüllt eine Spezifikation in drei geordneten Schritten: Zuschneiden zur Korrektur der
Orientierung, Hochskalieren zur Erreichung des Minimums, Herunterskalieren zur Einhaltung des Maximums.
Bereits erfüllte Einschränkungen lassen den Bildausschnitt unangetastet.
// A square image at least 1000x1000, capped at 2000x2000
{ "orientation": "square", "min_width": 1000, "min_height": 1000,
"max_width": 2000, "max_height": 2000 }Farbanpassung ist bewusst begrenzt
adjust_image beschränkt Helligkeit/Kontrast auf -0.5..0.5 und lehnt
alles außerhalb ab, anstatt zu klemmen. Werte über etwa ±0.15 verändern sichtbar
den Charakter eines Fotos, was wichtig ist, wenn ein Bild ein reales
Subjekt getreu darstellen soll. Es gibt bewusst keine Sättigungsverstärkung
oder "Auto-Verbesserung".
Verifizierung
Führe die Suite aus:
python -m pytest tests/ -vTests, die echte Bilder benötigen, werden übersprungen, es sei denn, du weist sie auf einige hin:
export GIMP_MCP_TEST_IMAGE=/path/to/photo.jpg # ideally EXIF-rotated
export GIMP_MCP_TEST_REFERENCE=/path/to/photo-square.jpgGIMP_MCP_TEST_REFERENCE sollte ein unabhängig erstelltes zentriertes Quadrat-Zuschneiden
von GIMP_MCP_TEST_IMAGE sein – zum Beispiel von Hand in GIMP zugeschnitten. Der
Haupttest behauptet, dass crop_square diese Referenz reproduziert, anstatt
nur ohne Fehler zu laufen.
Auf dem Referenzfoto, das während der Entwicklung verwendet wurde (ein 4000x3000-JPEG mit EXIF-Orientierung 6, das als 3000x4000 angezeigt wird):
crop_square vs hand-made reference : mean abs diff 0.236, max 18, outliers 0.0014%
same crop via the bridge backend : mean abs diff 0.236, max 18, outliers 0.0014%Dieser Rest ist JPEG-Neucodierungsrauschen – alleiniges Neucodieren ergibt ~0.5 Mittelwert – nicht ein Geometrieunterschied, und beide Backends stimmen exakt überein.
Die Suite deckt auch die Meldung der angezeigten Orientierung, Orientierungs- und Mindestgrößenspezifikationen, abgelehnte Zuschnitte außerhalb des Bereichs, abgelehnte Anpassungen außerhalb des Bereichs, Helligkeit, die Pixel in die richtige Richtung bewegt, verkettete Pipelines, Seitenverhältnis-Zuschneiden, Stapel über einen Ordner, die schreibgeschützte Prüfung, klare Fehler für fehlende Dateien und einen vollständigen Durchlauf über das echte MCP-stdio-Protokoll ab.
Fehlerbehebung
gimp-console not found – setze GIMP_CONSOLE auf den vollständigen Pfad von
gimp-console-3.exe.
Bridge-Werkzeuge schlagen mit "Could not reach the GIMP bridge" fehl – GIMP ist nicht
geöffnet, oder die Bridge wurde nicht gestartet. Führe Filter > Entwicklung > MCP-Bridge starten aus.
gimp_status zeigt beide Backends gleichzeitig an.
Das Menüelement fehlt nach der Installation – starte GIMP neu; es scannt Plug-ins
nur beim Start. Bestätige das Layout als
plug-ins/gimp-mcp-bridge/gimp-mcp-bridge.py (der Ordnername muss mit dem
Dateinamen übereinstimmen).
Diagnose des Plug-ins – ein GIMP-Plug-in ist ein separater Prozess, dessen stderr
unsichtbar ist, wenn GIMP als GUI-App unter Windows läuft. Die Bridge schreibt in
bridge.log neben dem installierten Plug-in.
Ein Farbprofil-Dialog blockiert GIMP beim Start, wenn ein Bild mit einem eingebetteten Profil im GUI-Modus geöffnet wird. Er erscheint nicht im Headless-Modus, was ein weiterer Grund ist, warum Stapelarbeiten das Headless-Backend verwenden.
Stapel-Timeout – der Standardwert ist 600s für den gesamten Lauf; sehr große Ordner benötigen möglicherweise mehr.
Bekannte Einschränkungen
Live-Steuerung ist nur wenig erprobt. Sie ist nachweislich funktionsfähig (Bild öffnen, auflisten, Screenshot, Live-Bearbeitung und Zuschneiden über die Bridge mit identischer Ausgabe wie im Headless-Modus), wurde aber weitaus weniger genutzt als der Headless-Pfad. Behandeln Sie Headless als die vertrauenswürdige Variante.
Die Bridge führt konstruktionsbedingt beliebigen Python-Code aus. Sie ist nur über Loopback erreichbar und wird manuell statt automatisch gestartet, aber alles, was localhost auf der Maschine erreichen kann, kann GIMP steuern, während sie läuft. Beenden Sie sie, wenn sie nicht verwendet wird.
Der Bridge-Start blockiert den eigenen Plug-in-Prozess — das hält ihn am Leben. Er friert die GIMP-Oberfläche nicht ein, aber GIMP zeigt das Plug-in als laufend an.
Der GUI-Menüpunkt selbst ist nicht durch automatisierte Tests abgedeckt. Das Verfahren, das er aufruft, ist verifiziert; der Klickpfad nicht.
Nur Windows ist verifiziert. Die Codepfade sind plattformübergreifend und der Installer behandelt Linux/macOS-Konfigurationsverzeichnisse, aber keines von beiden wurde getestet.
Das
mcp2.x SDK wird noch nicht unterstützt -- siehe Hinweis unter Install.Keine KI-Hintergrundentfernung oder Stilübertragung. Einige vergleichbare Projekte werben damit, ohne eine funktionierende Implementierung dahinter zu haben; hier wird das bewusst nicht behauptet.
Anmerkungen zu Vorarbeiten
Die Aufteilung zwischen einem GIMP-seitigen Plug-in, das eine Bridge bereitstellt, und einem eigenständigen MCP-Serverprozess, der sich als Client mit ihr verbindet, ist eine natürliche Form für dieses Problem und wird auch von anderen GIMP-MCP-Projekten verwendet. Stapelverarbeitung und Pipelines im Preset-Stil sind bei mehreren üblich. Live-Canvas-Steuerung existiert anderswo über D-Bus und wurde hier durch Loopback-TCP für Windows-Unterstützung ersetzt. Von keinem von ihnen wurde Code kopiert; die Windows-spezifischen Details — der tatsächliche Plug-in-Pfad, die Lebensdauer des Plug-in-Prozesses, die Run-Callback-Signatur und das EXIF-Verhalten — wurden direkt gegen GIMP 3.2.4 ermittelt.
Lizenz
MIT — siehe LICENSE.
Available Tools
18 toolsadjust_imageA
Brightness/contrast touch-up.
Both values run -1.0..1.0 -- GIMP 3's real range for this operation -- and
are rejected outside it rather than clamped. enhance_image uses the same
range for contrast; they wrap the same GIMP call.
Useful values are far smaller than the limits. GIMP 3 applies this in linear light, so it bites harder than the same number did in GIMP 2.x, and anything much past +/-0.1 visibly changes the character of a photo. Keep adjustments small when the image needs to represent a real subject faithfully.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| quality | No | ||
| contrast | No | ||
| brightness | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well. It discloses that out-of-range values are rejected rather than clamped, that the operation runs in linear light and therefore has a stronger effect than in GIMP 2.x, and that values past +/-0.1 visibly change the image. This goes well beyond a generic 'adjusts brightness and contrast.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then adds only high-value behavioral and practical details. Every sentence earns its place: the range/rejection rule, the relationship to enhance_image, the linear-light warning, and the guidance to keep adjustments small.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no annotations and 0% schema coverage, the description is nearly complete: it explains the main operation, the key parameters, edge-case behavior, and how values behave differently in GIMP 3. It misses only the optional backend and quality semantics, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides rich semantics for brightness and contrast (range, rejection behavior, practical limits), but says nothing about backend or quality, leaving the quality default of 0.92 unexplained. The required input_path and output_path are self-evident from their names, so the gap is mainly in the optional parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Brightness/contrast touch-up,' a specific verb and resource that immediately identifies the tool's function. It also references enhance_image, providing a sibling distinction even though the exact selection criteria are not spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful operational guidance about value ranges and suggests keeping adjustments small, but it does not explicitly state when to choose adjust_image over enhance_image or other siblings. The mention that both wrap the same GIMP call is informative but stops short of routeing the agent to the right tool in a specific scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_check_image_specA
Audit a folder: which images already satisfy a specification.
Read-only; writes nothing. Use it to triage a folder before editing.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| pattern | No | ||
| input_dir | Yes | ||
| max_width | No | ||
| min_width | No | ||
| recursive | No | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and explicitly states 'Read-only; writes nothing.' This is a valuable, non-obvious disclosure about side effects. It does not go further into permissions or edge cases, but it covers the primary behavioral concern.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and followed by a concise safety/usage note. Every word earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, safety, and a clear use case, and an output schema exists so return values need not be described. However, with 9 parameters and zero schema descriptions, the description leaves parameter semantics and default behavior underspecified. It is adequate for a first correct invocation but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the tool has 9 parameters. The description only mentions 'specification' in general terms and does not explain how parameters like pattern, recursive, orientation, or min/max dimensions map to that specification. Parameter names are suggestive, but the description itself provides almost no semantic value beyond the schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Audit') and resource ('a folder') plus the precise condition ('which images already satisfy a specification'). The batch/folder framing clearly distinguishes it from singular tools like check_image_spec.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit use context: 'Use it to triage a folder before editing.' This tells an agent when it is appropriate, though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_fit_to_specC
Make every image in a folder satisfy one dimension specification.
The common bulk case: point it at a folder of photos and get conforming copies, with each file's final dimensions and pass/fail reported.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| suffix | No | _out | |
| backend | No | ||
| pattern | No | ||
| quality | No | ||
| input_dir | Yes | ||
| max_width | No | ||
| min_width | No | ||
| recursive | No | ||
| max_height | No | ||
| min_height | No | ||
| output_dir | Yes | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose key behavior: it produces conforming copies, operates on every image in a folder, and reports final dimensions and pass/fail. However, it does not clarify how the dimension constraints interact, whether originals are left untouched, naming/overwrite behavior, or recursive handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core function and the second gives the primary use case and expected reporting. It is not padded, though it sacrifices useful detail to achieve that brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter batch tool with 0% schema coverage and no annotations, the description is too thin. It gives a scenario and outcome but leaves the agent without enough information about parameter semantics, alternatives, and behavior constraints; the presence of an output schema only partially offsets the missing return-value detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 13 parameters, and the description adds no parameter-level meaning. 'One dimension specification' hints at constraints but does not explain input_dir, output_dir, min/max width/height, orientation, pattern, quality, backend, suffix, recursive, or anchor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete operation: make every image in a folder satisfy a dimension specification, and adds bulk/copy/reporting context. It does not explicitly name a sibling such as fit_to_spec or batch_check_image_spec, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'common bulk case' implies this is for folder-wide operations and contrasts with single-image tools, but no explicit when-to-use/when-not-to-use guidance or alternative tool names are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_processA
Apply the same operations to every image in a folder.
All files are handled inside a single GIMP session, so a large folder
costs one GIMP startup rather than one per file. A file that fails does
not abort the run: it is reported in errors and the rest continue.
operations is a JSON list, same format as process_image. Set suffix to
e.g. "_out" to avoid overwriting, and output_format to "jpg"/"png" to
convert.
| Name | Required | Description | Default |
|---|---|---|---|
| suffix | No | ||
| backend | No | ||
| pattern | No | ||
| quality | No | ||
| input_dir | Yes | ||
| recursive | No | ||
| operations | Yes | ||
| output_dir | Yes | ||
| output_format | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses single-session processing, failure isolation with reported errors, and overwrite avoidance through suffix, all of which are valuable beyond the schema. It does not cover every behavioral nuance, but the core runtime traits are well described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: one sentence for purpose, one short paragraph for runtime behavior, and one for parameter guidance. Every sentence contributes useful information without repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Core behavior, error reporting, and key parameter usage are covered, and an output schema is indicated to exist. However, with 9 parameters and zero schema descriptions, several parameters such as backend, pattern, recursive, and quality remain under-specified for an agent to invoke the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain operations as a JSON list in process_image format, suffix for avoiding overwrites, and output_format for conversion, but it leaves backend, pattern, recursive, quality, input_dir, and output_dir semantically unexplained despite having 9 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: apply the same operations to every image in a folder. This clearly distinguishes the tool from single-image tools like process_image and from specialized batch tools like batch_fit_to_spec or batch_check_image_spec.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the batch usage context clear by explaining that a large folder costs one GIMP startup rather than one per file, which strongly implies the batch counterpart to process_image. It does not explicitly name alternatives or say when not to use this tool, but the context is not misleading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_image_specA
Check an image against a dimension specification.
Every constraint is optional: 0 means "no limit", and orientation "any" means no orientation requirement. Valid orientation values are: any, square, landscape, portrait, square_or_landscape, square_or_portrait.
Returns pass/fail with the actual measured dimensions and a plain-language reason for each failure. Useful for validating images against a publishing platform's requirements, a print size, or an asset pipeline's conventions before spending time editing.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| backend | No | ||
| max_width | No | ||
| min_width | No | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that every constraint is optional, that 0 means no limit, that 'any' means no orientation requirement, and that the tool returns pass/fail with measured dimensions and failure reasons. This is useful behavioral detail, though it does not cover error cases or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a clear one-line purpose, then a compact explanation of constraint semantics and orientation values, then return behavior and use cases. Every sentence earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 7 parameters, no annotations, and 0% schema coverage, the description covers the key semantics well: optional constraints, 0 meaning, orientation values, and return payload. It does not explain the backend parameter or explicitly compare itself with sibling validation tools, but an output schema exists to cover return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add real meaning by explaining the 0-as-no-limit convention for numeric constraints and enumerating valid orientation values. However, the optional 'backend' parameter is never explained, which is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check an image against a dimension specification.' It also states the return contract (pass/fail with dimensions and reasons), which clearly distinguishes it from siblings like crop_image, resize_image, or inspect_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: validating images against publishing requirements, print sizes, or asset pipeline conventions before editing. It does not explicitly exclude alternatives or name sibling tools such as fit_to_spec or inspect_image, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_imageB
Crop to an exact pixel rectangle.
x/y are the top-left offset in DISPLAYED orientation. Fails clearly if the rectangle falls outside the image rather than silently clamping.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| width | Yes | ||
| height | Yes | ||
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and delivers important details: x/y are interpreted in DISPLAYED orientation and out-of-bounds rectangles fail clearly instead of silently clamping. It could disclose overwrite behavior or backend semantics, but the core failure and orientation behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the most decision-relevant information is front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and 0% schema coverage, the description is too thin. It omits backend and quality meaning and does not state whether output_path overwrites existing files, leaving an agent to guess on non-obvious options even though required paths are given in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% parameter descriptions, so the description must compensate. It adds meaning for x/y as top-left offsets in displayed orientation and implies width/height are pixel-based, but backend and quality are left unexplained, and input_path/output_path semantics are still only inferable from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Crop to an exact pixel rectangle', which clearly identifies the operation and resource. It also distinguishes this tool from siblings like crop_to_aspect and crop_square by emphasizing exact pixel dimensions, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus crop_to_aspect, crop_square, or other siblings. The phrase 'exact pixel rectangle' implies a use case, but no when-to-use or when-not-to-use context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_squareB
Crop to the largest possible square.
anchor picks which part of the frame to keep: center (default), top, bottom, left, right, or a corner such as topleft.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the behavioral burden. It does explain that the crop keeps the largest possible square and that the anchor determines which part of the frame survives, including the center default. However, it does not disclose what backend or quality do, whether files are overwritten, or what side effects occur beyond writing output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. The main action appears in the first sentence, and the second sentence earns its place by clarifying the anchor parameter. There is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimally adequate for a simple crop operation: it states the core behavior and the key anchor option. But with five parameters and no annotations, important details like backend and quality semantics are missing, and no usage context versus sibling tools is provided. The presence of an output schema lowers the burden for return-value documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's silence. It adds helpful semantics for anchor by enumerating valid values and the default. It leaves backend and quality completely unexplained, and input_path/output_path relationships are only implicit from the tool's name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action and resource: 'Crop to the largest possible square.' It also adds meaningful detail about the anchor parameter. It does not explicitly name or differentiate from sibling tools like crop_to_aspect or crop_image, but the square-only behavior is reasonably distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use crop_square versus the many sibling cropping/resizing tools. The description does not state prerequisites, exclusions, or conditions that would help an agent select this tool over crop_to_aspect, crop_image, or adjust_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crop_to_aspectB
Crop to a target aspect ratio (width/height), keeping maximum area.
Use 1.0 for square, 1.3333 for 4:3, 1.5 for 3:2, 1.7778 for 16:9.
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | Yes | ||
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| input_path | Yes | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It discloses the 'keeping maximum area' behavior and explains the ratio meaning, which is useful. However, it does not mention important behaviors like default anchoring, quality handling, or whether the operation modifies the input image.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core behavior, and every sentence adds value. The ratio examples are practical and directly aid correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, 6 parameters, and zero schema description coverage, the description is too sparse. It fails to clarify optional parameters or provide enough context to confidently invoke the tool beyond the basic ratio and paths.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It meaningfully explains the ratio parameter with examples, but provides no additional semantics for anchor, backend, quality, input_path, or output_path. This leaves several parameters underdocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool crops to a target aspect ratio while keeping maximum area, with concrete ratio examples. It is specific enough to be understood, though it does not explicitly differentiate itself from sibling tools like crop_image or crop_square.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides common ratio values but offers no guidance on when to choose this tool over alternatives such as crop_image, crop_square, or fit_to_spec. There are no exclusions or explicit usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enhance_imageA
Tone and detail enhancement in one pass.
gamma lifts midtones and shadows via levels, leaving the black and white points alone so nothing clips. 1.0 = off. contrast GIMP 3 native -1..1. 0 = off. saturation -100..100. 0 = off. sharpen high-pass sharpen blended back at this percent opacity, 0 = off. Preferred over unsharp mask, which haloes. sharpen_radius blur radius in pixels for the high pass (default 8).
Contrast bites harder than the same nominal value did in GIMP 2.x, because GIMP 3 runs the operation in linear light: 2.x's "+12" is roughly 0.020 here, not 0.094. Calibrate against output rather than remapping an old number.
| Name | Required | Description | Default |
|---|---|---|---|
| gamma | No | ||
| backend | No | ||
| quality | No | ||
| sharpen | No | ||
| contrast | No | ||
| input_path | Yes | ||
| saturation | No | ||
| output_path | Yes | ||
| sharpen_radius | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it explains that gamma avoids clipping by preserving black/white points, sharpen is high-pass blended back at a given opacity, and contrast runs in GIMP 3 linear light, with explicit calibration caveats. This gives agents a strong model of the tool's actual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded, each parameter gets a compact line, and the GIMP 3 calibration note earns its place. There is no filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations, the description is thorough on the core enhancement behavior and parameter semantics. It is not fully complete because backend and quality are unexplained, and no sibling-tool comparison is given, leaving some ambiguity in tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates well by explaining gamma, contrast, saturation, sharpen, and sharpen_radius with scales and defaults. However, backend and quality are left undocumented in the description, so their semantics remain ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation ('Tone and detail enhancement in one pass') and elaborates on what each control does, so an agent understands the tool's role. It does not explicitly differentiate it from sibling tools like adjust_image or process_image, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to choose enhance_image over the many sibling image tools, nor exclusions for when not to use it. The parameter-level note about preferring high-pass sharpen over unsharp mask is useful but not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fit_to_specA
Transform an image until it satisfies a dimension specification.
Crops to fix the orientation if required, upscales to reach a minimum size, downscales to respect a maximum, and optionally applies a light touch-up -- all in one pass, so the JPEG is re-encoded only once. Images already satisfying a constraint keep their framing.
Example: to produce a square image at least 1000x1000, pass orientation="square" with min_width=1000 and min_height=1000.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | center | |
| backend | No | ||
| quality | No | ||
| upscale | No | ||
| contrast | No | ||
| max_width | No | ||
| min_width | No | ||
| brightness | No | ||
| input_path | Yes | ||
| max_height | No | ||
| min_height | No | ||
| orientation | No | any | |
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does well: it discloses cropping, upscaling, downscaling, optional touch-up, single re-encode of JPEG, and preservation of already-compliant framing. It leaves some specifics unstated (e.g., overwrite behavior, failure conditions, handling of non-JPEG inputs), which prevents a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, all informative: the first sentence states the purpose, the second explains the mechanism, and the example grounds the parameters. No filler or repetition of schema field names, and the key operations are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no annotations and no schema descriptions, the description covers the core workflow adequately but omits enough optional-parameter semantics to be fully self-contained. The presence of an output schema means return values need not be described, but the parameter gaps and lack of sibling differentiation leave clear holes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters; it gives meaning to orientation, min_width, min_height, and the max constraints via the example and operation summary. However, many of the 13 parameters (anchor, backend, quality, upscale, contrast, brightness, max_width/max_height defaults) are not explained in the description or schema, leaving significant inference required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Transform an image until it satisfies a dimension specification,' then enumerates the exact operations (crop for orientation, upscale, downscale, optional touch-up) and gives a concrete square-image example. This clearly separates the tool from generic resize/crop siblings by emphasizing the all-in-one constraint-satisfaction behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied clearly: call this tool when an image must meet dimension constraints (min/max width/height, orientation) in one pass. However, it never explicitly tells an agent when to prefer fit_to_spec over sibling tools like crop_to_aspect, resize_image, or process_image, nor states when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gimp_statusA
Check that GIMP is reachable and report its version.
Use this first if anything seems wrong. Reports both the headless backend and whether the live bridge plug-in is running inside an open GIMP.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses what the tool reports (headless backend status and live bridge plug-in state) beyond a simple reachability check, which adds useful behavioral context. It doesn't explicitly state non-destructive behavior, but 'check' and 'report' strongly imply a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the core purpose front-loaded and a clear usage directive. There is no filler or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status tool with one optional parameter, an output schema, and no required inputs, the description covers purpose, usage, and report contents. The only notable gap is the undocumented 'backend' parameter, but the tool can be correctly invoked with no arguments, so the description is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions the single optional 'backend' parameter or its meaning. An agent cannot determine what value to pass or why the parameter exists. The description mentions 'headless backend' as reported output, but that doesn't clarify the parameter's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Check that GIMP is reachable and report its version.' It also distinguishes itself from the many sibling processing tools by being a diagnostic/status tool, not an image manipulation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Use this first if anything seems wrong.' This tells the agent when to invoke it, though it doesn't name specific alternatives or exclusions. The first-step diagnostic role is clear enough for practical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_imageB
Report an image's dimensions, layers, and orientation.
Dimensions are reported as DISPLAYED (EXIF orientation applied), which is what a viewer sees -- not necessarily how the pixels are stored.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| backend | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that dimensions are reported in displayed form after EXIF orientation is applied, not as raw stored pixels. This is a meaningful behavioral nuance, though it could also mention that this is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief, front-loads the main purpose, and includes only the essential EXIF detail. Every sentence contributes value and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple inspection tool with an output schema, the description covers the core result and a key display nuance. However, it lacks an explanation of the backend parameter and gives no guidance on when this tool should be preferred over related tools, leaving some gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it adds no information about the path or backend parameters. The backend parameter is completely unexplained, and the description only implies the image is referenced by path without discussing either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Report') and a specific resource ('an image's dimensions, layers, and orientation'). This distinguishes it from sibling tools that crop, resize, or process images rather than inspect metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided, and it does not mention alternatives such as check_image_spec or other inspection-like siblings. The intended context must be inferred from the tool name and description, so there is no explicit usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_list_imagesA
List the images currently open in the running GIMP.
Requires the bridge plug-in (Filters > Development > Start MCP Bridge).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It usefully reveals that the tool depends on the bridge being started and reads current live GIMP state. 'List' implies a non-mutating operation, and the output schema covers return structure, so this is transparent enough for a simple read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the core purpose front-loaded and the required setup immediately after. Every sentence earns its place and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter live listing tool, the description covers what it does, the environment prerequisite, and leaves return details to the output schema. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is nothing for the description to clarify about arguments. Per the baseline for no-parameter tools, this is appropriately handled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a specific resource ('images currently open in the running GIMP'), making the tool's function immediately clear. It is naturally distinguishable from sibling tools like crop_image, gimp_status, and inspect_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys that this tool is for querying the current live set of open images in GIMP, and it flags the prerequisite of the bridge plug-in. It does not explicitly compare to alternatives, but the zero-parameter live-listing purpose is self-evident enough for an agent to know when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_run_pythonA
Execute Python inside the running GIMP and return result.
The Gimp module and every operation helper are already in scope. Assign to
a variable named result to return a value. Escape hatch for anything the
typed tools above do not cover; requires the bridge plug-in.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It usefully explains that GIMP modules and helpers are in scope and that a `result` variable must be assigned to return a value. However, it does not disclose that arbitrary Python execution can mutate or destroy GIMP state, crash the session, or have irreversible side effects, which is a significant transparency gap for an unbounded execution tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the main purpose, and every sentence earns its place: execution semantics, scope context, return-value convention, use-case, and prerequisite. No filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema, the description does not need to explain return structure. It covers the parameter, the execution environment, the return mechanism, and the prerequisite. The main missing piece is a warning about the destructive or uncontrolled nature of raw Python execution, which matters for a tool of this complexity and power.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does meaningfully: it tells the agent that `code` is Python to execute inside GIMP, that modules are already in scope, and that assigning to `result` controls the return value. This gives the single parameter real semantic grounding beyond the bare name 'code'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb+resource: 'Execute Python inside the running GIMP and return `result`'. It also differentiates itself from the typed sibling tools by describing itself as an 'escape hatch for anything the typed tools above do not cover', so an agent can distinguish it at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool: when the typed tools do not cover the needed operation. It also establishes the prerequisite that the bridge plug-in must be present. This provides clear when-to-use context and points to the sibling tools as the preferred alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_screenshotA
Save a flattened snapshot of an image open in the running GIMP.
Lets you see the current state of a document you are editing live, then iterate on it. Requires the bridge plug-in.
| Name | Required | Description | Default |
|---|---|---|---|
| max_edge | No | ||
| image_index | No | ||
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does disclose key behavioral traits: it saves a flattened/live snapshot and depends on the bridge plug-in. It does not discuss whether the original image is modified or whether an existing output file is overwritten, but 'snapshot' implies a non-destructive capture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the action, and the bridge plug-in requirement is worth stating. The sentence about seeing current state and iterating is somewhat redundant with 'snapshot' and 'live,' but it adds useful intent context without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters, no annotations, and 0% schema description coverage, so the description must compensate. It gives a clear purpose and prerequisite but leaves max_edge and image_index undefined, and does not explain output_path semantics or file format. An agent could invoke it with defaults, but not understand the full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain any of the three parameters. Output_path is only implied by 'Save', while max_edge and image_index—especially which open image is captured—are not described at all, so the description adds no parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Save a flattened snapshot of an image open in the running GIMP' names a specific verb, a concrete resource, and a key behavior (flattening). This clearly separates it from siblings like inspect_image and live_list_images, so an agent can distinguish the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use it when you need to see the current state of a live document and iterate on it. It also states a prerequisite ('Requires the bridge plug-in'), but it does not explicitly mention alternative tools or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_stop_bridgeA
Stop the bridge inside the running GIMP, leaving GIMP itself open.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that the bridge is stopped and that GIMP remains open, which are the two key behavioral consequences. It does not mention edge cases like stopping an already-stopped bridge, but for a zero-parameter control action this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the core action and adds the important clarifying detail about GIMP remaining open. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with a simple, well-defined action, the description is complete. It tells the agent exactly what happens and what does not happen. The output schema is present, so return-value details are not required in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description correctly includes no parameter-specific details because none are needed. There is no schema information to supplement or contradict.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Stop the bridge inside the running GIMP'. It also explicitly clarifies that GIMP itself remains open, which disambiguates this from closing GIMP. This is a clear, distinct purpose among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the context: the bridge is running inside GIMP, and the tool stops only the bridge. It does not explicitly name alternatives or when-not-to-use, but no sibling tool appears to perform a similar stop action, so the context is sufficient for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
process_imageA
Apply a custom sequence of operations in one pass.
operations is a JSON list, e.g.
[{"op":"crop_square","anchor":"center"},
{"op":"resize","max_edge":2000},
{"op":"adjust","brightness":0.05}]
Available ops: crop, crop_square, crop_aspect, resize, adjust, autocrop, flatten, fit_spec, enhance. Running them as one pipeline re-encodes the JPEG only once, which avoids stacking compression artefacts.
The optional spec arguments are checked against the FINAL result and
reported under spec, so a pipeline that both reshapes and edits an image
can be validated without a second pass over it.
| Name | Required | Description | Default |
|---|---|---|---|
| backend | No | ||
| quality | No | ||
| max_width | No | ||
| min_width | No | ||
| input_path | Yes | ||
| max_height | No | ||
| min_height | No | ||
| operations | Yes | ||
| orientation | No | any | |
| output_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It adds meaningful details: single-pass processing, one-time JPEG re-encoding, and that optional spec arguments are validated against the final result and reported under `spec`. It does not cover failure modes or input format restrictions, but the core execution behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: a one-sentence purpose, a concrete operations example, a concise list of available ops, and two short paragraphs explaining the pipeline benefit and spec-checking behavior. Every sentence earns its place, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no annotations and zero schema descriptions, the description is not fully complete. It covers the central `operations` parameter and the pipeline concept well, but it leaves the optional spec arguments and other tuning parameters under-specified, which an agent would need to invoke the tool correctly for advanced use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does a good job on `operations`, showing a full JSON example and listing valid op names, and it alludes to 'spec arguments'. However, it does not explain other parameters such as `quality`, `backend`, `max_width`, `min_width`, `max_height`, `min_height`, or `orientation`, leaving significant semantic gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Apply a custom sequence of operations in one pass.' It explicitly lists the available operations, which map directly to the sibling individual-operation tools, so an agent can tell that this is the composite/pipeline counterpart to crop_image, resize_image, adjust_image, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete reason to choose this tool over chaining siblings: running operations as one pipeline re-encodes the JPEG only once and avoids stacking compression artifacts. It does not explicitly state when to prefer a single-operation sibling, but the 'custom sequence' framing and op list make the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resize_imageA
Resize an image.
Give width, height, both, or max_edge (longest side, aspect preserved). With preserve_aspect and both dimensions, the image is fitted inside the box rather than distorted.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | ||
| height | No | ||
| backend | No | ||
| quality | No | ||
| max_edge | No | ||
| input_path | Yes | ||
| output_path | Yes | ||
| preserve_aspect | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It usefully explains aspect-ratio preservation and that both dimensions with preserve_aspect fits the image inside the box rather than distorting it. However, it says nothing about defaults, behavior when no sizing parameter is provided, backend handling, quality interpretation, or whether output files are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loaded, with 'Resize an image' first and only the most decision-relevant parameter guidance following. Every sentence earns its place and there is minimal fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and zero schema descriptions, the description covers the core sizing behavior well but leaves gaps around backend, quality, validation, and edge cases. The presence of an output schema reduces the need to explain return values, but the description is still only moderately complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It adds meaning for width, height, max_edge, and preserve_aspect, and the required input/output paths are reasonably self-explanatory from their names. But backend and quality receive no explanation, leaving two parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line 'Resize an image' states a specific verb and resource, and the parameter combinations (width, height, max_edge, preserve_aspect) make the tool's purpose clear. It does not explicitly distinguish itself from sibling tools like crop_image or adjust_image, but 'resize' is distinct enough on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to use the tool's sizing modes ('Give width, height, both, or max_edge') and the behavior of preserve_aspect, which is operational guidance. However, it does not say when to prefer this tool over alternatives such as crop_to_aspect, fit_to_spec, or process_image.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.1.0- First observed
adjust_image - First observed
batch_check_image_spec - First observed
batch_fit_to_spec - First observed
batch_process - First observed
check_image_spec - First observed
crop_image - First observed
crop_square - First observed
crop_to_aspect - First observed
enhance_image - First observed
fit_to_spec - First observed
gimp_status - First observed
inspect_image - First observed
live_list_images - First observed
live_run_python - First observed
live_screenshot - First observed
live_stop_bridge - First observed
process_image - First observed
resize_image
TDQS
Scored across 18 tools
Several tools overlap in purpose: adjust_image and enhance_image share the same contrast call, and process_image/batch_process can reproduce the effects of most individual editing tools. The descriptions generally clarify scope, but an agent could easily hesitate between a dedicated single-op tool and its pipeline equivalent.
Most tools follow a clear verb_object snake_case pattern such as crop_image, resize_image, inspect_image, and batch_check_image_spec. The batch_ and live_ prefixes are applied consistently, with only gimp_status and fit_to_spec deviating slightly from the otherwise predictable pattern.
At 18 tools the server is slightly above the ideal 3-15 range, but the count is justified by the distinct clusters: single-image operations, spec checking/fitting, batch variants, and live GIMP bridge tools. The single/batch pairs add surface area but each serves a real workload.
The toolset covers the core image pipeline well: inspect, validate, crop, resize, adjust, enhance, process in one pass, and batch over folders. Obvious gaps like rotation or flipping are absent, but live_run_python and process_image provide workarounds for most missing operations.
Maintenance
Related MCP Connectors
Image toolkit: resize, compress, crop, watermark, convert, rotate, EXIF read/strip.
- MochifyOAuthapp.mochify
Image and PDF toolkit: convert to AVIF/WebP/JXL, resize, crop, remove backgrounds, optimize PDFs.
Image toolkit: resize, convert, compress, crop, metadata, hashes, favicons, OCR, QR codes, barcodes.
181LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to perform GIMP-style image operations such as open, resize, crop, flip, rotate, blur, desaturate, text overlay, export, and batch processing via MCP tools, supporting both mock (Pillow) and live GIMP backends.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to control GIMP for image editing tasks such as opening, resizing, filtering, exporting, and batch processing images through Python-Fu scripting.36 npmMIT
- AlicenseBqualityBmaintenanceEnables controlling GIMP 3 locally through natural language, providing tools for image editing, layer management, selections, and PDB procedure invocation. Keeps all images and files on the user's machine with a local-first, secure design.401GPL 3.0
- AlicenseBqualityAmaintenanceEnables AI agents to operate GIMP 3 end-to-end: open and inspect images, call every PDB procedure, apply GEGL filters destructively or as layer effects, measure pixels, render before/after/diff comparisons, cut out subjects with AI segmentation, and run multi-step recipes across folders.3265 PyPI4Apache 2.0