Skip to main content
Glama

Editor Screenshot

editor_screenshot

Capture a screenshot from the Godot editor's 3D/2D viewport, active camera, or running game to inspect the current scene state.

Instructions

Capture a screenshot of the Godot editor viewport or running game.

Picking a source: the default "viewport" captures the editor's 3D viewport, which is empty if the edited scene has no Node3D anywhere in the tree (or no scene is open). Those cases return EDITOR_NOT_READY with error.data = {editor_state: "viewport_not_3d", scene_root_type} and an actionable error.message — switch to "cinematic" if the scene has a Camera3D, or open a scene with 3D content.

Sources:

  • "viewport" (default): editor 3D viewport. Requires Node3D content in the edited scene (root or any descendant); see above for the no-3D-content / no-scene error shape.

  • "viewport_2d": editor 2D viewport. Use for 2D scenes. Not compatible with view_target/coverage/elevation/azimuth/fov.

  • "cinematic": render edited scene through its active Camera3D (no editor gizmos). Prefers a Camera3D marked current; falls back to the first Camera3D found in a depth-first walk. NODE_NOT_FOUND only when the scene contains no Camera3D at all.

  • "game": running game's framebuffer (only when project is running). A backgrounded/minimized game window freezes its main loop; the capture then returns the last rendered frame with stale_frame: true and a note in the metadata — focus the game window and retry for a current frame. GAME_HELPER_TIMEOUT means the game process never replied at all (nothing rendered yet, main thread blocked, or helper dead) — focus the window and retry, or use game_command to confirm liveness.

include_image=True (default) returns an MCP ImageContent block. view_target (comma-separated Node3D paths) reframes editor camera; AABB metadata always returned. coverage=True with view_target captures perspective + orthographic top-down references.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fovNoCamera FOV in degrees. Tight 20-30 = zoom; 60-75 = context.
sourceNo"viewport" | "viewport_2d" | "cinematic" | "game". Default "viewport".viewport
azimuthNoCamera azimuth in degrees (0=front, 90=right).
coverageNoWith view_target, capture two reference shots + AABB.
elevationNoCamera elevation in degrees (0=level, 90=overhead).
session_idNoOptional Godot session to target. Empty = active session.
user_promptNoOptional context from the agent that requested the capture. With Vision Routing enabled it is sent alongside the image so the vision model can describe what the agent is looking for.
view_targetNoNode3D scene path(s) to frame, comma-separated.
include_imageNoReturn image data. Default True.
max_resolutionNoLongest-edge resolution. Default 640. 0 = full res.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv3.1.4
  2. Removedv3.0.7
  3. First observedv2.9.1

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden and delivers: error shapes (EDITOR_NOT_READY with data, NODE_NOT_FOUND, GAME_HELPER_TIMEOUT), fallback behavior (cinematic falls back to first Camera3D), stale-frame handling for backgrounded games, and compatibility constraints. It also discloses that AABB metadata is always returned and what include_image does. This is thorough and non-contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured, starting with a one-line summary, then a 'Picking a source' section with clear bullets for each source, followed by parameter clarifications. Every sentence provides unique value, but the length might be slightly intimidating; however, given the tool's complexity, it is justified and efficiently formatted with bullet lists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 10 parameters, no output schema, and no annotations. The description covers all necessary context: source behavior, error codes with data, fallbacks, compatibility, return format (MCP ImageContent, metadata), and special notes for backgrounded games. Parameters like user_prompt and session_id are adequately explained. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all 10 parameters with 100% coverage, so the baseline is 3. The description adds valuable semantic context: fov ranges ('Tight 20-30 = zoom; 60-75 = context'), view_target reframing, coverage capturing perspective + orthographic references, and source-specific constraints. This enhances understanding beyond the schema without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states the exact action and resources ('Capture a screenshot of the Godot editor viewport or running game'). It immediately distinguishes between editor and game captures, and the four named sources make the scope crystal clear. No sibling tool does screenshots, so there is no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit selection criteria for each source: 'use for 2D scenes' (viewport_2d), 'switch to cinematic if the scene has a Camera3D', and game source only when running. It also states explicit exclusions, e.g., viewport_2d is not compatible with view_target/coverage/elevation/azimuth/fov. Error states include actionable next steps, guiding the agent on retry or alternate sources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.