capture
Render an image or collage from the engine, or read the frame it presented (source='presented': the game view exactly as the player saw it, with its real TAA, motion blur, exposure and render scale; source='window': the whole window). Pick WHERE with 'source' (and aim it with 'viewpoint' + 'basis': front/top/left/iso against the subject's own axes or the world's), WHAT with 'pass' (motion_vectors to see whether something moves, normal to see whether a surface is correct), HOW it projects with 'projection' ('orthographic' keeps parallel edges parallel and equal sizes equal: the one to verify a shape under), and which render layers with 'renderLayers'. 'isolate' draws the subject without the other geometry. mode='collage' makes a grid whose cells vary by 'setups', by 'viewpoints', by 'passes', or over 'duration': a shot list's whole set-ups side by side, six sides of an object, or one view under four passes, in a single image. Every cell reports 'luma': the tone of the pixels it encoded, in code values on the 0-255 scale: min/max/mean, the percentiles p1 p5 p50 p95 p99, 'span' (max-min), 'spread' (p95-p5), and 'crushed'/'clipped', the shares sitting at 0 and at 255. Read 'spread' to answer whether a shot is legible: a subject can be modelled, lit and drawn and still arrive inside a handful of code values, which a mean cannot tell from a picture with something in it. A capture is taken whenever the engine reaches it rather than at the instant the call was sent, and the world keeps running between calls, so two shots of a running scene are separated by the real seconds between them; pause it or slow it when the picture has to catch a specific moment.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| as | No | Your own name, when you are one of several agents sharing one connection to this engine (the sub-agents one agent session runs share its connection, so without it the engine sees all of them as one caller). Pass the same name on every call you make. The engine then treats you as a caller of your own: your roster name is this name, and your play session, the play-shadow writes and refusals that say who wrote what, engine.modeOwner, your shell session and your notices are yours alone. Lowercase letters, digits, '-' and '_'. Leave it out when you are the only agent on your connection. | |
| fov | No | Vertical field of view in degrees, stated for this capture alone. Omit to take the lens of whatever is being captured: a capture of a camera ('main', 'screen', 'viewport', 'editor', 'camera') renders at that camera's own fov, and one standing at a position or orbiting an entity renders at 60. | |
| grid | No | Collage layout as 'COLSxROWS' (e.g. '3x3', '2x2', '4x2'). Omit to have it worked out from the number of cells; a grid too small for them is an error rather than a silent crop. | |
| mode | No | 'image' = one frame. 'collage' = a GRID of frames in one image, where what varies from cell to cell is what you pass: 'setups' for a cell per whole camera set-up (the contact sheet of a shot list), 'viewpoints' for a cell per named view of one subject, 'passes' for a cell per render pass, or 'duration' for a cell per sample over time. Naming a station axis ('setups' or 'viewpoints') alongside 'passes' lays out a 2D grid (a row per station, a column per pass). A collage with no axis is an error naming all of them — it does not default to time. | image |
| pass | No | Which render output to produce. 'final' = normal lit rendering (default). The built-in debug passes replace the fragment output with a diagnostic buffer — use these to verify geometry and material correctness without reasoning about final lit pixels: 'normal' (alias 'world_normal') = per-pixel world-space normal (RGB = xyz*0.5+0.5), the geometry sanity check; 'normal_texture' = raw normal map sample before TBN (flat blue if no normal map); 'depth' (alias 'linear_depth') = distance from the camera on a log curve (near black, far white), the fastest way to read scene layout; 'albedo' / 'roughness' / 'metallic' / 'ao' / 'emissive' = the corresponding PBR channel, drawn magenta on a surface whose shader lights itself and hands the engine no material to read; 'tangent' = world-space tangent; 'material_flags' = R=normal_map, G=roughness_metallic_tex, B=base_color_tex; 'shadow' = the share of the direct light reaching each fragment that survived every shadow test, over EVERY light standing on it — sun, point, spot and area alike (white = all of it, black = none of it), so a fragment the 'final' pass renders dark for being shadowed reads dark here; 'shadow_depth' = the map behind that answer for the light putting the most light on the fragment — R = the depth the fragment carries in that light's map, G = the depth the map already held there, B = which light answered (0 the sun's cascades, 0.25 a parallel light, which claims a map of no kind, 0.5 a spot or area light's atlas tile, 1 a point light's cube face); 'motion_vectors' = the frame's velocity buffer — each pixel's screen-space travel between the previous frame and now, whichever pass wrote it (a material, the splat pass, a render feature naming @scene.motion); static pixels flat black, moving pixels colour-biased. Beyond these, `pass` also accepts the name of any CONTENT-registered capture view (a render feature that publishes a view — e.g. a baked-lightmap view); an unknown name is resolved against that registry and, if it matches no view, the error lists what is available. Any pass but `final` is drawn under `all !sky !ui !EditorUI !debug` with the post-process chain off, so the buffer comes back as the value it encodes rather than as a graded picture of it; `renderLayers` and `postProcessing` each state that differently. | final |
| angle | No | Orbit direction as [yaw, pitch] in degrees (default [0, 20]). Only the viewing angle — distance is still auto-fit from bounds unless you set 'distance'. source='entity' only. | |
| basis | No | Which axes a 'viewpoint' is measured against. 'local' = the SUBJECT's own axes: 'front' is the side it faces however it is turned, and the frame is fitted to its own extents rather than the world-axis box around them. 'world' = the world axes: 'front' is whatever faces world +Z. Required alongside 'viewpoint' when framing an entity, because for anything rotated the two are different pictures. Framing a position instead of an entity, both values mean the world axes. | |
| owner | No | The key the holds this capture takes are stated under. `deterministic` pins the per-frame clock and `clearAir` holds the air, and both are cells every agent driving this engine renders through, so an agent can take one exclusively for a key of its own. A capture stating that key is admitted under the standing hold and photographs what its owner set; a capture stating another key, or none, is refused while that hold stands. It is the same key `renderer.temporal.hold` and `renderer.atmospherics.hold` take as `owner`, and their `release` hands the cell back by. | |
| width | No | Output width in pixels (image) or the whole SHEET's width (collage). A collage cell is rendered at the box the grid divides out of the sheet, so a cell holds the same picture a single capture at that cell's own width/height holds, and each cell reports the pixels it was rendered at. Omit to take the size of whatever is being captured: a capture of a camera ('main', 'screen', 'viewport', 'editor', 'camera') renders at the viewport's own pixel size, and one standing at a position, orbiting an entity or painting a UI window renders 1024 wide. | |
| camera | No | (source='camera', or with no source) A camera entity ref/id to render from. Renders that camera's authored pose offscreen - use it to validate any specific camera (a security cam, a cutscene cam) regardless of which camera is on screen. | |
| entity | No | What to frame (source='entity'): an entity name or id, OR an array of names/ids to frame several at once. The camera AUTO-FITS to the union of their world-space bounds, so the subject is guaranteed fully in frame at ANY scale: a 1cm prop and a 100m building both fill the shot. Single or list, name or id; you cannot pick the wrong shape. Framing is not isolation: this FITS THE FRAME to the subject, it does not decide what is drawn in it, and anything else standing there is still rendered, and a scene's default player spawn sits at the world ORIGIN, exactly where a prop is usually built. Pass 'isolate' to draw the subject without the other geometry. If the framed shot comes back empty, the geometry/material is the problem, not the camera. A name more than one entity carries is refused, and the refusal names each of them with its id; every cell reports the entities it framed under 'subjects'. | |
| format | No | Output image format. 'jpeg' (default) is smaller. 'png' is lossless. Ignored for source='ui_window' (always PNG). | jpeg |
| frames | No | source='presented' / 'window' only: how many presented frames to take, 1-32 (default 1). More than one comes back as a grid in the order the frames were presented, and each cell reports its frame number, so a run of consecutive frames reads as rising numbers at real gameplay speed: the measure of TAA convergence, motion blur, flicker and exposure adaptation over time. | |
| height | No | Output height in pixels (image) or the whole SHEET's height (collage). A collage cell is rendered at the box the grid divides out of the sheet, so a cell holds the same picture a single capture at that cell's own width/height holds, and each cell reports the pixels it was rendered at. Omit to take the size of whatever is being captured: a capture of a camera ('main', 'screen', 'viewport', 'editor', 'camera') renders at the viewport's own pixel size, and one standing at a position, orbiting an entity or painting a UI window renders 576 tall. | |
| lookAt | No | World point the camera aims at [x, y, z]. source='position' only; defaults to origin. | |
| passes | No | Collage pass axis: a cell per render pass, taking the same names as 'pass' (including content-registered capture views). One frame under final, albedo, normal and depth side by side is how a material problem separates from a geometry one. Combine with 'setups' or 'viewpoints' for a 2D grid. | |
| screen | No | (source='ui_window') Screen id passed to ui.registerScreen that contains the target Window. | |
| setups | No | Collage set-up axis: a cell per WHOLE camera set-up — the contact sheet of a shot list. Each entry takes the same options that aim a single capture ('source', 'camera', 'entity'/'entities', 'position', 'lookAt', 'rotation', 'distance', 'angle', 'margin', 'viewpoint', 'basis', 'projection', 'orthoHeight', 'fov', 'near', 'far', 'isolate'), plus 'label' to name the shot on the result; a field an entry leaves out is taken from the collage's own options, so a shot list writes down only what makes each shot differ. This is the axis for comparing several DIFFERENT stations and lenses of one scene at one instant — 'viewpoints' orbits one subject, 'duration' samples one camera over time. Exclusive with 'viewpoints' (both say where the camera stands); combine with 'passes' for a 2D grid, a row per set-up. | |
| source | No | Where the image comes from. 'presented' and 'window' READ the frame the engine presented; every other source RENDERS a frame of its own offscreen. 'presented' = the game view exactly as it was presented to the player, after the post chain, the upscale, the display transform and the UI, at the size it was composed at: the one source whose temporal antialiasing, motion blur, camera motion, auto exposure and render scale are the real ones, since nothing is re-rendered for it. 'window' = the whole presented window (or headless framebuffer) with the editor chrome and everything painted on the window. Both take 'frames' + 'interval' for a run of consecutive presented frames, report each frame's frame number, render scale, raster and composite sizes, projection jitter, exposure and history reset, and refuse the options that aim or shape a rendered frame (camera, entity, position, viewpoint, fov, pass, renderLayers, postProcessing). DEFAULT (omit, or 'main') = the MAIN SCENE CAMERA (the gameplay/PlayerPrototype camera or an agent-placed scene camera), rendered offscreen from its authored pose; it works in edit mode too (before you press play). Errors if the scene has no non-editor camera. 'editor' = the editor fly-camera you author with. 'screen' = a reproduction of the on-screen image, rendered offscreen from the camera holding the viewport under the render layers and diagnostic pass THAT camera draws the screen with, so a camera sitting on a diagnostic channel comes back as that diagnostic; the frame as presented is 'presented'. Passing 'renderLayers' or 'pass' alongside it states them for this capture instead. 'viewport' and 'gameplay' = aliases for the default (the scene camera). Pass 'camera' with a camera entity ref to render from any specific camera. 'entity' = orbit an ephemeral camera around an entity (entity + distance + angle). 'position' = render at an explicit world point (position + lookAt). 'ui_window' = render ONE egui Window into its own texture (screen + window). If 'source' is omitted: 'entity' when you pass 'entity', 'position' when you pass 'position', 'ui_window' when you pass both 'screen' and 'window', 'camera' when you pass 'camera', otherwise the main scene camera. | |
| window | No | (source='ui_window') Widget id of the Window inside that screen (e.g. 'system-tools-window'). | |
| isolate | No | Draw the subject WITHOUT the other geometry (source='entity' only). It changes exactly that: lighting, sky and post-process are untouched, so a final frame stays final — for a flat read of the geometry use 'pass'. Excluded geometry still casts shadows onto the subject and still bounces light into it, the same as any render-layer exclusion. The subject's layers are restored afterwards. In a 'setups' collage each cell shows its own set-up's subject alone; a set-up states 'isolate' of its own to override this for its cell. | |
| quality | No | JPEG quality 1-100 (format='jpeg' only). | |
| clearAir | No | How much of the air between the camera and a surface reaches this one frame. `true` takes it out entirely, so a surface renders in its own colour and its albedo, tint or material can be judged while another slice of a shared world drives the weather — a wash that no other option reaches, because aerial perspective and the volumetric media are render features rather than post-process or a render layer, so neither `postProcessing` nor `renderLayers` strips them. A number in [0, 1] keeps that share of the air instead. It reaches aerial perspective, height fog and volumetric light scattering wherever those are driven from, a component re-pushing them every frame included. The sky, the sun and the light they put on the surface are untouched, because those are what the surface's colour is made of, and the cloud layer and a placed volume draw as geometry that `renderLayers` and `isolate` already name. Scoped to this capture: the scene's own air is back the moment it returns, so a world several sessions share is never left in scratch weather. | |
| distance | No | Override the orbit distance (source='entity'). OMIT to auto-fit from world bounds (recommended — that's the guarantee that the subject is framed). Pass a value only to force a specific radius, which can push the subject out of frame. | |
| duration | No | Collage time axis: a cell per sample, spread across this many seconds of gameplay. Exclusive with 'setups', 'viewpoints' and 'passes' — a cell differing both in what it looks at and in when it was taken answers neither question. | |
| interval | No | source='presented' / 'window' only: how many presented frames apart two taken frames stand, 1-600 (default 1: consecutive frames). | |
| position | No | Explicit camera world position [x, y, z] (required when source='position'). | |
| save_path | No | Optional path to save the image to. Takes a VFS path in either spelling ('/source/tmp/shot.png' or '/zero/source/tmp/shot.png'), or a host-filesystem path ('/tmp/shot.png'). The image is still returned as base64 either way; the response reports 'savedTo' with the resolved destination, or 'saveError' explaining why there is no file. A VFS save is written without firing write side effects, so the file stays at the path 'savedTo' names and reads back from it: hand that path straight to 'set_world_cover', 'read_file' or 'vfs.read'. A VFS save under /zero/source made while play is running lands on the play shadow: the response then carries durable=false, a 'warning' naming the file a guarded play-exit deletes unless it is kept, and 'playShadow' naming the ways to keep it; a host-filesystem path is outside the VFS and is kept whatever play does. To make the image an imported texture asset instead, write it with 'write_file' or run the importer on it. | |
| viewpoint | No | A named camera station, instead of working out an 'angle'. front/back/left/right/top/bottom are the six sides; 'iso' is the corner view where all three axes project equally. REQUIRES 'basis' when framing an entity. Pair with projection='orthographic' to read a side as a true elevation. | |
| projection | No | 'perspective' (default) converges with distance, the way an eye sees. 'orthographic' covers a fixed world-space height at EVERY distance, so parallel edges stay parallel and two equal-size objects at different depths cover equal pixels — the projection to check a shape or compare sizes under, where perspective convergence hides both. | perspective |
| viewpoints | No | Collage view axis: a cell per named view of ONE subject. An array of viewpoint names, a space-separated string ('front top left'), or 'sides' for all six. Combine with 'passes' for a 2D grid; use 'setups' instead to vary the whole camera set-up from cell to cell. Exclusive with 'setups'. | |
| orthoHeight | No | World-space vertical extent an orthographic frame covers. Omit to fit it to whatever is being framed (recommended, the same way distance is auto-fit). | |
| renderLayers | No | Which render layers this capture includes — the control over both which objects draw and which passes run (sky/ui/EditorUI are built-in layers). Space-separated tokens: `all` seeds every layer, `name` adds a layer, `!name` drops one — e.g. `all !ui` for a clean shot with no UI overlay, `all !sky !ui` for a flat diagnostic background, or `default enemy` to render only those two layers. Post-processing is not a layer: pass `postProcessing = false` to capture the ungraded frame. Defaults omit `EditorUI` (the human editor's chrome) and `debug` (the debug-visualisation overlays — gizmos, light/probe icons, frustums, collider and bounds wireframes, several of which are on by default in edit mode) so agent captures aren't cluttered with them; pass `all` to include them. Naming a `pass` other than `final` states a stronger default than any of these: a diagnostic buffer's RGB is the reading, so that capture is drawn under `all !sky !ui !EditorUI !debug` and with the post-process chain off — the sky would paint its own colour where the buffer holds none, and the chain would grade the values being read. Stating `renderLayers` here replaces that spec, and `postProcessing = true` puts the chain back. Content `ui` is answered three different ways, by what aimed the shot. A capture of a camera the scene already holds — `main`/`viewport`/`gameplay` (the scene camera) and `editor` (the fly-cam) — keeps it, so an authored HUD is in the frame. A capture that BUILDS a camera to frame a subject — `entity` and `position` — drops it alongside the authoring layers, since a HUD drawn across the whole frame is unrelated to the subject being framed; pass `all !EditorUI !debug` to keep the HUD in a framed shot. Two sources answer from the camera they mirror rather than from either default: `screen` renders under the spec the on-screen camera draws the viewport with, verbatim, and `camera` under that named camera's own spec with the authoring overlays off, so each holds content `ui` exactly when the spec it mirrors does. `ui_window` paints one Window into a target of its own, and no render-layer spec applies to it. `presented` and `window` read the frame the engine presented, drawn under the layers the viewport camera draws with, and refuse a spec of their own. Overriding here scopes to THIS capture, unlike `debug.set`, which mutates the engine-wide state every other viewer shares. | |
| deterministic | No | Pin the clock every per-frame field is drawn against for the length of this capture. Film grain, an animated noise field and a jittered raymarch offset are redrawn each frame from that clock, so two shots of a scene in which nothing has moved otherwise differ over a large part of the frame — and a difference image cannot then answer whether anything in the shot moved. Pinned, they are redrawn identically, and every capture that asks for it pins at the same instant, so shots taken minutes apart still agree. It does not pin an accumulating temporal effect (TAA history, denoiser accumulation, auto-exposure adaptation), which settles by itself on a still scene. | |
| postProcessing | No | Whether this capture runs the post-process chain. Omitted, a capture naming a diagnostic `pass` drops the chain so its buffer is read flat, a capture from a specific camera mirrors that camera's own `postProcessing`, and every other capture keeps the chain, so a final frame stays a final frame. Pass false to read the scene ungraded: the chain comes off, and so does the film a render feature draws — the lens flare and grain, the bokeh a short focus throws, the shutter's smear — each declared part of the camera's photographic finish by the pass that draws it, so a graded frame and an ungraded one are different pictures. Post-processing is a Camera property, not a render layer, so `renderLayers` cannot strip it. |