Skip to main content
Glama

setup_camera

Create and aim a Blender scene camera so a chosen subject fills the frame exactly, with options for view angle, focal length, depth of field, target objects, and render resolution.

Instructions

Create/aim the scene camera so the subject fills the frame (exact fit to its geometry).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
dofNoDepth of field focused on the subject.
viewNothree_quarter (default hero angle), three_quarter_left, three_quarter_back, front, back, left, right, top, hero (low, heroic), high, or an [x,y,z] direction from subject to camera.three_quarter
fstopNo
marginNoFraming looseness (1.0 = tight).
targetNoObject name or list of names.
resolutionNoRender size [w,h], e.g. [1920,1080] or [1080,1350].
focal_lengthNomm: 35 wide/interiors, 50 natural, 85-100 product (less distortion).
orthographicNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies mutation of the scene camera but never says whether an existing camera is replaced or repurposed, whether framing is reversible, whether it requires a selected/visible subject, or what happens if target is omitted — all important for a state-changing setup tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence, front-loaded with the action and ending on the key behavioral constraint (exact fit to subject geometry). No filler, though the phrase 'scene camera' could be sharper about what object it acts on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Eight parameters, zero required, no annotations, and no output schema — the description should be doing much more work. It omits what the tool operates on by default, how the camera is selected/replaced, and what the agent should expect afterwards, leaving real ambiguity for a scene-mutating setup call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%: view, margin, target, resolution, focal_length, and dof are well documented in the schema, while fstop and orthographic have none. The description itself adds no parameter meaning beyond the framing intent, so it neither compensates for the two bare parameters nor improves on the rest — baseline 3 for a largely schema-documented tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (create/aim) and resource (scene camera) plus the intended outcome — the subject fills the frame with an exact fit to its geometry. That is clear enough to distinguish it from render_image or viewport_screenshot, though it does not explicitly contrast with the sibling 'look' tool, which plausibly also reorients the view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as 'look', viewport_screenshot, or manual transform_object on a camera. It never states prerequisites (e.g. whether a subject/target must exist first) or exclusions, so the agent must infer the usage context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.