Skip to main content
Glama

pixel_to_world

Convert picture pixel coordinates into 3D world points on a chosen plane using camera calibration, so you can measure floor positions and sizes from a photo in Blender.

Instructions

Turn picture pixels into world points. A ray goes from the camera through each pixel and hits the plane z=plane_z (the floor by default) or plane {"point": [x,y,z], "normal": [x,y,z]}. pixels is a list of [x, y]: x to the right, y down, from the top left of a picture of image_size [width, height]. It uses the field of view, the sensor fit, the lens shift and orthographic cameras. A point is null if the ray is parallel to the plane or the plane is behind the camera. Use it to read sizes and positions off a photo after match_camera (for example the width of something that stands on the floor). Heights above the plane cannot be read from one view.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
planeNo
cameraYes
pixelsYes
plane_zNo
image_sizeYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.4.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the exact failure semantics ('A point is null if the ray is parallel to the plane or the plane is behind the camera'), the default plane (the floor, z=0), that camera intrinsics — field of view, sensor fit, lens shift, orthographic — are honored, and a hard limitation on the reading of heights. That is genuine behavioral disclosure beyond anything in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then mechanism, parameter conventions, edge cases, and usage — a logical order with little waste. It is dense and multi-clause, and folding the parameter conventions into the same run-on paragraph makes it slightly harder to scan than a structured list would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description must stand alone, and it covers mechanism, parameters, edge-case nulls, usage and limits. The one remaining gap is the exact return shape — it implies a per-pixel list of points (some null) without saying so directly, and 'camera' is never defined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it documents the pixel coordinate convention (x right, y down, from top-left), the image_size [width, height] format, the plane_z default of floor, and the plane object shape {"point": [x,y,z], "normal": [x,y,z]}. Only the required 'camera' string parameter is left unexplained — it is inferable from the match_camera reference but its accepted form (name vs. handle) is never stated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Turn picture pixels into world points') and immediately grounds it in mechanism (ray from camera through each pixel, intersecting a plane), which an agent can use to tell it apart from the superficially similar place_at_pixel. However, it never explicitly contrasts itself with that sibling, so it stops short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real context ('Use it to read sizes and positions off a photo after match_camera, e.g. the width of something that stands on the floor') plus a when-not condition ('Heights above the plane cannot be read from one view'). It names a prerequisite sibling but not the alternative it should be chosen over, so no exclusion routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.