Skip to main content
Glama

overlay_reference

Draw a 3D model and reference photo on top of each other through a scene camera to compare alignment, with modes for blend, edges, difference, and split views.

Instructions

Draw the model and a reference picture on top of each other, seen through a scene camera: it shows how far the camera and the model are from the photo. The model is rendered from camera (the scene camera when empty; without one call set_camera or match_camera first) at the proportions of the reference; the long side is size pixels. names is the model (everything visible when empty). Modes: blend (the reference at alpha); edges (red reference outline, yellow reference inner edges, green model outline over the dimmed model: best for small shifts); difference (black is a match); split (reference left of a line at the split fraction of the width, model right: lines must continue across it). The answer has the path, the silhouette IoU and the picture. The reference outline comes from alpha or the corner colour; threshold is the colour distance. Give the original photo, not a crop. grid_height_m (the real height of the reference silhouette) draws a metric grid with millimetre labels; zero is the bottom centre of the silhouette box. It is true for a flat-on (orthographic) reference only. Before a model exists, read coordinates with measure_profile grid=true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
outNo
modeNoblend
sizeNo
alphaNo
namesNo
splitNo
cameraNo
referenceYes
thresholdNo
grid_height_mNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.4.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it describes each mode's visual behavior, what the answer returns (path, silhouette IoU, picture), how the outline and threshold are derived, the orthographic-only caveat for grid_height_m, and the 'give the original photo, not a crop' constraint. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action before enumerating modes and parameters, and nearly every sentence adds information. Some parenthetical clauses and run-on sentences ('green model outline over the dimmed model: best for small shifts') make it dense and slightly hard to parse, but there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with no annotations and no output schema, the description covers modes, prerequisites, return contents, and caveats well. The main gap is the unexplained `out` parameter; otherwise an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, so the description must compensate — and it explains mode, camera, size, names, alpha, split, threshold, grid_height_m, and the reference input. It omits any explanation of `out`, leaving one parameter undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — drawing the model and reference picture overlaid through a scene camera — and even explains what the overlay reveals (camera/model distance from the photo). It is clearly distinguishable from rendering siblings, though it never names the nearest alternatives (compare_view, fit_to_reference) explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives prerequisites (call set_camera or match_camera first when no scene camera exists) and routes the agent elsewhere for a related need ('Before a model exists, read coordinates with measure_profile grid=true'). Mode selection guidance is included ('edges ... best for small shifts'), but there is no explicit when-not-to-use statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.