Skip to main content
Glama

Caption a TOP (is the output alive?)

caption_top
Read-only

Render a TOP preview and return a plain-text description of its visual output. Use after a build to confirm the network is rendering instead of showing a black frame, via vision LLM or deterministic histogram fallback.

Instructions

Read-only: render a TOP's preview and return a plain-text description of it — the headless 'is the output alive?' primitive. Two paths: (a) a configured vision LLM endpoint when available, (b) a DETERMINISTIC luma/colour-histogram fallback decoded from the preview PNG pixels (always works, no model needed). Reports dominant colours, mean luma, near-black fraction, a coarse classification ('black'/'very dark'/'dark'/'bright'/'colorful'/'mid'), and a friendly caption. Returns {node_path, width, height, source:'vision'|'histogram', caption, stats{...}, warnings}. Use it after a build to confirm the network is actually rendering instead of a black frame. The vision path is currently inert (no vision field on the tool context) and falls back to the histogram.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
widthNoWidth to render the preview at before describing it. Smaller is faster.
heightNoHeight to render the preview at before describing it. Smaller is faster.
node_pathYesPath of the TOP to caption.
use_visionNoUse the configured vision LLM endpoint when available; else fall back to a deterministic histogram description.
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, openWorld, non-destructive), the description discloses crucial behaviors: the deterministic luma/color-histogram fallback that 'always works, no model needed', the current inert vision path, and the return structure including 'source: vision|histogram' and warnings. This provides rich context about reliability, fallback logic, and potential warnings, going well beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action ('Read-only: render a TOP's preview and return a plain-text description'), then efficiently explains the two paths, the outputs, the use case, and a caveat about the vision path. Every sentence contributes value without redundancy, and the structure flows logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explicitly enumerates the return fields ('{node_path, width, height, source:... caption, stats{...}, warnings}'). It explains both execution modes, the deterministic fallback, the intended usage scenario, and a current limitation. For a tool with 4 parameters and no output schema, this description is exceptionally complete and self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides thorough descriptions for all four parameters (width, height, node_path, use_vision), including defaults and the 'smaller is faster' note. The description adds context about the two paths (vision vs histogram) and the inert vision path, but most parameter-specific information is already in the schema. Since coverage is 100%, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'render a TOP's preview and return a plain-text description of it'. It also labels it as the 'headless is the output alive? primitive', and specifies the exact outputs (dominant colors, mean luma, near-black fraction, classification, caption). This is a specific verb+resource and distinguishes it from sibling tools like get_inline_preview or render_output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit use case: 'Use it after a build to confirm the network is actually rendering instead of a black frame.' It also clarifies the two execution paths and warns that the vision path is currently inert. However, it doesn't explicitly state when not to use this tool or compare it to alternative tools, so it lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/lucasmaher-hash/touch-designer-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server