Skip to main content
Glama

Look at Studio

look
Read-only

Ask a vision model about the Roblox Studio window and receive a concise text answer—no image enters your context. Use it for UI text, rendering, layout, glitches, and playtest visuals.

Instructions

Look at the Roblox Studio window through a vision model and get a short TEXT answer instead of an image — the screenshot never enters your context (a 768 px frame is ≈ 500 image tokens on the sidecar, a few hundred characters back to you). One-shot: {question, max_width? (768), region? {x,y,w,h} window px, model?} → {answer, model, provider, captured_ms, model_ms, usage, frame_path}. Ask concrete visual questions: 'is there a red error in the Output panel? quote it', 'where is the Play button (region or px)?', 'is the character standing on the platform or falling?'. The model sees only the screenshot and answers 'not visible' when it cannot tell. Watch: {watch: {question, interval_s (5, min 2; 15 on claude-cli), max_frames (60), stop_when? (substring, or /regex/i), diff_only? (true)}} → {watch_id}. Frames are analysed in the background, one model call at a time; each answer arrives as a Monitor event {type:'vision', watch_id, frame, answer, changed, provider}. With diff_only, frames whose bytes differ < 2% from the last analysed frame are skipped (no model call, no event). The watch ends on max_frames, when the answer matches stop_when, on {stop: watch_id | 'all'}, or after 3 consecutive failures; the last event has done:true and reason. {list: true} shows running watches with counts and the last answer. Prefer observe tree|props|find|player for state — they are exact and free. look is for what only pixels can tell: rendering, layout, UI text, visual glitches, what a playtest looks like. Works with an Anthropic API key (ANTHROPIC_API_KEY / ant auth login, ~2 s per look) OR a logged-in Claude Code install (claude on PATH, ~10–15 s per look on the subscription); STUDIO_LIVE_VISION_PROVIDER = auto (default) | api | claude-cli. Models: STUDIO_LIVE_VISION_MODEL (look; default claude-opus-5 / sonnet on the CLI) and STUDIO_LIVE_WATCH_MODEL (watch; default claude-sonnet-5 / haiku).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
listNoList running watches
stopNoStop a watch by watch_id, or 'all'
modelNoModel override: an id (default claude-opus-5 for look, claude-sonnet-5 for watch) or, on the claude-cli provider, an alias (sonnet / haiku / opus)
watchNoStart a background watch: frames → text events
regionNoCrop in window pixels before scaling (coordinates as in observe screenshot / windows)
questionNoOne-shot: what to look for in the Studio window
max_widthNoFrame width in px (default 768; image tokens ≈ w×h/750)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=false, and openWorldHint=true, covering the safety and side-effect profile. The description adds valuable behavioral context: the screenshot never enters the user context, the tool returns a text answer, it can return 'not visible' when inconclusive, and it details watch termination conditions and failure handling. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but information-dense, covering modes, parameters, providers, and examples. It front-loads the core purpose and the key benefit (text answer, no image tokens) before diving into details. Some redundancy, such as repeating model defaults, but each section serves a purpose. Not excessively verbose for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (one-shot, watch mode, region cropping, model overrides, provider differences), the description provides comprehensive coverage: it explains the output fields, event structure, diff_only behavior, stop conditions, and failure handling. It also references sibling tools for alternatives and lists concrete example questions. The schema and annotations cover the rest, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% as all parameters are described in the schema. The description adds some context beyond schema, such as default values for interval_s (5, min 2; 15 on claude-cli) and max_frames (60), and notes diff_only default true. However, this is mostly redundant with the schema descriptions and adds limited new meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool performs visual inspection of the Roblox Studio window via a vision model, returning text answers. It specifies the verb 'Look at' and resource 'the Roblox Studio window', and distinguishes from siblings by stating it is for pixel-level information, while exact state tools like observe are preferred for other queries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool versus alternatives: 'Prefer observe tree|props|find|player for state — they are exact and free. look is for what only pixels can tell.' It also explains when to use one-shot versus watch mode, and details the difference between providers and models.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools