Skip to main content
Glama

vision.vlm

Ask a local OpenAI-compatible vision model about a screenshot or cropped UI region to identify unknown icons and describe their purpose.

Instructions

本地视觉小模型问图(crop-and-ask):缺省整屏,传 box 只裁剪该区域提问(推荐:小图快且准)。需本地 OpenAI 兼容视觉端点:设 MAH_VLM_URL 与 MAH_VLM_MODEL(如 ollama:MAH_VLM_URL=http://127.0.0.1:11434/v1/chat/completions,MAH_VLM_MODEL=qwen2.5vl:3b)。定位:图集外未知图标的语义兜底层

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
boxNo[l,t,r,b] 只裁剪该区域
fromNo截图路径,缺省最近一次 vision.screenshot
promptNo问题;缺省=描述该界面元素及用途
timeoutNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a real prerequisite beyond structured data: a local OpenAI-compatible vision endpoint configured via MAH_VLM_URL/MAH_VLM_MODEL, with a concrete ollama example. That is valuable. But it omits return format, latency/failure behavior, and what happens if the endpoint is unconfigured, so a 3 fits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core crop-and-ask behavior, then setup requirements, then positioning. The env-var example is verbose but genuinely useful for invocation. Overall appropriately sized with little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must cover a lot, and it does well on setup and defaults. The main gap is the return-value shape (what the model returns: free text? structured?), which an agent calling a Q&A tool would benefit from knowing. Adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%. The description adds meaning beyond the schema: box semantics are reinforced ('only crop that region', 'recommended: small images fast and accurate') and the default-scope behavior (full screen when box absent) is explicit, which the schema only hints at. The from/prompt/timeout params are largely covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific capability: a local VLM 'crop-and-ask' that takes a screenshot (default full screen) or a cropped box and asks a vision model about it. It also positions itself as the 'semantic fallback layer for unknown icons outside the gallery', which helps distinguish it from template/icon-matching siblings. It does not explicitly name the nearest siblings (vision.ask, vision.ocr), so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives useful context: default is full screen, pass a box to crop that region, and cropping is recommended for speed/accuracy. The 'positioning' line implies it is a fallback when gallery/icon matching fails. However it never explicitly says when to choose this over vision.ask, vision.ocr, or ui2.scene_match, so usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.