Skip to main content
Glama

page_vision

Detect word bounding boxes in any element or viewport using local ML text detection. Returns pixel coordinates for captchas, canvas text, and scanned layouts.

Instructions

TIER 4 VISION — VISUAL CORTEX: detect word bounding boxes in an element (or viewport) via the built-in text-detection ML model. Returns JSON [{x, y, w, h}, ...] in image pixels. The first shipped visual-cortex model — see where text lives in ANY image (click-order captchas, canvas text, scanned layout). 100% local.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
refNoRef of the element to analyze. Omit for the whole viewport.
page_idYes
session_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.7.3

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries some burden. It states '100% local' (privacy), the input requirement (ref or viewport), and output format in image pixels. However, it does not disclose potential performance characteristics, model limitations, or whether the operation is read-only. The absence of annotations is mitigated somewhat by the '100% local' note, but a clear read-only hint is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is fairly concise and front-loaded: the main purpose, output, and usage are in the first two lines. The subsequent sentences provide context on model capabilities and use cases. It's not overly verbose and every sentence adds value, though the 'TIER 4 VISION — VISUAL CORTEX' could be seen as redundant marketing fluff, but it helps set context for a specialized tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (vision model) and no output schema or annotations, the description covers the essential aspects: what it does, what it returns, how to invoke (ref or viewport), and key use cases. It doesn't fully explain input format details (e.g., how to specify a specific element vs viewport other than 'ref'), but it's adequate for an agent to understand and call the tool correctly. The example use cases aid in routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with 'ref' having a description; page_id and session_id are undocumented. The description explains that 'ref' can be omitted for viewport, which adds meaning. For page_id and session_id, the description doesn't add anything—they are common identifiers, but given the low coverage, some compensation is expected. It's borderline acceptable because the main behavioral parameter (ref) is explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool detects word bounding boxes via ML model and returns JSON array. It names the resource (element or viewport) and the output format in image pixels. The 'visual cortex' metaphor and use cases (captchas, canvas text) make it clear what it is for. It's clearly distinct from OCR tools in the sibling list since it focuses on bounding boxes, not text extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions 'see where text lives in ANY image' and gives example use cases (click-order captchas, canvas text, scanned layout), implying when to use it. However, it does not explicitly state alternatives or when NOT to use it. For example, it doesn't say 'use page_ocr for text content, not bounding boxes.' It's implied but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.