Skip to main content
Glama

desktop_read_text

Read-only

Extract on-screen text from desktop UI elements using accessibility metadata, with configurable character limits and whitespace preservation. Enables AI agents to read visible text for automation.

Instructions

Read accessible text and representation metadata, preserving whitespace. limit counts Unicode code points (default 16000, maximum 1000000), not bytes. Opaque embedded objects are not exact logical plain text: check plain_text_verification_supported. Normalization reads the bounded full field; a smaller limit does not enable streaming. Protected fields are refused. Generic native readback reports composition known=false, active=null; pending preedit is not checked.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
element_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.3.4
    • addedInput schema / additionalProperties
      Added value: +false
  2. First observedv0.1.0

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral details: preservation of whitespace, code-point-based limits, the non-exactness of opaque embedded objects, normalization reading the full bounded field, refusal of protected fields, and specific native readback behavior (composition known=false, active=null, pending preedit not checked). This far exceeds what annotations provide and gives the agent critical edge-case knowledge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and technical, but every sentence adds necessary nuance. It is front-loaded with the primary action and then layers on important caveats. While not as brief as a one-liner, the length is justified by the complexity of the behavior, and the structure flows logically from purpose to edge cases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the absence of an output schema, the description covers a wide range of scenarios: limit semantics, streaming behavior, normalization, opaque objects, protected fields, and native readback details. This is comprehensive enough for an agent to understand what to expect and avoid common pitfalls when invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It thoroughly explains 'limit' (default 16000, max 1000000, counts code points, not bytes) and its behavioral implications. 'element_id' is not elaborated, but its meaning is obvious from the name and required status. The description compensates well for the lack of schema documentation on the key parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb-resource pairing: 'Read accessible text and representation metadata, preserving whitespace.' This is specific and distinguishes the tool from OCR and other text-reading tools by focusing on accessibility metadata and whitespace preservation. It is not a tautology and immediately tells the agent what the tool accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides several usage constraints (e.g., 'limit counts Unicode code points', 'a smaller limit does not enable streaming', 'check plain_text_verification_supported') but does not explicitly state when to use this tool over siblings like desktop_ocr or desktop_inspect. The guidance is implicit rather than explicit, with no mention of alternatives or exclusions, so it falls short of a clear decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.