Skip to main content
Glama

android_screenshot

Capture the current Android device screen as a base64 PNG image for visual analysis. Enables AI agents to inspect UI state or diagnose issues.

Instructions

Capture the current screen of the Android device and return it as a base64 PNG image for visual analysis.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the return encoding ('base64 PNG'), but does not state device prerequisites (e.g. connection state), whether it is a side-effect-free read, image dimensions/size, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every clause (capture, target device, output format, purpose) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description covers what the agent most needs: what is returned and in what encoding. Only minor gaps remain (device-connection precondition, error behavior).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document and the baseline is 4. The description correctly adds no spurious parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Capture'), resource ('current screen of the Android device'), and output form ('base64 PNG image'), which makes it clearly distinct from text-oriented siblings like android_get_ui_tree. It does not explicitly name a sibling, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for visual analysis' implies the intended use case, distinguishing it from tree/text inspection tools, but there is no explicit when-to-use/when-not-to-use guidance or named alternative. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.