Skip to main content
Glama
FZ2000

android-phone-control

by FZ2000

take_screenshot

Capture the phone screen to inspect visual content like photos, maps, or CAPTCHAs when text-based reading returns nothing useful or lacks confidence.

Instructions

Capture the phone's screen and look at it.

Try read_screen first: it names things so you can act on the result, and it costs far less. Take a screenshot when the screen is genuinely visual -- a photo, a map, a game, a CAPTCHA, an icon-only toolbar -- when read_screen returns nothing useful, or when a decision came back with low confidence.

Args: max_width: Downscale the image to this width in pixels before returning it. Lower it to spend fewer tokens; 0 keeps full resolution.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
max_widthNoDownscale the image to this width in pixels. Lower costs fewer tokens to look at.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It accurately explains that the screenshot is returned, can be downscaled, and incurs a token cost, which is meaningful operational behavior. It does not mention potential limitations like image format or screen capture permissions, but those are not critical for this tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a one-sentence purpose, a focused usage paragraph that earns its place by routing the agent correctly, and a compact args explanation. No filler or redundant content; the structure supports fast parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only tool with no annotations or output schema, this description covers what an agent needs: when to use it, how to control image size, and what the result is (an image to look at). It also names the sibling tool and the tradeoff, making it fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents `max_width` at 100% coverage, so baseline is 3. The description adds a useful semantic detail not in the schema: `0 keeps full resolution`, and reiterates the token-cost tradeoff in a more decision-oriented way. This exceeds baseline by providing practical guidance on how to choose the value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool captures the phone screen and looks at it, using a specific verb and resource. It explicitly differentiates itself from the sibling `read_screen` by naming what it is not and when the alternative is preferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives direct, actionable guidance: try `read_screen` first, use a screenshot only for genuinely visual content or when `read_screen` is unhelpful or low-confidence. This explicitly defines when to use the tool versus its sibling, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.