mobile-mcp-opengl
用于OpenGL Android开发和自动化的MCP
一个用于AI编码代理(Claude Code、Cursor等)的MCP服务器,用于测试整个UI绘制在单个不透明OpenGL/Vulkan/Metal表面内的Android应用——Cocos2d-x、Unity、Unreal、原始OpenGL、libGDX及类似引擎。
它解决的问题
adb shell uiautomator dump以及所有基于无障碍树的自动化工具(包括大多数MCP移动自动化服务器)通过检查原生Android视图层次结构来工作——按钮、标签、它们的文本和坐标。这对于由原生视图构建的正常Android UI效果很好。
但对于将整个UI作为纹理渲染在单个GLSurfaceView内的游戏或应用,它不起作用。从无障碍树的角度来看,屏幕上只有一个不透明视图,没有子视图、没有标签、没有任何内部元素的坐标。没有任何可检查的内容——无论屏幕上实际有多少UI,它都是一个黑盒。
剩下的唯一真实观察渠道是截图。本服务器正是围绕这一事实构建的,将其视为常规情况,而非偶尔的备用方案。
与mobile-mcp的区别
mobile-next/mobile-mcp是通用的MCP移动自动化服务器,对于普通原生应用来说是一个很好的默认选择:优先使用无障碍树(快速、廉价、无需视觉模型、无图像令牌),仅在树无法提供所需信息时才回退到截图+坐标。
对于OpenGL画布应用,这种回退并非偶尔——它是唯一始终有效的路径。mobile-mcp-opengl正是针对这种情况构建的,因此做出了两个不同的设计选择:
完全不尝试无障碍树。 尝试毫无意义——对于这些应用,它总是返回空——因此这里的每个工具都直接使用截图+视觉。
视觉分析通过可插拔的独立提供程序进行(见下文),而不是通过运行调用代理的模型。对游戏进行功能QA循环很容易在每次会话中达到数百次截图检查;将所有这些都通过主编码代理自身的视觉进行路由,既花费真金白银,也消耗你更愿意花在实际编码工作上的令牌/上下文。在这里,截图字节完全不会进入调用代理的上下文——只有提供程序的简短文本答案会。
Related MCP server: Android-MCP
为什么使用组合的动作+观察工具,而不是单独的原语
天真的设计将tap、screenshot和ask暴露为三个独立的工具。这迫使调用代理为每次交互编排多步循环:点击→截图→交给视觉步骤→读取结果→决定下一步。每一步都是单独的工具调用和单独的轮次——在协调上浪费令牌,而不是在实际测试逻辑上,并且给代理更多机会来遗漏步骤、错误排序或在调用之间基于过时状态进行推理。
相反,本服务器暴露组合工具——tap_and_ask、swipe_and_ask、long_press_and_ask——它们执行操作、短暂等待、截图、询问视觉提供程序,并返回一个简短答案,全部作为一次工具调用。多步测试场景最终每个有意义的检查大约花费一个代理轮次,而不是三四个。
普通的screenshot_ask(仅观察,无操作)和廉价的非视觉工具(type_text、press_key、logcat_grep)也可用于测试流程中不需要此模式的部分。
工具
工具 | 功能 | 视觉调用? |
| 截图,然后询问关于它的简短问题 | 是 |
| 点击(x, y),等待,截图,询问 | 是 |
| 滑动/拖动(x1,y1)→(x2,y2),等待,截图,询问 | 是 |
| 长按(x, y)持续一段时间,等待,截图,询问 | 是 |
| 可选操作,然后按时间间隔拍摄N张截图,对每帧询问相同问题 | 是(N次调用) |
| 在当前聚焦的字段中输入 | 否 |
| 发送Android | 否 |
| 读取最近的logcat,可选地按正则表达式过滤 | 否 |
| 报告今日累计视觉支出和阈值 | 否 |
只要所需信息已经在日志行中(崩溃、你自己的调试输出、网络错误),就优先使用logcat_grep而不是视觉调用——它免费且精确,而视觉调用两者都不是。
检查动画:record_and_ask
单帧工具无法告诉你某物是否正确动画(力量指示器是否平滑脉动,标签是否飞起并淡出,精灵是否弹回起始位置)。record_and_ask执行一个可选操作(点击或滑动,或两者都不),等待waitMs(与tap_and_ask/swipe_and_ask中的waitMs含义相同——在首帧之前让UI开始反应的时间),然后按intervalMs间隔捕获frameCount张截图,并返回每帧一个简短答案——调用代理在一次工具调用中获得时间线,而不是自己编排N次单独的截图+询问往返。
为什么每帧一次视觉调用,而不是将所有帧捆绑在一次调用中。 Runware的imageCaption被证明接受一个未记录的inputImages数组(复数),与文档化的单个inputImage并列——已直接针对API测试。对于恰好2张图像,它工作正常(同一请求中的前后对比返回正确且连贯)。在3张及以上图像的一次请求中,该数组参数和手动合成的并排“胶片条”图像在测试中都产生了截断或格式错误的答案——这个小的7B视觉模型显然在单次调用中超过一定的组合视觉+指令负载后会失去连贯性。顺序单图像调用(本工具的方法)在测试的任何帧数下都可靠,并且不会显著更昂贵:成本主要由响应长度(见下文)而非调用次数决定,因此N个简短的顺序答案与一个长的多图像答案成本大致相同或更低。如果你自己的提供程序能更可靠地处理多图像请求,这是一个明显的优化点——参见“自带模型”部分。
设置
git clone <this repo>
cd mobile-mcp-opengl
npm install
cp .env.example .env
# edit .env: at minimum set RUNWARE_API_KEY (or switch VISION_PROVIDER, see below)需要adb在PATH中(或在.env中设置ADB_PATH),以及一个运行中/已连接的设备或模拟器。如果连接了多个,请设置ADB_DEVICE_SERIAL(参见adb devices)。
注册到Claude Code
在项目根目录添加.mcp.json(此文件通常是项目本地的且被git忽略,因为它通常指向特定机器的路径或包含特定机器的环境覆盖):
{
"mcpServers": {
"mobile-opengl": {
"command": "node",
"args": ["/absolute/path/to/mobile-mcp-opengl/src/server.js"]
}
}
}Claude Code会自动为项目获取此文件。服务器读取自己的.env(位于此仓库中package.json旁边)以获取所有配置——调用代理本身永远不需要知道或传递任何API密钥。
成本模型——在运行长时间QA会话之前请阅读
成本由响应长度驱动,而非图像大小。 这是针对默认的Runware/Qwen2.5-VL-7B-Instruct提供程序进行实证测量的:相同的问题,强制一个单词的答案,在从360×360到1600×2400(视网膜级)的图像大小下成本相同($0.0006)。同一张1024×1024图像,使用开放式“描述这个”提示,成本为$0.0013–0.0019——高出2-3倍——纯粹是因为模型写了更长的答案,而不是因为图像更大。
实际影响:
发送前不要费心对截图进行降采样——对于此提供程序,它不会显著降低成本,而且你会丢失可能需要的细节。
始终将问题表述为强制简短答案:是/否、数字、简短标签、包含几个字段的小型JSON对象。本服务器中的每个工具都会自动附加简短答案指令,但模糊的开放式问题(“你看到了什么?”)仍可能促使模型给出比具体问题(“错误对话框可见吗?是/否”)更长的答案。
对于格式良好的简短问题,每次调用约$0.0006,500次调用的QA会话大约花费$0.30。同样数量的开放式“描述屏幕”问题可能花费2-3倍。
内置支出护栏
每次视觉调用都会记录到.vision-log.jsonl(JSONL,每次调用一条记录:时间戳、问题、答案、成本)。该日志之上有两层独立的保护,两者都与提供程序无关(它们基于提供程序报告的任何costUsd工作):
每次调用警报(
VISION_ALERT_USD,默认$0.0015):如果单次调用返回超过此值,工具响应将包含[COST ALERT]注释,告诉你模型可能忽略了简短答案指令——这是重新表述问题的信号,而不是默默接受的事情。每日上限(
VISION_SESSION_CAP_USD,默认$2.00):一旦今日累计记录支出达到此值,每次进一步的视觉调用都会被直接拒绝(在到达提供程序之前),直到上限提高或日期翻转。这是对失控循环的硬性停止,而不仅仅是警告。
随时调用vision_spend_report检查今日总额,无需进行设备或视觉调用。
如果提供程序无法报告成本(参见下面的openai-compatible),来自它的调用将以costUsd: null记录,并且永远不会触发警报或计入上限——护栏无法保护它们无法看到的支出。
自带模型
视觉分析通过src/providers/visionProvider.js进行,它从.env中的VISION_PROVIDER按名称选择提供程序。内置两个:
runware(默认)——直接与Runware.ai的imageCaption任务通信,默认使用Qwen2.5-VL-7B-Instruct(AIR idrunware:152@2)。Runware和OpenRouter是两个独立服务,具有独立的API密钥和模型目录——这直接与Runware通信,而不是通过OpenRouter。openai-compatible— 一个通用提供程序,适用于任何使用OpenAI聊天补全视觉格式(image_url内容部分)的服务。适用于OpenRouter、运行视觉模型的本地Ollama/LM Studio服务器、Groq、Together.ai或任何其他兼容端点。在.env中配置OPENAI_COMPATIBLE_BASE_URL、OPENAI_COMPATIBLE_API_KEY、OPENAI_COMPATIBLE_MODEL。大多数OpenAI兼容API报告令牌使用量而不是固定美元成本;如果你希望此提供程序从中估算costUsd,请设置OPENAI_COMPATIBLE_PRICE_PER_1M_INPUT/_OUTPUT(否则,根据上面的说明,此提供程序的成本跟踪/护栏将无效)。
要添加完全自定义的提供程序(自托管模型、完全不同的API形状),请复制src/providers/openaiCompatibleProvider.js作为起点,实现:
async function ask(imageBuffer, mimeType, question) {
// return { text: string, costUsd: number | null }
}
module.exports = { ask };并在src/providers/visionProvider.js的loadProvider()中按名称注册它。
许可证
MIT
由Kinect.PRO开发
Available Tools
9 toolslogcat_grepRead recent logcat, filteredA
Read the last N logcat lines, optionally filtered by a regex (e.g. your app's tag, or "Exception|FATAL"). No vision call, no cost - prefer this over screenshot_ask whenever what you need is already in a log line (crashes, your own debug prints, network errors).
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | How many recent lines to fetch (default 200). | |
| filterRegex | No | Optional regex; only matching lines are returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly frames the operation as a read ('Read the last N logcat lines'), implying no mutation, and adds resource-related behavior ('No vision call, no cost'). It does not detail empty-result behavior or regex error handling, but for a non-destructive log reader this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, each earning its place: the first states the core operation, the second provides selection guidance and cost context. No redundant phrases or unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only two optional parameters and no output schema, the description covers the action, filtering, selection criteria, and cost trade-off. It is complete enough for an agent to invoke correctly, though it does not spell out behavior for empty results or invalid regex.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents both parameters with 100% coverage, so the baseline is 3. The description adds practical regex examples ('your app's tag, Exception|FATAL') and clarifies that the filter is optional, providing contextual guidance beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a clear resource ('logcat lines'), and it explicitly distinguishes itself from a sibling tool (screenshot_ask) by stating 'No vision call, no cost'. An agent can immediately tell what this tool does and how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use rule: 'prefer this over screenshot_ask whenever what you need is already in a log line,' followed by concrete examples (crashes, debug prints, network errors). It also explains the cost advantage, making the selection decision clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
long_press_and_askLong-press + screenshot + askA
Long-press at (x, y) for durationMs, wait briefly, take a screenshot, and ask a short question about the result.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| waitMs | No | Milliseconds to wait after releasing before screenshotting (default 500). | |
| question | Yes | ||
| durationMs | No | Hold duration in ms (default 800). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. It transparently lists the operation sequence and references default waitMs/durationMs defaults in the schema. However, it does not disclose side effects of long-pressing (e.g., opening context menus or triggering navigation), what the 'ask' returns or to whom, or any required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core action and includes the full workflow without filler. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no annotations, and no output schema, the description leaves important gaps: the return/response behavior is ambiguous ('ask a short question about the result'), coordinate system is unspecified, and side effects are not mentioned. An agent would need additional implicit knowledge to call this tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions only cover waitMs and durationMs (40% coverage). The description helps by framing x and y as long-press coordinates and question as a short question about the result. Still, it does not specify coordinate units/origin or any constraints on the question, so it only partially compensates for the schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action sequence: long-press at (x, y) for durationMs, wait, screenshot, and ask a question. The long-press gesture clearly differentiates it from sibling tools like tap_and_ask, swipe_and_ask, and screenshot_ask, even without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: perform a long-press and inspect the resulting screen via a screenshot and question. However, it does not explicitly state when to prefer this over tap_and_ask, swipe_and_ask, or other siblings, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyPress hardware/virtual keyA
Send an Android keyevent code (e.g. 4 = BACK, 66 = ENTER, 187 = APP_SWITCH). No vision call.
| Name | Required | Description | Default |
|---|---|---|---|
| keycode | Yes | Android KEYCODE_* integer value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description is the only source of behavior. It discloses the core action (sending a keycode) and that it does not use vision, but it does not clarify whether the key is pressed and released with a single event or describe timing/duration. Lacks details on side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with two clauses; the key action is front-loaded, and each part (action, examples, vision exclusion) adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple tool with one required parameter and no output schema. The description covers what it does and gives examples, so an agent can invoke it correctly. It does not explain return behavior or errors, but those are likely unnecessary for a fire-and-forget key event. Minor missing context about when to use it is covered under usage guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents keycode as an Android KEYCODE_* integer. The description adds specific example values (4=BACK, 66=ENTER, 187=APP_SWITCH), which clarify the range and meaning significantly beyond the schema's generic description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb 'send' and resource 'Android keyevent code', gives concrete examples distinguishing it from vision-based siblings like screenshot_ask and tap_and_ask, and explicitly notes 'No vision call.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides minimal guidance on when to use; the 'No vision call' implies it is not for visual tasks, but it does not explicitly name alternatives or conditions for selection. The examples imply use for system keys but lack explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_and_askRecord a timed screenshot sequence + ask about each frameA
For checking an ANIMATION or any effect that plays out over time (e.g. "does the strength indicator pulse smoothly?", "does the XP label fly up and fade out?", "does the sprite return to its start position?"). Optionally performs one action first (tap or swipe, or neither), then takes frameCount screenshots spaced intervalMs apart, and asks the SAME short question about each frame separately (each frame gets its own vision call, with its frame number in the prompt) - returns one answer per frame in order.
Sequential single-frame calls were chosen over sending several frames in one request: Runware's imageCaption does accept an undocumented multi-image array, and it works fine for exactly 2 frames, but degrades noticeably at 3+ (truncated/malformed answers in testing) - sequential calls are both more reliable and, per-frame, no more expensive. Keep frameCount modest (3-6) - each frame is a full separate vision call and cost scales linearly with it.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Required for action=tap or action=swipe (swipe start x). | |
| y | No | Required for action=tap or action=swipe (swipe start y). | |
| x2 | No | Required for action=swipe (end x). | |
| y2 | No | Required for action=swipe (end y). | |
| action | Yes | Action to perform before starting the capture sequence. | |
| waitMs | No | Milliseconds to wait after the action before the FIRST screenshot (default 500) - same meaning as waitMs in tap_and_ask/swipe_and_ask, separate from intervalMs which spaces out the frames after that. | |
| question | Yes | The same short question asked about every captured frame (e.g. "Is the indicator visible? yes/no"). | |
| frameCount | Yes | How many screenshots to take, spaced intervalMs apart (2-8; keep modest, see description). | |
| intervalMs | Yes | Milliseconds between each screenshot (i.e. the sampling interval of the sequence). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers: it reveals that each frame gets its own vision call with the frame number in the prompt, that answers come back one per frame in order, and that sequential calls were deliberately chosen over multi-image requests due to reliability degradation. This gives the agent accurate expectations about cost and behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence contributes: purpose, action semantics, per-frame behavior, return ordering, rationale for sequential calls, and cost guidance. The key use case is front-loaded, and the engineering rationale is placed where it helps the agent decide rather than adding noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, no output schema, and no annotations, the description covers the core interaction fully: what triggers the sequence, what each frame does, how the question is applied, what the result order is, and cost implications. The remaining parameter details are already well documented in the input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra value by advising frameCount stay modest (3-6), explaining linear cost scaling, and clarifying that waitMs is distinct from intervalMs and has the same meaning as in tap_and_ask/swipe_and_ask. This goes beyond the schema's structural descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific use case—checking an animation or effect that plays out over time—and clearly states the mechanism: perform an optional action, take frameCount screenshots spaced intervalMs apart, and ask the same question about each frame. This distinguishes it from single-shot siblings like screenshot_ask or tap_and_ask without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool: for animations or time-based effects. It also explains when the optional action is tap, swipe, or neither. However, it does not explicitly name alternative tools or state when NOT to use this one, so the guidance is clear but not fully contrastive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_askScreenshot + askA
Take a screenshot of the current screen and ask a short question about it (e.g. "Is there an error dialog visible?", "How many word icons are on screen?", "What color is the strength indicator?"). Use this when you need to check state WITHOUT performing an action first. Phrase the question so a short answer is possible (yes/no, a number, a short label) - see this server's README "Cost model".
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | A short, specific question about the current screen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly communicates a non-action read-only behavior and implies the response is short ('so a short answer is possible'). It points to the README for cost details, adding context. While it doesn't describe the exact return format, the answer is implied by the question-asking purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (about 3 sentences) and front-loaded with the main action. Every sentence earns its place: statement of action, examples, usage guidance, and cost reference. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool without an output schema, the description is largely complete. It covers purpose, usage timing, question phrasing, and cost considerations. It could explicitly mention that the result is an answer to the question, but that is reasonably implied. The description is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter description already clarifies the question. The tool description adds value by providing examples of valid questions and guidance on phrasing for short answers, which enriches the parameter's semantics beyond the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair ('Take a screenshot of the current screen and ask a short question about it') and provides clear examples. It also distinguishes itself from siblings by specifying 'WITHOUT performing an action first', which separates it from action-based tools like tap_and_ask or swipe_and_ask.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool ('when you need to check state WITHOUT performing an action first') and gives concrete guidance on phrasing questions for short answers. The reference to the README 'Cost model' provides additional usage context. This fully addresses when to use instead of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
swipe_and_askSwipe/drag + screenshot + askA
Swipe (or drag, for drag-and-drop UIs) from (x1, y1) to (x2, y2), wait briefly, take a screenshot, and ask a short question about the result - all in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| x1 | Yes | ||
| x2 | Yes | ||
| y1 | Yes | ||
| y2 | Yes | ||
| waitMs | No | Milliseconds to wait after the swipe before screenshotting (default 500). | |
| question | Yes | A short, specific question about the screen after the swipe. | |
| durationMs | No | Swipe duration in ms (default 300; use longer for drag-and-drop hold gestures). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does disclose the full ordered behavior: swipe, wait, screenshot, ask. It also notes the duration nuance for drag-and-drop holds. It doesn't detail coordinate units or what 'ask' returns, but the step sequence is clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence captures the entire workflow with no filler. The core action is front-loaded, and the drag-and-drop nuance is efficiently folded into the gesture description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple composite gesture tool, but given no output schema and no annotations, it omits the return/answer semantics of 'ask', the coordinate system, and any cost or side-effect implications. These gaps prevent it from being fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate. It explains x1/y1/x2/y2 as the swipe's start and end points, and clarifies that question should be short and about the post-swipe screen. The optional waitMs and durationMs already have schema descriptions, and the prose adds the 'hold for drag-and-drop' nuance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: swipe/drag from one coordinate to another, wait, screenshot, and ask a question. It clearly distinguishes itself from sibling tools like tap_and_ask and screenshot_ask by naming the gesture ('Swipe') and the compound workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction to use drag 'for drag-and-drop UIs' gives some contextual guidance, but there is no explicit when-to-use/when-not-to-use statement or reference to alternatives. The appropriate context is implied by the swipe gesture rather than directly contrasted with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tap_and_askTap + screenshot + askA
Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call. Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in device pixels. | |
| y | Yes | Y coordinate in device pixels. | |
| waitMs | No | Milliseconds to wait after the tap before screenshotting (default 500). | |
| question | Yes | A short, specific question about the screen after the tap. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes the sequence of actions (tap, wait, screenshot, ask) and mentions a wait period before screenshotting. However, it does not disclose what happens if the tap fails, the exact return format (e.g., does it return an image, a text answer, or both?), or any side effects like requiring a running app. The description is adequate for basic behavior but lacks depth on error handling or output, which is significant for a composite tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy. The first sentence front-loads the action sequence, and the second sentence provides direct usage guidance. Every word serves a purpose, and the structure is clean and immediately understandable. It avoids jargon and is well-scoped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a composite tool with 4 parameters and no output schema, the description should cover the response format and any prerequisites. While it clearly states the purpose and usage, it does not mention what the tool returns (e.g., an answer to the question, a screenshot reference) or any necessary preconditions (e.g., the device being interactive). Since there is no output schema, the description's silence on return values leaves an agent without complete information for correctly interpreting the tool's outcome. This is a notable gap, so a 3.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal semantic value beyond the schema: it refers to 'wait briefly' which maps to waitMs, and characterizes the question as 'short and specific', but does not explain coordinate units or default wait behavior beyond what the schema provides. The description does not go beyond the schema definitions, so a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's verb and resource: 'Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call.' It distinguishes itself from the siblings by defining its specific action (tap) and explicitly contrasting with calling separate tap/screenshot/ask tools. The title also reinforces the composite nature, so there is no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: 'Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.' This clearly defines the usage context and names an alternative (separate tools). It does not explicitly mention sibling tools like swipe_and_ask or long_press_and_ask, but the reference to 'tap' inherently implies a distinction from those. This is strong guidance, but not exhaustive about exclusions from all siblings, hence a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textType textA
Type text into whatever field currently has focus. No screenshot/vision call - pair with screenshot_ask if you need to confirm the result.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that typing targets the focused field and that screenshot/vision is not part of the operation, but it does not cover edge cases such as no focused field, whether existing text is replaced, or how special characters are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the core behavior is front-loaded, and the follow-up guidance about screenshot_ask earns its place. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool, the description is nearly sufficient: it states the target, the action, and the verification route. Missing failure-mode detail, such as what happens when no field has focus, keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not add meaning beyond the property name 'text'. The parameter is simple, but nothing explains format, newline behavior, limits, or encoding, so the description fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action ('Type text') and a specific target ('whatever field currently has focus'), and explicitly warns against treating it as a screenshot/vision operation. This clearly differentiates it from screenshot_ask and the other _and_ask siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational guidance: use it when a field has focus, and pair it with screenshot_ask when confirmation is needed. It does not explicitly contrast with press_key or other input tools, but the focus-based behavior is enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vision_spend_reportReport today's vision spendA
Report the cumulative vision-provider spend for today and the configured alert/cap thresholds, without making any device or vision call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It openly states 'without making any device or vision call', which discloses side-effect-free behavior, but it does not describe the output format, potential delays, or any other behavioral aspects. This is a reasonable disclosure but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the verb and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description covers the essential information: what is reported and the guarantee of no side effects. It does not specify the return format or any prerequisites, but given the simplicity, it is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description adds meaning about what the report contains (spend and thresholds) beyond the empty schema, which is exactly what is needed for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Report' and identifies the exact resource ('cumulative vision-provider spend for today') plus the alert/cap thresholds. It is clearly distinct from the sibling action-oriented tools (screenshot, tap, etc.) by stating it makes no device or vision call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for checking spend information and explicitly notes it does not make any device or vision call, but it does not name alternative tools or provide explicit when-to-use guidance. The use case is somewhat obvious given the sibling list, but the guidance is not explicit enough for a higher score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
logcat_grep - First observed
long_press_and_ask - First observed
press_key - First observed
record_and_ask - First observed
screenshot_ask - First observed
swipe_and_ask - First observed
tap_and_ask - First observed
type_text - First observed
vision_spend_report
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose: screenshot_ask is passive state-checking, while tap/swipe/long_press_and_ask each combine a specific gesture with screenshot-and-ask. record_and_ask targets animations, type_text and press_key are direct input without vision, logcat_grep handles logs, and vision_spend_report tracks cost. No two tools overlap in function.
All tool names follow a consistent snake_case pattern, with a clear <action>_and_ask convention for vision-verifying interactions and simple verb_noun for the rest. The naming logically separates gesture tools from non-vision tools, making the set easy to navigate.
Nine tools is well-scoped for a mobile automation/verification server. Each tool addresses a concrete need—actions, verification, logging, cost monitoring—and none feel redundant or purely decorative. The count fits the domain without bloat or sparsity.
The tool surface covers the full cycle of mobile UI interaction and verification: direct input (type_text, press_key), gestures (tap/swipe/long_press), visual state checking (screenshot_ask, record_and_ask), log inspection (logcat_grep), and cost governance (vision_spend_report). No obvious dead ends or missing operations for the stated purpose.
Maintenance
Related MCP Connectors
Disposable cloud Android emulators for coding agents: run an APK or PR build, tap, type, screenshot.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Cloud Android phones for AI agents: create a phone, read the screen, tap, type, install APKs, park.
Related MCP Servers
- AlicenseCqualityCmaintenanceA lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.142,158 PyPI880MIT
- AlicenseBqualityBmaintenanceEnables AI agents to control Android devices and emulators through direct UI interaction, allowing app navigation, automated testing, and real-world task execution via ADB without computer vision or scripts.182MIT
- AlicenseAqualityFmaintenanceProvides AI agents with real-time vision and control over Android devices through screen streaming, UI automation, and fast input control via scrcpy protocol.3318MIT
- AlicenseBqualityDmaintenanceEnables AI agents to control Android devices via ADB, supporting gestures, input, screenshots, UI analysis, and app management.1919 npmISC