mobile-mcp-opengl
OpenGL Android開発と自動化のためのMCP
AIコーディングエージェント(Claude Code、Cursorなど)が、UI全体が単一の不透明なOpenGL/Vulkan/Metalサーフェス内に描画されるAndroidアプリ(Cocos2d-x、Unity、Unreal、生のOpenGL、libGDXなど)をテストするためのMCPサーバーです。
これが解決する問題
adb shell uiautomator dumpや、アクセシビリティツリーに基づくすべての自動化ツール(ほとんどのMCPモバイル自動化サーバーを含む)は、ネイティブのAndroidビュー階層(ボタン、ラベル、そのテキストと座標)を検査することで動作します。これは、ネイティブビューで構築された通常のAndroid UIではうまく機能します。
しかし、UI全体を1つのGLSurfaceView内のテクスチャとしてレンダリングするゲームやアプリでは機能しません。アクセシビリティツリーの観点からは、画面上には子要素もラベルも、内部の何かの座標もない、単一の不透明なビューしか存在しません。検査するものは何もありません。実際にどれだけのUIが表示されていても、画面はブラックボックスです。
残された唯一の実際の観測手段はスクリーンショットです。このサーバーは、その事実を通常のケースとして、まれなフォールバックとしてではなく、中心に据えて構築されています。
mobile-mcpとの違い
mobile-next/mobile-mcpは汎用のMCPモバイル自動化サーバーであり、通常のネイティブアプリには良いデフォルトの選択肢です。アクセシビリティツリーを優先し(高速、低コスト、ビジョンモデル不要、画像トークン不要)、ツリーが必要な情報を提供しない場合にのみスクリーンショット+座標にフォールバックします。
OpenGLキャンバスアプリの場合、そのフォールバックはまれなものではなく、毎回機能する唯一の経路です。mobile-mcp-openglはそのケースに特化して構築されており、その結果として2つの異なる設計上の選択を行っています:
アクセシビリティツリーの試行は一切行いません。 試しても無駄です。これらのアプリでは常に空で返ってくるため、ここにあるすべてのツールは直接スクリーンショット+ビジョンに進みます。
ビジョン分析は、プラグイン可能な別のプロバイダーを経由します(下記参照)。呼び出し元エージェントを実行しているモデルではありません。ゲームに対する機能QAループは、セッションあたり数百回のスクリーンショットチェックに簡単に達する可能性があります。そのすべてをメインのコーディングエージェント自身のビジョンにルーティングすると、実際のコーディング作業に使いたいお金とトークン/コンテキストの両方がかかります。ここでは、スクリーンショットのバイトは呼び出し元エージェントのコンテキストに一切入りません。プロバイダーの短いテキスト回答だけが入ります。
Related MCP server: Android-MCP
なぜ個別のプリミティブではなく、アクション+オブザーブを組み合わせたツールなのか
素朴な設計では、tap、screenshot、askを3つの別々のツールとして公開します。これにより、呼び出し元エージェントはすべての操作に対してマルチステップのループを調整する必要があります:タップ→スクリーンショットを撮る→ビジョンステップに渡す→結果を読む→次に何をするか決定する。それぞれが別々のツール呼び出しであり、別々のターンです。実際のテストロジックではなく調整にトークンを消費し、エージェントがステップを落としたり、順序を間違えたり、呼び出し間で古い状態について推論したりする余地が広がります。
代わりに、このサーバーは組み合わせたツール(tap_and_ask、swipe_and_ask、long_press_and_ask)を公開します。これらはアクションを実行し、少し待って、スクリーンショットを撮り、ビジョンプロバイダーに質問し、1つの短い回答を返します。すべて単一のツール呼び出しです。マルチステップのテストシナリオは、意味のあるチェックごとに約1エージェントターンで済み、3つや4つにはなりません。
このパターンを必要としないテストフローの部分には、単純なscreenshot_ask(観察のみ、アクションなし)と安価な非ビジョンツール(type_text、press_key、logcat_grep)も利用できます。
ツール
ツール | 機能 | ビジョン呼び出し? |
| スクリーンショットを撮り、それについて短い質問をする | はい |
| (x, y)をタップし、待って、スクリーンショットを撮り、質問する | はい |
| (x1,y1)→(x2,y2)にスワイプ/ドラッグし、待って、スクリーンショットを撮り、質問する | はい |
| (x, y)を一定時間長押しし、待って、スクリーンショットを撮り、質問する | はい |
| オプションのアクション、その後時間を空けてN枚のスクリーンショットを撮り、各フレームについて同じ質問をする | はい(N回) |
| 現在フォーカスされているフィールドに入力する | いいえ |
| Androidの | いいえ |
| 最近のlogcatを読み、オプションで正規表現でフィルタリングする | いいえ |
| 今日の累積ビジョン支出としきい値を報告する | いいえ |
必要な情報がすでにログ行にある場合(クラッシュ、独自のデバッグ出力、ネットワークエラー)は、ビジョン呼び出しよりもlogcat_grepを優先してください。無料で正確ですが、ビジョン呼び出しはどちらでもありません。
アニメーションの確認:record_and_ask
単一フレームのツールでは、何かが正しくアニメーションしているかどうかを知ることはできません(強度インジケーターが滑らかに脈動するか、ラベルが飛び上がってフェードアウトするか、スプライトが開始位置に跳ね返るか)。record_and_askは、1つのオプションのアクション(タップまたはスワイプ、またはどちらでもない)を実行し、waitMs待って(tap_and_ask/swipe_and_askのwaitMsと同じ意味 — 最初のフレームの前にUIが反応し始めるまでの時間)、その後intervalMs間隔でframeCount枚のスクリーンショットを撮り、フレームごとに1つの短い回答を返します。呼び出し元エージェントは、N回の個別のスクリーンショット+質問の往復を自分で調整する代わりに、1回のツール呼び出しでタイムラインを取得します。
なぜフレームごとに1回のビジョン呼び出しなのか、すべてのフレームをまとめて1回の呼び出しではないのか。 RunwareのimageCaptionは、文書化された単一のinputImageに加えて、文書化されていないinputImages配列(複数形)を受け入れることが判明しました — APIに対して直接テスト済みです。正確に2枚の画像では問題なく機能します(同じリクエスト内での前後比較は正しく一貫した結果が返りました)。1リクエストで3枚以上の画像の場合、その配列パラメータと手動で合成したサイドバイサイドの「フィルムストリップ」画像の両方が、テストで切り詰められた、または不正な回答を生成しました。小さな7Bビジョンモデルは、1回の呼び出しで視覚+指示の負荷が一定を超えると、どうやら一貫性を失うようです。逐次的な単一画像呼び出し(このツールのアプローチ)は、テストしたどのフレーム数でも信頼性が高く、コストも実質的に高くありません。コストは呼び出し回数ではなく応答長によって支配されるため(下記参照)、N回の短い逐次回答は、1回の長い複数画像回答とほぼ同じか、それ以下のコストになります。独自のプロバイダーが複数画像リクエストをより確実に処理できる場合は、これが明らかに最適化のポイントです — 「独自モデルを使用する」を参照してください。
セットアップ
git clone <this repo>
cd mobile-mcp-opengl
npm install
cp .env.example .env
# edit .env: at minimum set RUNWARE_API_KEY (or switch VISION_PROVIDER, see below)PATHにadbが必要です(または.envでADB_PATHを設定)。また、実行中/接続済みのデバイスまたはエミュレーターが必要です。複数接続されている場合は、ADB_DEVICE_SERIALを設定してください(adb devicesを参照)。
Claude Codeに登録する
プロジェクトのルートに.mcp.jsonを追加します(このファイルは通常プロジェクトローカルでgit無視されます。通常はマシン固有のパスを指すか、マシン固有の環境変数の上書きを保持するためです):
{
"mcpServers": {
"mobile-opengl": {
"command": "node",
"args": ["/absolute/path/to/mobile-mcp-opengl/src/server.js"]
}
}
}Claude Codeはこれをプロジェクト用に自動的に取得します。サーバーはすべての設定を独自の.env(このリポジトリのpackage.jsonの隣)から読み取ります。呼び出し元エージェントはAPIキーを知ったり渡したりする必要は一切ありません。
コストモデル — 長時間のQAセッションを実行する前にこれを読んでください
コストを左右するのは応答長であり、画像サイズではありません。 これはデフォルトのRunware/Qwen2.5-VL-7B-Instructプロバイダーに対して経験的に測定されました。同じ質問を強制的に1語で回答させた場合、360×360から1600×2400(retinaクラス)までの画像サイズで同じコスト($0.0006)でした。同じ1024×1024の画像で、自由回答の「これを説明してください」というプロンプトでは$0.0013〜0.0019かかりました — 2〜3倍です。これは純粋にモデルがより長い回答を書いたためであり、画像が大きいためではありません。
実際的な影響:
送信前にスクリーンショットをダウンサンプリングする必要はありません — このプロバイダーではコストを実質的に削減できず、必要な詳細を失う可能性があります。
常に短い回答を強制するように質問を組み立ててください:はい/いいえ、数字、短いラベル、いくつかのフィールドを持つ小さなJSONオブジェクト。このサーバーのすべてのツールは自動的に短い回答の指示を追加しますが、曖昧な自由回答の質問(「何が見えますか?」)は、具体的な質問(「エラーダイアログは表示されていますか?はい/いいえ」)よりも長い回答にモデルを押しやる可能性があります。
適切に構成された短い質問で約$0.0006/呼び出しの場合、500回のQAセッションは約$0.30かかります。同じ量の自由回答の「画面を説明してください」という質問は、その2〜3倍になる可能性があります。
組み込みの支出ガードレール
すべてのビジョン呼び出しは.vision-log.jsonl(JSONL、呼び出しごとに1エントリ:タイムスタンプ、質問、回答、コスト)に記録されます。そのログの上に2つの独立した保護が置かれ、どちらもプロバイダーに依存しません(プロバイダーが報告するcostUsdに基づいて機能します):
呼び出しごとのアラート(
VISION_ALERT_USD、デフォルト$0.0015):単一の呼び出しがこれを超えて返された場合、ツールの応答には[COST ALERT]メモが含まれ、モデルが短い回答の指示を無視した可能性があることを知らせます。これは質問を再構成するためのシグナルであり、黙って受け入れるものではありません。日次上限(
VISION_SESSION_CAP_USD、デフォルト$2.00):今日の累積記録支出がこれに達すると、それ以降のすべてのビジョン呼び出しは(プロバイダーに到達する前に)完全に拒否されます。上限が引き上げられるか、日が変わるまでです。これは暴走ループに対するハードストップであり、単なる警告ではありません。
vision_spend_reportをいつでも呼び出して、デバイスやビジョン呼び出しを行わずに今日の合計を確認できます。
プロバイダーがコストを報告できない場合(下記のopenai-compatibleを参照)、そのプロバイダーからの呼び出しはcostUsd: nullで記録され、アラートをトリガーしたり上限にカウントされたりすることはありません。ガードレールは、可視性のない支出を保護することはできません。
独自モデルを使用する
ビジョン分析はsrc/providers/visionProvider.jsを経由します。これは.envのVISION_PROVIDERから名前でプロバイダーを選択します。組み込みは2つです:
runware(デフォルト)— Runware.aiのimageCaptionタスクに直接通信し、デフォルトでQwen2.5-VL-7B-Instruct(AIR IDrunware:152@2)を使用します。RunwareとOpenRouterは別々のサービスで、別々のAPIキーとモデルカタログを持っています。これはOpenRouterを介さずにRunwareに直接通信します。openai-compatible— OpenAIチャット完了ビジョン形式(image_urlコンテンツパーツ)を話すもののための汎用プロバイダーです。OpenRouter、ビジョンモデルを実行するローカルのOllama/LM Studioサーバー、Groq、Together.ai、その他の互換性のあるエンドポイントで動作します。.envでOPENAI_COMPATIBLE_BASE_URL、OPENAI_COMPATIBLE_API_KEY、OPENAI_COMPATIBLE_MODELを設定します。ほとんどのOpenAI互換APIは、固定のドルコストではなくトークン使用量を報告します。このプロバイダーにcostUsdを推定させたい場合は、OPENAI_COMPATIBLE_PRICE_PER_1M_INPUT/_OUTPUTを設定してください(そうしないと、上記の注記のとおり、このプロバイダーではコスト追跡/ガードレールは無効になります)。
完全にカスタムなプロバイダー(自己ホストモデル、まったく異なるAPI形状)を追加するには、src/providers/openaiCompatibleProvider.jsを出発点としてコピーし、以下を実装します:
async function ask(imageBuffer, mimeType, question) {
// return { text: string, costUsd: number | null }
}
module.exports = { ask };そして、src/providers/visionProvider.jsのloadProvider()に名前を付けて登録します。
ライセンス
MIT
Kinect.PROによって開発されました
Available Tools
9 toolslogcat_grepRead recent logcat, filteredA
Read the last N logcat lines, optionally filtered by a regex (e.g. your app's tag, or "Exception|FATAL"). No vision call, no cost - prefer this over screenshot_ask whenever what you need is already in a log line (crashes, your own debug prints, network errors).
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | How many recent lines to fetch (default 200). | |
| filterRegex | No | Optional regex; only matching lines are returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly frames the operation as a read ('Read the last N logcat lines'), implying no mutation, and adds resource-related behavior ('No vision call, no cost'). It does not detail empty-result behavior or regex error handling, but for a non-destructive log reader this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, each earning its place: the first states the core operation, the second provides selection guidance and cost context. No redundant phrases or unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only two optional parameters and no output schema, the description covers the action, filtering, selection criteria, and cost trade-off. It is complete enough for an agent to invoke correctly, though it does not spell out behavior for empty results or invalid regex.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents both parameters with 100% coverage, so the baseline is 3. The description adds practical regex examples ('your app's tag, Exception|FATAL') and clarifies that the filter is optional, providing contextual guidance beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a clear resource ('logcat lines'), and it explicitly distinguishes itself from a sibling tool (screenshot_ask) by stating 'No vision call, no cost'. An agent can immediately tell what this tool does and how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use rule: 'prefer this over screenshot_ask whenever what you need is already in a log line,' followed by concrete examples (crashes, debug prints, network errors). It also explains the cost advantage, making the selection decision clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
long_press_and_askLong-press + screenshot + askA
Long-press at (x, y) for durationMs, wait briefly, take a screenshot, and ask a short question about the result.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| waitMs | No | Milliseconds to wait after releasing before screenshotting (default 500). | |
| question | Yes | ||
| durationMs | No | Hold duration in ms (default 800). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. It transparently lists the operation sequence and references default waitMs/durationMs defaults in the schema. However, it does not disclose side effects of long-pressing (e.g., opening context menus or triggering navigation), what the 'ask' returns or to whom, or any required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core action and includes the full workflow without filler. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no annotations, and no output schema, the description leaves important gaps: the return/response behavior is ambiguous ('ask a short question about the result'), coordinate system is unspecified, and side effects are not mentioned. An agent would need additional implicit knowledge to call this tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions only cover waitMs and durationMs (40% coverage). The description helps by framing x and y as long-press coordinates and question as a short question about the result. Still, it does not specify coordinate units/origin or any constraints on the question, so it only partially compensates for the schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action sequence: long-press at (x, y) for durationMs, wait, screenshot, and ask a question. The long-press gesture clearly differentiates it from sibling tools like tap_and_ask, swipe_and_ask, and screenshot_ask, even without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: perform a long-press and inspect the resulting screen via a screenshot and question. However, it does not explicitly state when to prefer this over tap_and_ask, swipe_and_ask, or other siblings, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyPress hardware/virtual keyA
Send an Android keyevent code (e.g. 4 = BACK, 66 = ENTER, 187 = APP_SWITCH). No vision call.
| Name | Required | Description | Default |
|---|---|---|---|
| keycode | Yes | Android KEYCODE_* integer value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description is the only source of behavior. It discloses the core action (sending a keycode) and that it does not use vision, but it does not clarify whether the key is pressed and released with a single event or describe timing/duration. Lacks details on side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with two clauses; the key action is front-loaded, and each part (action, examples, vision exclusion) adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple tool with one required parameter and no output schema. The description covers what it does and gives examples, so an agent can invoke it correctly. It does not explain return behavior or errors, but those are likely unnecessary for a fire-and-forget key event. Minor missing context about when to use it is covered under usage guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents keycode as an Android KEYCODE_* integer. The description adds specific example values (4=BACK, 66=ENTER, 187=APP_SWITCH), which clarify the range and meaning significantly beyond the schema's generic description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb 'send' and resource 'Android keyevent code', gives concrete examples distinguishing it from vision-based siblings like screenshot_ask and tap_and_ask, and explicitly notes 'No vision call.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides minimal guidance on when to use; the 'No vision call' implies it is not for visual tasks, but it does not explicitly name alternatives or conditions for selection. The examples imply use for system keys but lack explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_and_askRecord a timed screenshot sequence + ask about each frameA
For checking an ANIMATION or any effect that plays out over time (e.g. "does the strength indicator pulse smoothly?", "does the XP label fly up and fade out?", "does the sprite return to its start position?"). Optionally performs one action first (tap or swipe, or neither), then takes frameCount screenshots spaced intervalMs apart, and asks the SAME short question about each frame separately (each frame gets its own vision call, with its frame number in the prompt) - returns one answer per frame in order.
Sequential single-frame calls were chosen over sending several frames in one request: Runware's imageCaption does accept an undocumented multi-image array, and it works fine for exactly 2 frames, but degrades noticeably at 3+ (truncated/malformed answers in testing) - sequential calls are both more reliable and, per-frame, no more expensive. Keep frameCount modest (3-6) - each frame is a full separate vision call and cost scales linearly with it.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Required for action=tap or action=swipe (swipe start x). | |
| y | No | Required for action=tap or action=swipe (swipe start y). | |
| x2 | No | Required for action=swipe (end x). | |
| y2 | No | Required for action=swipe (end y). | |
| action | Yes | Action to perform before starting the capture sequence. | |
| waitMs | No | Milliseconds to wait after the action before the FIRST screenshot (default 500) - same meaning as waitMs in tap_and_ask/swipe_and_ask, separate from intervalMs which spaces out the frames after that. | |
| question | Yes | The same short question asked about every captured frame (e.g. "Is the indicator visible? yes/no"). | |
| frameCount | Yes | How many screenshots to take, spaced intervalMs apart (2-8; keep modest, see description). | |
| intervalMs | Yes | Milliseconds between each screenshot (i.e. the sampling interval of the sequence). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers: it reveals that each frame gets its own vision call with the frame number in the prompt, that answers come back one per frame in order, and that sequential calls were deliberately chosen over multi-image requests due to reliability degradation. This gives the agent accurate expectations about cost and behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence contributes: purpose, action semantics, per-frame behavior, return ordering, rationale for sequential calls, and cost guidance. The key use case is front-loaded, and the engineering rationale is placed where it helps the agent decide rather than adding noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, no output schema, and no annotations, the description covers the core interaction fully: what triggers the sequence, what each frame does, how the question is applied, what the result order is, and cost implications. The remaining parameter details are already well documented in the input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra value by advising frameCount stay modest (3-6), explaining linear cost scaling, and clarifying that waitMs is distinct from intervalMs and has the same meaning as in tap_and_ask/swipe_and_ask. This goes beyond the schema's structural descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific use case—checking an animation or effect that plays out over time—and clearly states the mechanism: perform an optional action, take frameCount screenshots spaced intervalMs apart, and ask the same question about each frame. This distinguishes it from single-shot siblings like screenshot_ask or tap_and_ask without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool: for animations or time-based effects. It also explains when the optional action is tap, swipe, or neither. However, it does not explicitly name alternative tools or state when NOT to use this one, so the guidance is clear but not fully contrastive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshot_askScreenshot + askA
Take a screenshot of the current screen and ask a short question about it (e.g. "Is there an error dialog visible?", "How many word icons are on screen?", "What color is the strength indicator?"). Use this when you need to check state WITHOUT performing an action first. Phrase the question so a short answer is possible (yes/no, a number, a short label) - see this server's README "Cost model".
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | A short, specific question about the current screen. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly communicates a non-action read-only behavior and implies the response is short ('so a short answer is possible'). It points to the README for cost details, adding context. While it doesn't describe the exact return format, the answer is implied by the question-asking purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (about 3 sentences) and front-loaded with the main action. Every sentence earns its place: statement of action, examples, usage guidance, and cost reference. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool without an output schema, the description is largely complete. It covers purpose, usage timing, question phrasing, and cost considerations. It could explicitly mention that the result is an answer to the question, but that is reasonably implied. The description is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter description already clarifies the question. The tool description adds value by providing examples of valid questions and guidance on phrasing for short answers, which enriches the parameter's semantics beyond the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair ('Take a screenshot of the current screen and ask a short question about it') and provides clear examples. It also distinguishes itself from siblings by specifying 'WITHOUT performing an action first', which separates it from action-based tools like tap_and_ask or swipe_and_ask.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool ('when you need to check state WITHOUT performing an action first') and gives concrete guidance on phrasing questions for short answers. The reference to the README 'Cost model' provides additional usage context. This fully addresses when to use instead of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
swipe_and_askSwipe/drag + screenshot + askA
Swipe (or drag, for drag-and-drop UIs) from (x1, y1) to (x2, y2), wait briefly, take a screenshot, and ask a short question about the result - all in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| x1 | Yes | ||
| x2 | Yes | ||
| y1 | Yes | ||
| y2 | Yes | ||
| waitMs | No | Milliseconds to wait after the swipe before screenshotting (default 500). | |
| question | Yes | A short, specific question about the screen after the swipe. | |
| durationMs | No | Swipe duration in ms (default 300; use longer for drag-and-drop hold gestures). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and does disclose the full ordered behavior: swipe, wait, screenshot, ask. It also notes the duration nuance for drag-and-drop holds. It doesn't detail coordinate units or what 'ask' returns, but the step sequence is clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence captures the entire workflow with no filler. The core action is front-loaded, and the drag-and-drop nuance is efficiently folded into the gesture description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple composite gesture tool, but given no output schema and no annotations, it omits the return/answer semantics of 'ask', the coordinate system, and any cost or side-effect implications. These gaps prevent it from being fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate. It explains x1/y1/x2/y2 as the swipe's start and end points, and clarifies that question should be short and about the post-swipe screen. The optional waitMs and durationMs already have schema descriptions, and the prose adds the 'hold for drag-and-drop' nuance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: swipe/drag from one coordinate to another, wait, screenshot, and ask a question. It clearly distinguishes itself from sibling tools like tap_and_ask and screenshot_ask by naming the gesture ('Swipe') and the compound workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction to use drag 'for drag-and-drop UIs' gives some contextual guidance, but there is no explicit when-to-use/when-not-to-use statement or reference to alternatives. The appropriate context is implied by the swipe gesture rather than directly contrasted with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tap_and_askTap + screenshot + askA
Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call. Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in device pixels. | |
| y | Yes | Y coordinate in device pixels. | |
| waitMs | No | Milliseconds to wait after the tap before screenshotting (default 500). | |
| question | Yes | A short, specific question about the screen after the tap. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes the sequence of actions (tap, wait, screenshot, ask) and mentions a wait period before screenshotting. However, it does not disclose what happens if the tap fails, the exact return format (e.g., does it return an image, a text answer, or both?), or any side effects like requiring a running app. The description is adequate for basic behavior but lacks depth on error handling or output, which is significant for a composite tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy. The first sentence front-loads the action sequence, and the second sentence provides direct usage guidance. Every word serves a purpose, and the structure is clean and immediately understandable. It avoids jargon and is well-scoped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a composite tool with 4 parameters and no output schema, the description should cover the response format and any prerequisites. While it clearly states the purpose and usage, it does not mention what the tool returns (e.g., an answer to the question, a screenshot reference) or any necessary preconditions (e.g., the device being interactive). Since there is no output schema, the description's silence on return values leaves an agent without complete information for correctly interpreting the tool's outcome. This is a notable gap, so a 3.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal semantic value beyond the schema: it refers to 'wait briefly' which maps to waitMs, and characterizes the question as 'short and specific', but does not explain coordinate units or default wait behavior beyond what the schema provides. The description does not go beyond the schema definitions, so a 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's verb and resource: 'Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call.' It distinguishes itself from the siblings by defining its specific action (tap) and explicitly contrasting with calling separate tap/screenshot/ask tools. The title also reinforces the composite nature, so there is no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: 'Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.' This clearly defines the usage context and names an alternative (separate tools). It does not explicitly mention sibling tools like swipe_and_ask or long_press_and_ask, but the reference to 'tap' inherently implies a distinction from those. This is strong guidance, but not exhaustive about exclusions from all siblings, hence a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textType textA
Type text into whatever field currently has focus. No screenshot/vision call - pair with screenshot_ask if you need to confirm the result.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that typing targets the focused field and that screenshot/vision is not part of the operation, but it does not cover edge cases such as no focused field, whether existing text is replaced, or how special characters are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the core behavior is front-loaded, and the follow-up guidance about screenshot_ask earns its place. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool, the description is nearly sufficient: it states the target, the action, and the verification route. Missing failure-mode detail, such as what happens when no field has focus, keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not add meaning beyond the property name 'text'. The parameter is simple, but nothing explains format, newline behavior, limits, or encoding, so the description fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action ('Type text') and a specific target ('whatever field currently has focus'), and explicitly warns against treating it as a screenshot/vision operation. This clearly differentiates it from screenshot_ask and the other _and_ask siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational guidance: use it when a field has focus, and pair it with screenshot_ask when confirmation is needed. It does not explicitly contrast with press_key or other input tools, but the focus-based behavior is enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vision_spend_reportReport today's vision spendA
Report the cumulative vision-provider spend for today and the configured alert/cap thresholds, without making any device or vision call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It openly states 'without making any device or vision call', which discloses side-effect-free behavior, but it does not describe the output format, potential delays, or any other behavioral aspects. This is a reasonable disclosure but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the verb and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description covers the essential information: what is reported and the guarantee of no side effects. It does not specify the return format or any prerequisites, but given the simplicity, it is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description adds meaning about what the report contains (spend and thresholds) beyond the empty schema, which is exactly what is needed for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Report' and identifies the exact resource ('cumulative vision-provider spend for today') plus the alert/cap thresholds. It is clearly distinct from the sibling action-oriented tools (screenshot, tap, etc.) by stating it makes no device or vision call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for checking spend information and explicitly notes it does not make any device or vision call, but it does not name alternative tools or provide explicit when-to-use guidance. The use case is somewhat obvious given the sibling list, but the guidance is not explicit enough for a higher score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
logcat_grep - First observed
long_press_and_ask - First observed
press_key - First observed
record_and_ask - First observed
screenshot_ask - First observed
swipe_and_ask - First observed
tap_and_ask - First observed
type_text - First observed
vision_spend_report
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose: screenshot_ask is passive state-checking, while tap/swipe/long_press_and_ask each combine a specific gesture with screenshot-and-ask. record_and_ask targets animations, type_text and press_key are direct input without vision, logcat_grep handles logs, and vision_spend_report tracks cost. No two tools overlap in function.
All tool names follow a consistent snake_case pattern, with a clear <action>_and_ask convention for vision-verifying interactions and simple verb_noun for the rest. The naming logically separates gesture tools from non-vision tools, making the set easy to navigate.
Nine tools is well-scoped for a mobile automation/verification server. Each tool addresses a concrete need—actions, verification, logging, cost monitoring—and none feel redundant or purely decorative. The count fits the domain without bloat or sparsity.
The tool surface covers the full cycle of mobile UI interaction and verification: direct input (type_text, press_key), gestures (tap/swipe/long_press), visual state checking (screenshot_ask, record_and_ask), log inspection (logcat_grep), and cost governance (vision_spend_report). No obvious dead ends or missing operations for the stated purpose.
Maintenance
Related MCP Connectors
Disposable cloud Android emulators for coding agents: run an APK or PR build, tap, type, screenshot.
Control real Android and iOS devices with LLM agents — tap, swipe, type, automate flows.
Cloud Android phones for AI agents: create a phone, read the screen, tap, type, install APKs, park.
Related MCP Servers
- AlicenseCqualityCmaintenanceA lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.142,158 PyPI880MIT
- AlicenseBqualityBmaintenanceEnables AI agents to control Android devices and emulators through direct UI interaction, allowing app navigation, automated testing, and real-world task execution via ADB without computer vision or scripts.182MIT
- AlicenseAqualityFmaintenanceProvides AI agents with real-time vision and control over Android devices through screen streaming, UI automation, and fast input control via scrcpy protocol.3318MIT
- AlicenseBqualityDmaintenanceEnables AI agents to control Android devices via ADB, supporting gestures, input, screenshots, UI analysis, and app management.1919 npmISC