Skip to main content
Glama
antonpinchuk

mobile-mcp-opengl

by antonpinchuk

OpenGL Android 개발 및 자동화를 위한 MCP

AI 코딩 에이전트(Claude Code, Cursor 등)가 전체 UI가 단일 불투명 OpenGL/Vulkan/Metal 표면 내부에 그려지는 Android 앱(Cocos2d-x, Unity, Unreal, 원시 OpenGL, libGDX 및 유사 엔진)을 테스트할 수 있게 해주는 MCP 서버입니다.

이 서버가 해결하는 문제

adb shell uiautomator dump 및 모든 접근성 트리 기반 자동화 도구(대부분의 MCP 모바일 자동화 서버 포함)는 네이티브 Android 뷰 계층(버튼, 라벨, 텍스트 및 좌표)을 검사하여 작동합니다. 이는 네이티브 뷰로 구성된 일반적인 Android UI에는 훌륭하게 작동합니다.

하지만 전체 UI를 하나의 GLSurfaceView 내부에 텍스처로 렌더링하는 게임이나 앱에는 작동하지 않습니다. 접근성 트리의 관점에서 보면 화면에는 자식, 라벨, 내부 좌표가 전혀 없는 단 하나의 불투명 뷰만 존재합니다. 검사할 것이 없습니다. 화면은 실제로 얼마나 많은 UI가 있든 관계없이 블랙박스입니다.

남은 유일한 실제 관찰 채널은 스크린샷입니다. 이 서버는 이를 가끔 사용하는 폴백이 아니라 일반적인 경우로 삼아 구축되었습니다.

mobile-mcp와의 차이점

mobile-next/mobile-mcp는 범용 MCP 모바일 자동화 서버이며, 일반 네이티브 앱에는 좋은 기본 선택입니다. 접근성 트리를 우선 사용하고(빠르고 저렴하며 비전 모델이나 이미지 토큰이 필요 없음), 트리가 필요한 정보를 제공하지 못할 때만 스크린샷 + 좌표로 폴백합니다.

OpenGL 캔버스 앱의 경우 이 폴백은 가끔 발생하는 것이 아니라 매번 작동하는 유일한 경로입니다. mobile-mcp-opengl은 이 특정 경우를 위해 구축되었으며, 그 결과 두 가지 다른 설계 선택을 합니다:

  1. 접근성 트리 시도를 전혀 하지 않습니다. 시도해도 이 앱에서는 항상 비어 있으므로 얻을 이득이 없습니다. 따라서 여기의 모든 도구는 바로 스크린샷 + 비전으로 진행합니다.

  2. 비전 분석은 호출 에이전트를 실행하는 모델이 아니라 플러그 가능한 별도 공급자(아래 참조)를 통해 진행됩니다. 게임에 대한 기능 QA 루프는 세션당 수백 번의 스크린샷 검사에 쉽게 도달할 수 있습니다. 이 모든 것을 기본 코딩 에이전트의 자체 비전으로 라우팅하면 실제 비용과 실제 코딩 작업에 사용하고 싶은 토큰/컨텍스트가 모두 소모됩니다. 여기서 스크린샷 바이트는 호출 에이전트의 컨텍스트에 전혀 들어가지 않습니다. 공급자의 짧은 텍스트 답변만 들어갑니다.

Related MCP server: Android-MCP

별도의 기본 도구가 아닌 결합된 액션+관찰 도구를 사용하는 이유

단순한 설계는 tap, screenshot, ask를 세 개의 별도 도구로 노출합니다. 그러면 호출 에이전트가 모든 상호작용에 대해 다단계 루프를 조율해야 합니다: 탭 → 스크린샷 촬영 → 비전 단계에 전달 → 결과 읽기 → 다음 작업 결정. 각각은 별도의 도구 호출이자 별도의 턴입니다. 실제 테스트 로직 대신 조정에 토큰을 소모하고, 에이전트가 단계를 빠뜨리거나 순서를 잘못 지정하거나 호출 사이에 오래된 상태에 대해 추론할 가능성이 더 커집니다.

대신 이 서버는 결합된 도구(tap_and_ask, swipe_and_ask, long_press_and_ask)를 노출합니다. 이 도구들은 액션을 수행하고, 잠시 기다린 후, 스크린샷을 찍고, 비전 공급자에게 질문하고, 하나의 짧은 답변을 반환합니다. 모두 단일 도구 호출로 처리됩니다. 다단계 테스트 시나리오는 의미 있는 검사당 대략 한 번의 에이전트 턴이 소요되며, 3~4턴이 아닙니다.

일반 screenshot_ask(액션 없이 관찰만) 및 저렴한 비비전 도구(type_text, press_key, logcat_grep)도 테스트 흐름에서 이 패턴이 필요하지 않은 부분에 사용할 수 있습니다.

도구

도구

기능

비전 호출?

screenshot_ask

스크린샷을 찍은 후, 그에 대한 짧은 질문을 합니다

예

tap_and_ask

(x, y)를 탭하고, 기다린 후, 스크린샷을 찍고, 질문합니다

예

swipe_and_ask

(x1,y1)→(x2,y2)로 스와이프/드래그하고, 기다린 후, 스크린샷을 찍고, 질문합니다

예

long_press_and_ask

(x, y)를 일정 시간 길게 누르고, 기다린 후, 스크린샷을 찍고, 질문합니다

예

record_and_ask

선택적 액션 후, 시간 간격을 두고 N개의 스크린샷을 찍고, 각 프레임에 대해 같은 질문을 합니다

예 (N회 호출)

type_text

현재 포커스된 필드에 텍스트를 입력합니다

아니요

press_key

Android KEYCODE_* 이벤트(뒤로, 엔터 등)를 보냅니다

아니요

logcat_grep

최근 logcat을 읽고, 선택적으로 정규식으로 필터링합니다

아니요

vision_spend_report

오늘의 누적 비전 지출 및 임계값을 보고합니다

아니요

필요한 것이 이미 로그 라인(크래시, 자체 디버그 출력, 네트워크 오류)에 있다면 비전 호출보다 logcat_grep을 선호하세요. 무료이고 정확하며, 비전 호출은 둘 다 아닙니다.

애니메이션 확인: record_and_ask

단일 프레임 도구는 무언가가 애니메이션되는지 여부를 알 수 없습니다(강도 표시기가 부드럽게 맥동하는지, 라벨이 위로 날아가며 사라지는지, 스프라이트가 시작 위치로 튕겨 돌아오는지). record_and_ask는 하나의 선택적 액션(탭 또는 스와이프, 또는 둘 다 아님)을 수행하고, waitMs(첫 프레임 전에 UI가 반응하기 시작할 시간 — tap_and_ask/swipe_and_ask의 waitMs와 같은 의미)를 기다린 후, intervalMs 간격으로 frameCount개의 스크린샷을 캡처하고, 각 프레임에 대해 하나의 짧은 답변을 반환합니다. 호출 에이전트는 N개의 별도 스크린샷+질문 왕복을 직접 조율하는 대신 한 번의 도구 호출로 타임라인을 얻습니다.

모든 프레임을 하나의 호출로 묶는 대신 프레임당 하나의 비전 호출을 사용하는 이유. Runware의 imageCaption은 문서화된 단일 inputImage와 함께 문서화되지 않은 inputImages 배열(복수)을 허용하는 것으로 밝혀졌습니다. API에 직접 테스트했습니다. 정확히 2개의 이미지(같은 요청 내 before/after 비교)에서는 깔끔하게 작동했습니다. 한 요청에 3개 이상의 이미지가 있으면 해당 배열 매개변수와 수동으로 합성한 나란히 붙인 "필름스트립" 이미지 모두 테스트에서 잘리거나 잘못된 답변을 생성했습니다. 작은 7B 비전 모델은 한 번의 호출에서 결합된 시각+명령 부하가 일정 수준을 넘으면 일관성을 잃는 것으로 보입니다. 순차적 단일 이미지 호출(이 도구의 접근 방식)은 테스트한 모든 프레임 수에서 안정적이었고, 비용이 크게 증가하지도 않습니다. 비용은 호출 수보다 응답 길이(아래 참조)에 의해 지배되므로, N개의 짧은 순차 답변은 하나의 긴 다중 이미지 답변과 비슷하거나 더 저렴합니다. 자체 공급자가 다중 이미지 요청을 더 안정적으로 처리한다면 이는 분명히 최적화할 부분입니다. "자체 모델 사용"을 참조하세요.

설정

git clone <this repo>
cd mobile-mcp-opengl
npm install
cp .env.example .env
# edit .env: at minimum set RUNWARE_API_KEY (or switch VISION_PROVIDER, see below)

PATH에 adb가 필요하거나(또는 .env에 ADB_PATH 설정), 실행 중/연결된 기기 또는 에뮬레이터가 필요합니다. 둘 이상이 연결된 경우 ADB_DEVICE_SERIAL을 설정하세요(adb devices 참조).

Claude Code에 등록

프로젝트 루트에 .mcp.json을 추가하세요(이 파일은 일반적으로 프로젝트 로컬이며 git-ignored입니다. 일반적으로 머신별 경로를 가리키거나 머신별 환경 변수 재정의를 보유하기 때문입니다):

{
  "mcpServers": {
    "mobile-opengl": {
      "command": "node",
      "args": ["/absolute/path/to/mobile-mcp-opengl/src/server.js"]
    }
  }
}

Claude Code는 이 파일을 프로젝트에 대해 자동으로 인식합니다. 서버는 모든 구성에 대해 자체 .env(이 저장소의 package.json 옆)를 읽습니다. 호출 에이전트는 API 키를 알거나 전달할 필요가 없습니다.

비용 모델 — 긴 QA 세션을 실행하기 전에 읽어보세요

비용을 결정하는 것은 이미지 크기가 아니라 응답 길이입니다. 이는 기본 Runware/Qwen2.5-VL-7B-Instruct 공급자에 대해 경험적으로 측정되었습니다. 강제로 한 단어 답변을 요구한 동일한 질문은 360×360에서 1600×2400(레티나급)까지 모든 이미지 크기에서 동일한 비용($0.0006)이 들었습니다. 동일한 1024×1024 이미지에 개방형 "이것을 설명하세요" 프롬프트는 $0.0013–0.0019가 들었습니다. 이는 이미지가 더 커서가 아니라 모델이 더 긴 답변을 작성했기 때문에 2~3배 더 비쌌습니다.

실용적 의미:

  • 보내기 전에 스크린샷을 다운샘플링하지 마세요. 이 공급자에게는 비용을 의미 있게 줄이지 않으며, 필요할 수 있는 세부 정보를 잃게 됩니다.

  • 항상 짧은 답변을 강제하도록 질문을 구성하세요: 예/아니요, 숫자, 짧은 라벨, 몇 개의 필드가 있는 작은 JSON 객체. 이 서버의 모든 도구는 자동으로 짧은 답변 지침을 추가하지만, 모호한 개방형 질문("무엇이 보이나요?")은 특정 질문("오류 대화상자가 보이나요? 예/아니요")보다 모델이 더 긴 답변을 하도록 유도할 수 있습니다.

잘 구성된 짧은 질문의 경우 호출당 약 $0.0006이며, 500회 호출 QA 세션은 약 $0.30입니다. 동일한 양의 개방형 "화면 설명" 질문은 2~3배 더 비쌀 수 있습니다.

내장 지출 가드레일

모든 비전 호출은 .vision-log.jsonl(JSONL, 호출당 하나의 항목: 타임스탬프, 질문, 답변, 비용)에 기록됩니다. 이 로그 위에는 두 가지 독립적인 보호 장치가 있으며, 둘 다 공급자에 구애받지 않습니다(공급자가 보고하는 costUsd를 기반으로 작동):

  • 호출당 경고 (VISION_ALERT_USD, 기본 $0.0015): 단일 호출이 이 값을 초과하면 도구 응답에 [COST ALERT] 메모가 포함되어 모델이 짧은 답변 지침을 무시했을 가능성이 있음을 알려줍니다. 이는 조용히 넘어갈 것이 아니라 질문을 다시 구성하라는 신호입니다.

  • 일일 상한 (VISION_SESSION_CAP_USD, 기본 $2.00): 오늘의 누적 기록 지출이 이 값에 도달하면 이후의 모든 비전 호출은 (공급자에 도달하기 전에) 완전히 거부됩니다. 상한을 올리거나 날짜가 바뀔 때까지입니다. 이는 경고가 아니라 폭주 루프에 대한 하드 스톱입니다.

vision_spend_report를 호출하여 기기나 비전 호출 없이 오늘의 총액을 확인하세요.

공급자가 비용을 보고할 수 없는 경우(아래 openai-compatible 참조), 해당 공급자의 호출은 costUsd: null로 기록되며 경고를 트리거하거나 상한에 포함되지 않습니다. 가드레일은 가시성이 없는 지출을 보호할 수 없습니다.

자체 모델 사용

비전 분석은 src/providers/visionProvider.js를 통해 진행되며, .env의 VISION_PROVIDER에서 이름으로 공급자를 선택합니다. 두 가지가 내장되어 있습니다:

  • runware (기본값) — Runware.ai의 imageCaption 작업에 직접 연결하며, 기본적으로 Qwen2.5-VL-7B-Instruct(AIR id runware:152@2)를 사용합니다. Runware와 OpenRouter는 별도의 API 키와 모델 카탈로그를 가진 두 개의 별도 서비스입니다. 이 서버는 OpenRouter를 통하지 않고 Runware에 직접 연결합니다.

  • openai-compatible — OpenAI 채팅 완성 비전 형식(image_url 콘텐츠 부분)을 말하는 모든 것에 대한 일반 공급자입니다. OpenRouter, 비전 모델을 실행하는 로컬 Ollama/LM Studio 서버, Groq, Together.ai 또는 기타 호환 엔드포인트에서 작동합니다. .env에서 OPENAI_COMPATIBLE_BASE_URL, OPENAI_COMPATIBLE_API_KEY, OPENAI_COMPATIBLE_MODEL을 구성하세요. 대부분의 OpenAI 호환 API는 고정 달러 비용 대신 토큰 사용량을 보고합니다. 이 공급자가 costUsd를 추정하도록 하려면 OPENAI_COMPATIBLE_PRICE_PER_1M_INPUT/_OUTPUT을 설정하세요(그렇지 않으면 위 메모에 따라 이 공급자에 대해 비용 추적/가드레일이 비활성화됩니다).

완전히 사용자 정의 공급자(자체 호스팅 모델, 완전히 다른 API 형태)를 추가하려면 src/providers/openaiCompatibleProvider.js를 시작점으로 복사하고 다음을 구현하세요:

async function ask(imageBuffer, mimeType, question) {
  // return { text: string, costUsd: number | null }
}
module.exports = { ask };

그리고 src/providers/visionProvider.js의 loadProvider()에 이름으로 등록하세요.

라이선스

MIT


Kinect.PRO 개발

Available Tools

9 tools
logcat_grepRead recent logcat, filteredA

Read the last N logcat lines, optionally filtered by a regex (e.g. your app's tag, or "Exception|FATAL"). No vision call, no cost - prefer this over screenshot_ask whenever what you need is already in a log line (crashes, your own debug prints, network errors).

ParametersJSON Schema
NameRequiredDescriptionDefault
linesNoHow many recent lines to fetch (default 200).
filterRegexNoOptional regex; only matching lines are returned.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It clearly frames the operation as a read ('Read the last N logcat lines'), implying no mutation, and adds resource-related behavior ('No vision call, no cost'). It does not detail empty-result behavior or regex error handling, but for a non-destructive log reader this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, each earning its place: the first states the core operation, the second provides selection guidance and cost context. No redundant phrases or unnecessary details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given only two optional parameters and no output schema, the description covers the action, filtering, selection criteria, and cost trade-off. It is complete enough for an agent to invoke correctly, though it does not spell out behavior for empty results or invalid regex.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents both parameters with 100% coverage, so the baseline is 3. The description adds practical regex examples ('your app's tag, Exception|FATAL') and clarifies that the filter is optional, providing contextual guidance beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and a clear resource ('logcat lines'), and it explicitly distinguishes itself from a sibling tool (screenshot_ask) by stating 'No vision call, no cost'. An agent can immediately tell what this tool does and how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use rule: 'prefer this over screenshot_ask whenever what you need is already in a log line,' followed by concrete examples (crashes, debug prints, network errors). It also explains the cost advantage, making the selection decision clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

long_press_and_askLong-press + screenshot + askA

Long-press at (x, y) for durationMs, wait briefly, take a screenshot, and ask a short question about the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
waitMsNoMilliseconds to wait after releasing before screenshotting (default 500).
questionYes
durationMsNoHold duration in ms (default 800).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing behavior. It transparently lists the operation sequence and references default waitMs/durationMs defaults in the schema. However, it does not disclose side effects of long-pressing (e.g., opening context menus or triggering navigation), what the 'ask' returns or to whom, or any required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the core action and includes the full workflow without filler. Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, no annotations, and no output schema, the description leaves important gaps: the return/response behavior is ambiguous ('ask a short question about the result'), coordinate system is unspecified, and side effects are not mentioned. An agent would need additional implicit knowledge to call this tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions only cover waitMs and durationMs (40% coverage). The description helps by framing x and y as long-press coordinates and question as a short question about the result. Still, it does not specify coordinate units/origin or any constraints on the question, so it only partially compensates for the schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action sequence: long-press at (x, y) for durationMs, wait, screenshot, and ask a question. The long-press gesture clearly differentiates it from sibling tools like tap_and_ask, swipe_and_ask, and screenshot_ask, even without naming them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: perform a long-press and inspect the resulting screen via a screenshot and question. However, it does not explicitly state when to prefer this over tap_and_ask, swipe_and_ask, or other siblings, nor does it mention any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyPress hardware/virtual keyA

Send an Android keyevent code (e.g. 4 = BACK, 66 = ENTER, 187 = APP_SWITCH). No vision call.

ParametersJSON Schema
NameRequiredDescriptionDefault
keycodeYesAndroid KEYCODE_* integer value.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description is the only source of behavior. It discloses the core action (sending a keycode) and that it does not use vision, but it does not clarify whether the key is pressed and released with a single event or describe timing/duration. Lacks details on side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with two clauses; the key action is front-loaded, and each part (action, examples, vision exclusion) adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Simple tool with one required parameter and no output schema. The description covers what it does and gives examples, so an agent can invoke it correctly. It does not explain return behavior or errors, but those are likely unnecessary for a fire-and-forget key event. Minor missing context about when to use it is covered under usage guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents keycode as an Android KEYCODE_* integer. The description adds specific example values (4=BACK, 66=ENTER, 187=APP_SWITCH), which clarify the range and meaning significantly beyond the schema's generic description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific verb 'send' and resource 'Android keyevent code', gives concrete examples distinguishing it from vision-based siblings like screenshot_ask and tap_and_ask, and explicitly notes 'No vision call.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides minimal guidance on when to use; the 'No vision call' implies it is not for visual tasks, but it does not explicitly name alternatives or conditions for selection. The examples imply use for system keys but lack explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_and_askRecord a timed screenshot sequence + ask about each frameA

For checking an ANIMATION or any effect that plays out over time (e.g. "does the strength indicator pulse smoothly?", "does the XP label fly up and fade out?", "does the sprite return to its start position?"). Optionally performs one action first (tap or swipe, or neither), then takes frameCount screenshots spaced intervalMs apart, and asks the SAME short question about each frame separately (each frame gets its own vision call, with its frame number in the prompt) - returns one answer per frame in order.

Sequential single-frame calls were chosen over sending several frames in one request: Runware's imageCaption does accept an undocumented multi-image array, and it works fine for exactly 2 frames, but degrades noticeably at 3+ (truncated/malformed answers in testing) - sequential calls are both more reliable and, per-frame, no more expensive. Keep frameCount modest (3-6) - each frame is a full separate vision call and cost scales linearly with it.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNoRequired for action=tap or action=swipe (swipe start x).
yNoRequired for action=tap or action=swipe (swipe start y).
x2NoRequired for action=swipe (end x).
y2NoRequired for action=swipe (end y).
actionYesAction to perform before starting the capture sequence.
waitMsNoMilliseconds to wait after the action before the FIRST screenshot (default 500) - same meaning as waitMs in tap_and_ask/swipe_and_ask, separate from intervalMs which spaces out the frames after that.
questionYesThe same short question asked about every captured frame (e.g. "Is the indicator visible? yes/no").
frameCountYesHow many screenshots to take, spaced intervalMs apart (2-8; keep modest, see description).
intervalMsYesMilliseconds between each screenshot (i.e. the sampling interval of the sequence).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it reveals that each frame gets its own vision call with the frame number in the prompt, that answers come back one per frame in order, and that sequential calls were deliberately chosen over multi-image requests due to reliability degradation. This gives the agent accurate expectations about cost and behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average, but every sentence contributes: purpose, action semantics, per-frame behavior, return ordering, rationale for sequential calls, and cost guidance. The key use case is front-loaded, and the engineering rationale is placed where it helps the agent decide rather than adding noise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters, no output schema, and no annotations, the description covers the core interaction fully: what triggers the sequence, what each frame does, how the question is applied, what the result order is, and cost implications. The remaining parameter details are already well documented in the input schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds extra value by advising frameCount stay modest (3-6), explaining linear cost scaling, and clarifying that waitMs is distinct from intervalMs and has the same meaning as in tap_and_ask/swipe_and_ask. This goes beyond the schema's structural descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific use case—checking an animation or effect that plays out over time—and clearly states the mechanism: perform an optional action, take frameCount screenshots spaced intervalMs apart, and ask the same question about each frame. This distinguishes it from single-shot siblings like screenshot_ask or tap_and_ask without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: for animations or time-based effects. It also explains when the optional action is tap, swipe, or neither. However, it does not explicitly name alternative tools or state when NOT to use this one, so the guidance is clear but not fully contrastive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshot_askScreenshot + askA

Take a screenshot of the current screen and ask a short question about it (e.g. "Is there an error dialog visible?", "How many word icons are on screen?", "What color is the strength indicator?"). Use this when you need to check state WITHOUT performing an action first. Phrase the question so a short answer is possible (yes/no, a number, a short label) - see this server's README "Cost model".

ParametersJSON Schema
NameRequiredDescriptionDefault
questionYesA short, specific question about the current screen.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly communicates a non-action read-only behavior and implies the response is short ('so a short answer is possible'). It points to the README for cost details, adding context. While it doesn't describe the exact return format, the answer is implied by the question-asking purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (about 3 sentences) and front-loaded with the main action. Every sentence earns its place: statement of action, examples, usage guidance, and cost reference. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only tool without an output schema, the description is largely complete. It covers purpose, usage timing, question phrasing, and cost considerations. It could explicitly mention that the result is an answer to the question, but that is reasonably implied. The description is sufficient for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description already clarifies the question. The tool description adds value by providing examples of valid questions and guidance on phrasing for short answers, which enriches the parameter's semantics beyond the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair ('Take a screenshot of the current screen and ask a short question about it') and provides clear examples. It also distinguishes itself from siblings by specifying 'WITHOUT performing an action first', which separates it from action-based tools like tap_and_ask or swipe_and_ask.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use this tool ('when you need to check state WITHOUT performing an action first') and gives concrete guidance on phrasing questions for short answers. The reference to the README 'Cost model' provides additional usage context. This fully addresses when to use instead of alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

swipe_and_askSwipe/drag + screenshot + askA

Swipe (or drag, for drag-and-drop UIs) from (x1, y1) to (x2, y2), wait briefly, take a screenshot, and ask a short question about the result - all in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
x1Yes
x2Yes
y1Yes
y2Yes
waitMsNoMilliseconds to wait after the swipe before screenshotting (default 500).
questionYesA short, specific question about the screen after the swipe.
durationMsNoSwipe duration in ms (default 300; use longer for drag-and-drop hold gestures).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and does disclose the full ordered behavior: swipe, wait, screenshot, ask. It also notes the duration nuance for drag-and-drop holds. It doesn't detail coordinate units or what 'ask' returns, but the step sequence is clearly communicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence captures the entire workflow with no filler. The core action is front-loaded, and the drag-and-drop nuance is efficiently folded into the gesture description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple composite gesture tool, but given no output schema and no annotations, it omits the return/answer semantics of 'ask', the coordinate system, and any cost or side-effect implications. These gaps prevent it from being fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description must compensate. It explains x1/y1/x2/y2 as the swipe's start and end points, and clarifies that question should be short and about the post-swipe screen. The optional waitMs and durationMs already have schema descriptions, and the prose adds the 'hold for drag-and-drop' nuance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: swipe/drag from one coordinate to another, wait, screenshot, and ask a question. It clearly distinguishes itself from sibling tools like tap_and_ask and screenshot_ask by naming the gesture ('Swipe') and the compound workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction to use drag 'for drag-and-drop UIs' gives some contextual guidance, but there is no explicit when-to-use/when-not-to-use statement or reference to alternatives. The appropriate context is implied by the swipe gesture rather than directly contrasted with siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tap_and_askTap + screenshot + askA

Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call. Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesX coordinate in device pixels.
yYesY coordinate in device pixels.
waitMsNoMilliseconds to wait after the tap before screenshotting (default 500).
questionYesA short, specific question about the screen after the tap.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It describes the sequence of actions (tap, wait, screenshot, ask) and mentions a wait period before screenshotting. However, it does not disclose what happens if the tap fails, the exact return format (e.g., does it return an image, a text answer, or both?), or any side effects like requiring a running app. The description is adequate for basic behavior but lacks depth on error handling or output, which is significant for a composite tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no redundancy. The first sentence front-loads the action sequence, and the second sentence provides direct usage guidance. Every word serves a purpose, and the structure is clean and immediately understandable. It avoids jargon and is well-scoped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a composite tool with 4 parameters and no output schema, the description should cover the response format and any prerequisites. While it clearly states the purpose and usage, it does not mention what the tool returns (e.g., an answer to the question, a screenshot reference) or any necessary preconditions (e.g., the device being interactive). Since there is no output schema, the description's silence on return values leaves an agent without complete information for correctly interpreting the tool's outcome. This is a notable gap, so a 3.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds minimal semantic value beyond the schema: it refers to 'wait briefly' which maps to waitMs, and characterizes the question as 'short and specific', but does not explain coordinate units or default wait behavior beyond what the schema provides. The description does not go beyond the schema definitions, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb and resource: 'Tap at device screen coordinates (x, y), wait briefly for the UI to react, take a screenshot, and ask a short question about the result - all in one call.' It distinguishes itself from the siblings by defining its specific action (tap) and explicitly contrasting with calling separate tap/screenshot/ask tools. The title also reinforces the composite nature, so there is no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Use this for any "tap here, then check what happened" step instead of calling separate tap/screenshot/ask tools.' This clearly defines the usage context and names an alternative (separate tools). It does not explicitly mention sibling tools like swipe_and_ask or long_press_and_ask, but the reference to 'tap' inherently implies a distinction from those. This is strong guidance, but not exhaustive about exclusions from all siblings, hence a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textType textA

Type text into whatever field currently has focus. No screenshot/vision call - pair with screenshot_ask if you need to confirm the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that typing targets the focused field and that screenshot/vision is not part of the operation, but it does not cover edge cases such as no focused field, whether existing text is replaced, or how special characters are handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the core behavior is front-loaded, and the follow-up guidance about screenshot_ask earns its place. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool, the description is nearly sufficient: it states the target, the action, and the verification route. Missing failure-mode detail, such as what happens when no field has focus, keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not add meaning beyond the property name 'text'. The parameter is simple, but nothing explains format, newline behavior, limits, or encoding, so the description fails to compensate for the schema's lack of documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete action ('Type text') and a specific target ('whatever field currently has focus'), and explicitly warns against treating it as a screenshot/vision operation. This clearly differentiates it from screenshot_ask and the other _and_ask siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational guidance: use it when a field has focus, and pair it with screenshot_ask when confirmation is needed. It does not explicitly contrast with press_key or other input tools, but the focus-based behavior is enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vision_spend_reportReport today's vision spendA

Report the cumulative vision-provider spend for today and the configured alert/cap thresholds, without making any device or vision call.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It openly states 'without making any device or vision call', which discloses side-effect-free behavior, but it does not describe the output format, potential delays, or any other behavioral aspects. This is a reasonable disclosure but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the verb and resource. Every word adds value, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description covers the essential information: what is reported and the guarantee of no side effects. It does not specify the return format or any prerequisites, but given the simplicity, it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description adds meaning about what the report contains (spend and thresholds) beyond the empty schema, which is exactly what is needed for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Report' and identifies the exact resource ('cumulative vision-provider spend for today') plus the alert/cap thresholds. It is clearly distinct from the sibling action-oriented tools (screenshot, tap, etc.) by stating it makes no device or vision call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for checking spend information and explicitly notes it does not make any device or vision call, but it does not name alternative tools or provide explicit when-to-use guidance. The use case is somewhat obvious given the sibling list, but the guidance is not explicit enough for a higher score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observedlogcat_grep
    • First observedlong_press_and_ask
    • First observedpress_key
    • First observedrecord_and_ask
    • First observedscreenshot_ask
    • First observedswipe_and_ask
    • First observedtap_and_ask
    • First observedtype_text
    • First observedvision_spend_report

TDQS

A4.3/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: screenshot_ask is passive state-checking, while tap/swipe/long_press_and_ask each combine a specific gesture with screenshot-and-ask. record_and_ask targets animations, type_text and press_key are direct input without vision, logcat_grep handles logs, and vision_spend_report tracks cost. No two tools overlap in function.

Naming Consistency5/5

All tool names follow a consistent snake_case pattern, with a clear <action>_and_ask convention for vision-verifying interactions and simple verb_noun for the rest. The naming logically separates gesture tools from non-vision tools, making the set easy to navigate.

Tool Count5/5

Nine tools is well-scoped for a mobile automation/verification server. Each tool addresses a concrete need—actions, verification, logging, cost monitoring—and none feel redundant or purely decorative. The count fits the domain without bloat or sparsity.

Completeness5/5

The tool surface covers the full cycle of mobile UI interaction and verification: direct input (type_text, press_key), gestures (tap/swipe/long_press), visual state checking (screenshot_ask, record_and_ask), log inspection (logcat_grep), and cost governance (vision_spend_report). No obvious dead ends or missing operations for the stated purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    C
    maintenance
    A lightweight bridge enabling AI agents to perform real-world tasks on Android devices such as app navigation, UI interaction, and automated QA testing without requiring computer-vision pipelines or preprogrammed scripts.
    14
    2,158 PyPI
    880
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Enables AI agents to control Android devices and emulators through direct UI interaction, allowing app navigation, automated testing, and real-world task execution via ADB without computer vision or scripts.
    18
    2
    MIT