DroidMind
✨DroidMind🤖
通过模型上下文协议使用 AI 控制 Android 设备
DroidMind 是 AI 助手与 Android 设备之间的强大桥梁,支持通过自然语言进行控制、调试和系统分析。通过实现模型上下文协议 (MCP),DroidMind 允许 AI 模型通过 ADB 以安全、结构化的方式直接与 Android 设备交互。当作为代理编码工作流程的一部分使用时,DroidMind 可以让您的助手直接在设备循环中构建和调试。
💫 功能
📱设备控制- 通过 USB 或 TCP/IP 连接到设备,运行 shell 命令,重新启动
📊系统分析- 检查设备属性,查看硬件信息,分析系统日志
🔍文件系统访问- 浏览目录内容并管理设备上的文件
📷可视化诊断- 捕获设备屏幕截图以进行分析和调试
📦应用程序管理- 在连接的设备上安装、卸载、启动、停止和清除应用程序数据
🔄多设备支持- 控制并在多个连接的设备之间切换
👆 UI 自动化- 通过点击、滑动、文本输入和按键与设备交互
🔍应用程序检查- 查看应用程序清单、共享首选项和特定于应用程序的日志
🔒安全框架- 通过全面的命令验证保护设备
💬 MCP 集成- 无缝连接到 Claude、Cursor、Cline 等
Related MCP server: MCP Android Agent
🚀 安装
# Clone the repository
git clone https://github.com/hyperbliss/droidmind.git
cd droidmind
# Set up a virtual environment with UV
uv venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install dependencies with UV
uv sync📋 先决条件
Python 3.13 或更高版本
已启用 USB 调试的 Android 设备
ADB(Android 调试桥)已安装并位于 PATH 中
UV 包管理器(推荐用于依赖项管理)
对于网络控制:启用了 ADB over TCP/IP 的 Android 设备
🔮 快速入门
运行 DroidMind 服务器
运行 DroidMind 作为服务器,通过 MCP 连接 AI 助手:
# Start DroidMind as a network server
droidmind --transport sse与人工智能助手一起使用
以 SSE 模式启动 DroidMind:
droidmind --transport sse使用 MCP 协议 URI 连接您的 AI 助手:
sse://localhost:4256/sse人工智能现在可以通过自然语言控制您的 Android 设备!
🛠️ 可用的 MCP 资源和工具
资源
devices://list- 列出所有连接的设备device://{serial}/properties- 获取详细的设备属性logs://{serial}/logcat- 从设备获取最近的日志logs://{serial}/anr- 获取应用程序无响应 (ANR) 跟踪logs://{serial}/crashes- 获取应用程序崩溃日志logs://{serial}/battery- 获取电池统计数据和历史记录logs://{serial}/app/{package}- 获取应用程序特定的日志fs://{serial}/list/{path}- 列出设备上的目录内容fs://{serial}/read/{path}- 从设备读取文件内容fs://{serial}/stats/{path}- 获取详细的文件/目录统计信息app://{serial}/{package}/manifest- 获取 AndroidManifest.xml 内容app://{serial}/{package}/data- 列出应用程序数据目录中的文件app://{serial}/{package}/shared_prefs- 获取应用程序的共享偏好设置
工具
devicelist- 列出所有已连接的 Android 设备device_properties- 获取特定设备的详细属性device_logcat- 从设备获取最近的 logcat 输出list_directory- 列出设备上目录的内容connect_device通过 TCP/IP 连接到设备disconnect_device- 断开与 Android 设备的连接shell_command- 在设备上运行 shell 命令install_app- 在设备上安装 APKuninstall_app- 从设备上卸载应用程序start_app- 在设备上启动应用程序stop_app- 强制停止设备上的应用程序clear_app_data- 清除应用程序数据和缓存list_packages- 列出设备上已安装的软件包get_app_manifest- 获取应用程序的 AndroidManifest.xml 内容get_app_permissions- 获取应用程序请求的权限get_app_activities- 获取应用中定义的活动get_app_info- 获取有关应用程序的详细信息reboot_device- 重启设备(正常、恢复或引导加载程序)screenshot- 从设备获取屏幕截图capture_bugreport- 从设备生成全面的错误报告dump_heap- 从正在运行的进程创建堆转储以进行内存分析push_file- 将文件上传到设备pull_file- 从设备下载文件delete_file- 从设备中删除文件或目录create_directory- 在设备上创建目录file_exists- 检查设备上是否存在文件read_file- 读取设备上文件的内容write_file- 将文本内容写入设备上的文件file_stats- 获取有关文件或目录的详细信息tap- 点击设备屏幕上的特定坐标swipe- 在屏幕上从一个点到另一个点执行滑动手势input_text- 在设备上输入文本,就像从键盘输入一样press_key- 按下硬件或软件键(例如 HOME、BACK、VOLUME)start_intent- 使用 Android Intent 启动应用活动
📊 AI 助手查询示例
将 AI 助手连接到 DroidMind 后,尝试以下查询:
“列出所有已连接的 Android 设备并显示其属性”
“连接到我的手机 192.168.1.100 并检查其电池状态”
“截取我的 Pixel 的屏幕截图并向我显示当前屏幕上的内容”
“检查我的设备上的可用存储空间”
“显示我设备的 ANR 跟踪和崩溃日志”
“查看最近的日志并告诉我是否有任何错误”
“在我的设备上安装这个 APK 文件并告诉我是否成功”
“显示我手机上所有已安装应用程序的列表”
“将我的设备重启至恢复模式”
“我的手机运行的是哪个版本的 Android?”
“检查我的设备是否已 root 并告诉我其安全补丁级别”
“显示 com.android.settings 的清单文件”
“检查我的应用程序的共享偏好设置”
“点击坐标 500,1000 处的“设置”图标”
“从屏幕顶部向下滑动即可打开通知栏”
“在当前文本字段中输入我的密码”
“按三次返回按钮即可返回主屏幕”
“通过启动 com.android.settings 包来打开“设置”应用”
🔒 安全功能
DroidMind 包含一个全面的安全框架来保护您的设备,同时仍允许 AI 助手发挥表达能力:
命令验证:所有 shell 命令都根据安全命令的允许列表进行验证
风险评估:命令按风险级别分类(安全、低、中、高、严重)
命令清理:对输入进行清理,以防止命令注入攻击
受保护的路径:系统目录和关键路径受到保护,不得修改
全面日志记录:所有命令均记录其风险级别以供审计
可疑模式检测:具有潜在危险模式的命令将被阻止
ADB 命令安全性:通过适当的异步验证对 ADB 特定命令进行特殊处理
安全系统设计得足够宽松,允许常见操作,同时防止破坏性操作。高风险命令会在执行前向用户显示警告,关键操作会被完全阻止,无需明确覆盖。
💻 开发
DroidMind 使用 UV 进行依赖项管理和开发工作流程。UV 是一个快速、可靠的 Python 包管理器和解析器。
# Update dependencies
uv sync
# Run tests
pytest
# Run linting
ruff check .
# Run type checking
mypy .🤝 贡献
欢迎贡献代码!欢迎提交 Pull 请求。
分叉存储库
创建你的功能分支(
git checkout -b feature/amazing-feature)使用 UV 设置您的开发环境
进行更改
运行测试和 linting
提交您的更改(
git commit -m 'Add some amazing feature')推送到分支(
git push origin feature/amazing-feature)打开拉取请求
📝 许可证
该项目根据 Apache 许可证获得许可 - 有关详细信息,请参阅 LICENSE 文件。
如果您发现 DroidMind 有用,请给我买一个 Monster Ultra Violet ⚡️
Available Tools
8 toolsandroid-appA
Perform various application management operations on an Android device.
This single tool consolidates various app-related actions. The 'action' parameter determines the operation.
Args:
serial: Device serial number.
action: The specific app operation to perform.
ctx: MCP Context for logging and interaction.
package (Optional[str]): Package name for the target application. Required by most actions.
apk_path (Optional[str]): Path to the APK file (local to the server). Used by install_app.
reinstall (Optional[bool]): Whether to reinstall if app exists. Used by install_app.
grant_permissions (Optional[bool]): Whether to grant all requested permissions. Used by install_app.
keep_data (Optional[bool]): Whether to keep app data and cache directories. Used by uninstall_app.
activity (Optional[str]): Optional activity name to start. Used by start_app.
extras (Optional[dict[str, str]]): Optional intent extras. Used by start_intent.
include_system_apps (Optional[bool]): Whether to include system apps. Used by list_packages.
include_app_name (Optional[bool]): Whether to include app labels (best-effort). Used by list_packages.
include_apk_path (Optional[bool]): Whether to include APK paths. Used by list_packages.
max_packages (Optional[int]): Max packages to return. Used by list_packages.
Returns: A string message indicating the result or status of the operation.
Available Actions and their specific argument usage:
action="install_app"Requires:
apk_pathOptional:
reinstall,grant_permissions
action="uninstall_app"Requires:
packageOptional:
keep_data
action="start_app"Requires:
packageOptional:
activity3b.action="start_intent"Requires:
package,activityOptional:
extras
action="stop_app"Requires:
package
action="clear_app_data"Requires:
package
action="list_packages"Optional:
include_system_apps,include_app_name,include_apk_path,max_packages
action="get_app_manifest"Requires:
package
action="get_app_permissions"Requires:
package
action="get_app_activities"Requires:
package
action="get_app_info"Requires:
package
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| action | Yes | ||
| package | No | ||
| apk_path | No | ||
| reinstall | No | ||
| grant_permissions | No | ||
| keep_data | No | ||
| activity | No | ||
| extras | No | ||
| include_system_apps | No | ||
| include_app_name | No | ||
| include_apk_path | No | ||
| max_packages | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries full burden. It details required vs optional parameters per action but does not disclose preconditions (e.g., device connection), side effects, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a summary, parameter list, action list, and per-action details. Though lengthy, it is justified given the complexity of 11 actions; however, the inclusion of 'ctx' parameter not in schema is a minor inconsistency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and low schema coverage, the description covers required/optional params per action and notes the return type. However, it omits device prerequisites, error scenarios, and includes an undocumented 'ctx' parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds significant meaning by explaining each parameter's purpose and mapping them to specific actions (e.g., 'apk_path' for install, 'package' for uninstall).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs 'various application management operations on an Android device' and lists 11 specific actions, making it distinct from sibling tools like android-shell or android-ui.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the tool consolidates app-related actions and lists each action with required parameters, but does not explicitly guide when to use this tool over alternatives or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-deviceA
Perform various device management operations on Android devices.
This single tool consolidates various device-related actions. The 'action' parameter determines the operation.
Args: action: The specific device operation to perform. ctx: MCP Context for logging and interaction. serial (Optional[str]): Device serial number. Required by most actions except connect/list. ip_address (Optional[str]): IP address for 'connect_device' action. port (Optional[int]): Port for 'connect_device' action (default: 5555). mode (Optional[str]): Reboot mode for 'reboot_device' action (default: "normal").
Returns: A string message indicating the result or status of the operation.
Available Actions and their specific argument usage:
action="list_devices"No specific arguments required beyond
ctx.
action="connect_device"Requires:
ip_addressOptional:
port
action="disconnect_device"Requires:
serial
action="reboot_device"Requires:
serialOptional:
mode(e.g., "normal", "recovery", "bootloader")
action="device_properties"Requires:
serial
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| serial | No | ||
| ip_address | No | ||
| port | No | ||
| mode | No | normal |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It details each action's required and optional parameters and notes return type. However, it lacks disclosure of side effects (e.g., reboot disconnects device) and error cases. Still, it is fairly transparent, earning a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a brief intro, standard arg list, and a clear enumeration of actions with their specific arguments. Every sentence adds value; no fluff. Score 5 for efficient organization.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, 5 actions, and no annotations, the description covers action-parameter dependencies and return type. It lacks error handling or examples, but is otherwise complete for a moderately complex tool. Score 4.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It effectively maps parameters to actions, clarifying which parameters are needed for which operation. This adds significant meaning beyond the schema, though detailed constraints (e.g., valid formats) are omitted. Score 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it performs device management operations and lists actions via an enum. While it does not explicitly differentiate from sibling tools like android-shell or android-diag, the action list provides clarity on scope. A score of 4 reflects clear purpose with minor sibling differentiation gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by enumerating actions but does not provide explicit when-to-use or when-not-to-use guidance compared to siblings. No alternatives or exclusions are mentioned. Score 3 for implied but unguided usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-diagA
Perform diagnostic operations like capturing bug reports or heap dumps.
Args: ctx: MCP Context. serial: Device serial number. action: The diagnostic action to perform. output_path: Optional. Path to save the output file. For bugreport: host path for adb to write the .zip. If empty, a temp file is used & summarized. For dump_heap: local path to save the .hprof. If empty, a temp file is used. include_screenshots: For CAPTURE_BUGREPORT. Default True. package_or_pid: For DUMP_HEAP. App package name or process ID. native: For DUMP_HEAP. True for native (C/C++) heap, False for Java. Default False. timeout_seconds: Max time for the operation. If 0, action-specific defaults are used (bugreport: 300s, dump_heap: 120s).
Returns: A string message indicating the result or path to the output.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| action | Yes | ||
| output_path | No | ||
| include_screenshots | No | ||
| package_or_pid | No | ||
| native | No | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It explains that output_path defaults to a temp file and is summarized, and mentions timeout defaults. However, it does not disclose potential side effects like performance impact or device lock requirements. It is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with separate Args and Returns sections, and front-loads the purpose. It is detailed but not excessively verbose; every sentence adds value. Minor redundancy in separately explaining actions, but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 params, no annotations, output schema exists), the description covers all parameters, default behaviors, and return value. It omits prerequisites like device connectivity and error handling, but is reasonably complete for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description adds meaning to all parameters: serial, action, output_path, include_screenshots, package_or_pid, native, and timeout_seconds. It explains defaults and behavior for optional paths. This compensates well for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Perform diagnostic operations like capturing bug reports or heap dumps.' It uses a specific verb and resource, and distinguishes from sibling tools that focus on other aspects like app, device, file, log, screenshot, shell, and UI. This leaves no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lacks explicit guidance on when to use this tool versus alternatives. It does not mention context such as troubleshooting device issues or any exclusions. The intended usage is only implied by the parameter descriptions, but no explicit when-to-use or when-not-to-use advice is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-fileA
Perform file and directory operations on an Android device.
This single tool consolidates various file system actions. The 'action' parameter determines the operation.
Args: serial: Device serial number. action: The specific file operation to perform. See available actions below. ctx: MCP Context for logging and interaction. path (Optional[str]): General path argument on the device. Used by: list_directory, delete_file, create_directory, file_exists, file_stats. Can also be used by read_file and write_file as an alternative to 'device_path'. local_path (Optional[str]): Path on the DroidMind server machine. Used by: push_file (source), pull_file (destination). device_path (Optional[str]): Path on the Android device. Used by: push_file (destination), pull_file (source), read_file (source), write_file (destination). If 'path' is also provided for read/write, 'device_path' takes precedence. content (Optional[str]): Text content to write. Used by: write_file. max_size (Optional[int]): Maximum file size in bytes for read_file (default: 100KB). Used by: read_file.
Returns: Union[str, bool]: A string message indicating the result or status for most actions. Returns a boolean for the 'file_exists' action.
Available Actions and their specific argument usage:
action="list_directory": Lists contents of a directory.Requires:
path(directory path on device).Returns: Formatted string of directory contents.
action="push_file": Uploads a file from the local server to the device.Requires:
local_path(source on server),device_path(destination on device).Returns: String message confirming upload.
action="pull_file": Downloads a file from the device to the local server.Requires:
device_path(source on device),local_path(destination on server).Returns: String message confirming download.
action="delete_file": Deletes a file or directory from the device.Requires:
path(path to delete on device).Returns: String message confirming deletion.
action="create_directory": Creates a directory on the device.Requires:
path(directory path to create on device).Returns: String message confirming creation.
action="file_exists": Checks if a file or directory exists on the device.Requires:
path(path to check on device).Returns:
Trueif exists,Falseotherwise.
action="read_file": Reads the contents of a file from the device.Requires:
device_path(orpath) for the file on device.Optional:
max_size(defaults to 100KB).Returns: String containing file contents or error message.
action="write_file": Writes text content to a file on the device.Requires:
device_path(orpath) for the file on device,content(text to write).Returns: String message confirming write.
action="file_stats": Gets detailed statistics for a file or directory.Requires:
path(path on device).Returns: Markdown-formatted string of file/directory statistics.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| action | Yes | ||
| path | No | ||
| local_path | No | ||
| device_path | No | ||
| content | No | ||
| max_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description fully covers behavior: explains each action, parameter precedence (device_path over path), defaults (max_size=100KB), and return types. Destructive operations are implied by action names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with sections and bullet points, front-loaded with purpose. Slightly verbose due to repeated 'Used by' lines, but every sentence adds value for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, 2 required, no annotations, and an output schema exists, the description provides complete guidance for all actions, including parameter mapping and return types, ensuring correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% parameter descriptions; description compensates thoroughly by explaining each parameter's purpose, which actions use them, precedence rules, and defaults, adding significant meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it performs file and directory operations on an Android device, listing 9 distinct actions, which differentiates it from sibling tools like android-app or android-shell.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied through action descriptions, but no explicit guidance on when to use this tool versus alternatives (e.g., android-shell for commands) or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-logA
Perform various log retrieval operations on an Android device.
This single tool consolidates various log-related actions. The 'action' parameter determines the operation.
Args:
serial: Device serial number.
action: The specific log operation to perform.
ctx: MCP Context for logging and interaction.
package (Optional[str]): Package name for get_app_logs action.
lines (int): Number of lines to fetch for logcat actions (default: 1000).
filter_expr (Optional[str]): Logcat filter expression for get_device_logcat.
buffer (Optional[str]): Logcat buffer for get_device_logcat (default: "main").
format_type (Optional[str]): Logcat output format for get_device_logcat (default: "threadtime").
max_size (Optional[int]): Max output size for get_device_logcat (default: 100KB).
Returns: A string message containing the requested logs or status.
Available Actions and their specific argument usage:
action="get_device_logcat"Optional:
lines,filter_expr,buffer,format_type,max_size.
action="get_app_logs"Requires:
package.Optional:
lines.
action="get_anr_logs"No specific arguments beyond
serialandctx.
action="get_crash_logs"No specific arguments beyond
serialandctx.
action="get_battery_stats"No specific arguments beyond
serialandctx.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| action | Yes | ||
| package | No | ||
| lines | No | ||
| filter_expr | No | ||
| buffer | No | main | |
| format_type | No | threadtime | |
| max_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes each action and its parameters, but lacks details on performance implications, error conditions, or prerequisites (e.g., USB debugging enabled). For a read-heavy tool, this is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat lengthy but well-structured with bullet points for each action, front-loading the purpose. Every sentence adds value, though a slightly more succinct summary of all actions could improve conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, 5 actions), the description covers all necessary usage details. The output schema is not described, but the return type ('string message') is mentioned. This is largely sufficient for an agent to decide and execute.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description adds significant meaning: it maps each parameter to specific actions, explains optional vs required, and provides defaults. This fully compensates for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs 'various log retrieval operations on an Android device' and lists five specific actions via the 'action' parameter. This distinguishes it from siblings like android-app, android-device, etc., which handle different domains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the 'action' parameter determines the operation and details which arguments are required or optional for each action. While it does not explicitly contrast with siblings, the context is sufficient for an agent to choose and invoke the tool correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-screenshotA
Get a screenshot from a device.
Args: serial: Device serial number ctx: MCP context quality: JPEG quality (1-100, lower means smaller file size)
Returns: The device screenshot as an image
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| quality | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided. The description only states it gets a screenshot but does not disclose behavioral traits such as whether the device needs to be unlocked, any side effects, or latency considerations. For a read operation, this minimal disclosure is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and uses a structured Args/Returns format, front-loading the purpose. It is slightly verbose due to docstring conventions but remains efficient. The inclusion of 'ctx' parameter not in schema is a minor inconsistency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (screenshot) and lack of output schema, the description adequately covers functionality, parameters, and return type. It could mention prerequisites like device connection, but overall completeness is satisfactory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds value by explaining 'serial: Device serial number' and 'quality: JPEG quality (1-100, lower means smaller file size)'. It partially compensates for missing schema descriptions, though the 'ctx' parameter in the docstring is not in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Get a screenshot from a device,' specifying the action and resource. It distinctly separates from sibling tools like android-app, android-device, etc., which focus on other device aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention contexts where a screenshot is appropriate or when to use other tools like android-shell or android-file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-shellA
Run a shell command on the device.
Args: serial: Device serial number command: Shell command to run max_lines: Maximum lines of output to return (default: 1000) Use positive numbers for first N lines, negative for last N lines Set to None for unlimited (not recommended for large outputs) max_size: Maximum output size in characters (default: 100000) Limits total response size regardless of line count
Returns: Command output
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| command | Yes | ||
| max_lines | No | ||
| max_size | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries full burden. It adds behavioral context about output limits (max_lines with sign logic, max_size, defaults) but does not disclose potential destructive effects, security implications, or error conditions of running shell commands.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with Args and Returns sections, making it easy to parse. Slightly verbose but each sentence adds value; no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that runs arbitrary commands, the description adequately covers input parameters but lacks details on output format (despite having output schema) and potential risks or prerequisites, such as device connectivity or authentication.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains all four parameters: serial, command, max_lines (with positive/negative line logic and default), and max_size (default and limit). This adds significant meaning beyond the schema's type-only definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Run a shell command on the device.' It uses a specific verb (Run) and resource (shell command) with context (on the device), and distinguishes from sibling tools that manage apps, devices, diagnostics, files, logs, screenshots, and UI.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor any conditions for appropriate use or exclusions. It only lists parameters without contextual usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
android-uiB
Perform various UI interaction operations on an Android device.
Args: ctx: MCP Context. serial: Device serial number. action: The UI action to perform. x: X coordinate (for tap). y: Y coordinate (for tap). start_x: Starting X coordinate (for swipe). start_y: Starting Y coordinate (for swipe). end_x: Ending X coordinate (for swipe). end_y: Ending Y coordinate (for swipe). duration_ms: Duration of the swipe in milliseconds (default: 300). text: Text to input (for input_text). keycode: Android keycode to press (for press_key). package: Package name (for start_intent). activity: Activity name (for start_intent). extras: Optional intent extras (for start_intent).
Returns: A string message indicating the result of the operation.
| Name | Required | Description | Default |
|---|---|---|---|
| serial | Yes | ||
| action | Yes | ||
| x | No | ||
| y | No | ||
| start_x | No | ||
| start_y | No | ||
| end_x | No | ||
| end_y | No | ||
| duration_ms | No | ||
| text | No | ||
| keycode | No | ||
| package | No | ||
| activity | No | ||
| extras | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should disclose behavioral traits. It only lists parameters and actions but does not describe side effects (e.g., app changes), prerequisites (e.g., unlocked device), or limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured as a docstring with Args and Returns sections, but it is somewhat lengthy and could be more concise by grouping related parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 14 parameters, no schema descriptions, and no annotations, the description provides basic parameter explanations but lacks usage context, error handling, or return value details beyond 'a string message'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to each parameter (e.g., 'x: X coordinate (for tap)'), compensating for the 0% schema description coverage. However, it could provide more detail on ranges or formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs various UI interaction operations on an Android device and lists the supported actions (tap, swipe, etc.). This distinguishes it from sibling tools like android-device or android-app.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not mention when not to use it or what other tools exist for different tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- Changed
android-app6 fields changed- changed
Input schema / $defs / AppAction / enumPrevious value: -[ - "install_app", - "uninstall_app", - "start_app", - "stop_app", - "clear_app_data", - "list_packages", - "get_app_manifest", - "get_app_permissions", - "get_app_activities", - "get_app_info" -]New value: +[ + "install_app", + "uninstall_app", + "start_app", + "start_intent", + "stop_app", + "clear_app_data", + "list_packages", + "get_app_manifest", + "get_app_permissions", + "get_app_activities", + "get_app_info" +] - added
Input schema / properties / extrasAdded value: +{ + "anyOf": [ + { + "additionalProperties": { + "type": "string" + }, + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Extras" +} - added
Input schema / properties / include_apk_pathAdded value: +{ + "default": true, + "title": "Include Apk Path", + "type": "boolean" +} - added
Input schema / properties / include_app_nameAdded value: +{ + "default": false, + "title": "Include App Name", + "type": "boolean" +} - added
Input schema / properties / max_packagesAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": 200, + "title": "Max Packages" +} - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "app_operationsOutput", + "type": "object" +}
- Changed
android-device1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "android_deviceOutput", + "type": "object" +}
- Changed
android-diag1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "android_diagOutput", + "type": "object" +}
- Changed
android-file1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "boolean" + } + ], + "title": "Result" + } + }, + "required": [ + "result" + ], + "title": "file_operationsOutput", + "type": "object" +}
- Changed
android-log1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "android_logOutput", + "type": "object" +}
- Changed
android-shell1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "shell_commandOutput", + "type": "object" +}
- Changed
android-ui1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "android_uiOutput", + "type": "object" +}
8 tool updates
- First observed
android-app - First observed
android-device - First observed
android-diag - First observed
android-file - First observed
android-log - First observed
android-screenshot - First observed
android-shell - First observed
android-ui
TDQS
Scored across 8 tools
The top-level tools are clearly categorized, but the internal actions within each mega-tool can overlap (e.g., start_intent appears in both android-app and android-ui), causing potential ambiguity for an agent. Additionally, the bundling of many operations under one tool name makes it less obvious which tool handles a specific action.
All tool names follow a consistent 'android-<category>' pattern, and internal actions use a uniform snake_case style. The naming is predictable and readable across the entire server.
With 8 tools, the count is reasonable for an Android device management server. However, the bundling of many actions into each tool makes the surface feel slightly under-tooled, though still within the effective range.
The tool set covers major areas: app management, device control, diagnostics, file operations, logs, screenshots, shell, and UI. Some gaps exist (e.g., network or settings management), but the core workflows for common Android tasks are well addressed.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
MCP server for building and testing AI agents with multi-model experimentation and insights.
The Telnyx MCP server is an official implementation of the Model Context Protocol that enables AI clients (like Claude Desktop, Cursor, and OpenAI Agents) to interact with Telnyx's telephony, messaging, and AI assistant APIs. It provides comprehensive capabilities including making and managing phone calls, sending SMS/MMS messages, purchasing and configuring phone numbers, creating AI assistants with custom instructions, managing cloud storage buckets, scraping and embedding website content, and handling integration secrets. The server exists as both a local implementation and a remotely hosted version, allowing developers to integrate real-world communication infrastructure directly into AI applications.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI assistants to interact with Android devices through ADB, allowing for automated device management, app installation, file transfers, and screenshot capture.480 npm38ISC
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI agents to control and automate Android devices through natural language, supporting actions like app management, UI interactions, and device monitoring.59MIT
- FlicenseAqualityDmaintenanceA MCP server that enables AI assistants to control Android devices via ADB, supporting device info, screen control, input simulation, app management, shell execution, file transfer, and UI parsing.20-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that wraps Android ADB functionality into AI assistant tools, enabling device management, shell execution, file operations, app management, media capture, and log analysis via natural language.21 npm2MIT