mcp-remote-macos-use
MCP 서버 - 원격 MacOS 사용
AI가 원격 macOS 시스템을 완벽하게 제어할 수 있도록 하는 최초의 오픈소스 MCP 서버입니다.
OpenAI Operator에 대한 직접적인 대안으로, 완전한 데스크톱 기능을 갖춘 자율 AI 에이전트에 특별히 최적화되었으며, 추가 소프트웨어 설치가 필요 없습니다.
쇼케이스
트위터를 조사하고 트위터에 게시하세요( https://www.youtube.com/watch?v=--QHz2jcvcs )
CapCut을 사용하여 짧은 하이라이트 영상을 만들어 보세요( https://www.youtube.com/watch?v=RKAqiNoU8ec )
AI 채용 담당자: 메일 앱을 사용하여 후보자 정보 수집을 자동화하고, 지원서 심사 및 스크리닝 세션을 전송합니다.
AI 마케팅 인턴: LinkedIn 참여 - 관련 사용자를 자동으로 팔로우하고, 좋아요를 누르고, 댓글을 달기
AI 마케팅 인턴: 트위터 참여 - 관련 사용자에 대한 자동 팔로우, 좋아요 및 댓글 달기
할 일 목록(우선순위 지정)
성능 최적화 - Ubuntu 데스크톱 대안의 속도 맞추기
Apple Scripts 생성 - 유연성을 유지하면서 실행 시간 단축
VNC 커서 가시성 - 디버깅 및 데모 경험 개선
기여를 환영합니다!
Related MCP server: macos-control-mcp
특징
추가 API 비용 없음 : 기존 Claude Pro 플랜으로 무료 화면 처리
최소 설정 : 대상 Mac에서 화면 공유를 활성화하기만 하면 됩니다. 추가 소프트웨어가 필요하지 않습니다.
범용 호환성 : 현재 및 미래의 모든 macOS 버전과 호환됩니다.
우리가 이것을 만든 이유
타협 없는 네이티브 macOS 경험
macOS 네이티브 생태계는 오늘날 사용자 경험 측면에서 독보적인 위치를 차지하고 있으며, 앞으로도 오랫동안 최고의 기준으로 자리매김할 것입니다. 바로 이 부분에서 인간의 능력이 진정으로 발휘되며, 이제 AI도 이러한 환경에서 동일한 수준의 유창함을 유지할 수 있습니다.
디자인에 따른 개방형 아키텍처
범용 LLM 호환성 : 선택한 모든 MCP 클라이언트와 함께 작업 가능
모델 유연성 : OpenAI, Anthropic 또는 기타 LLM 공급자와 원활하게 통합
미래 지향적 통합 : MCP 생태계와 함께 발전하도록 설계됨
간편한 배포
대상 컴퓨터에서 제로 설정 : macOS에서 백그라운드 애플리케이션이나 에이전트가 필요하지 않습니다.
화면 공유만 있으면 됩니다 . 화면 공유를 활성화하여 모든 Mac을 제어하세요.
백엔드 복잡성 제거 : Python 애플리케이션이나 백그라운드 서비스 실행이 필요한 다른 솔루션과 달리
간소화된 부트스트랩 프로세스
Claude Desktop의 세련된 UI 활용 : 개발자 스타일의 Python 인터페이스가 필요 없습니다.
직관적인 사용자 경험 : 친숙하고 사용자 친화적인 인터페이스를 통해 AI가 제어하는 Mac과 상호 작용하세요.
즉각적인 생산성 : 구성의 번거로움 없이 즉시 작업을 시작하세요
건축학
설치
MacOs에서 화면 공유 활성화 macstadium.com에서 Mac을 대여한 경우 이 단계를 건너뛸 수 있습니다.
Claude Desktop에 이 MCP 서버를 추가합니다. Claude 구성에 다음을 추가하여 Docker 이미지를 사용하도록 Claude Desktop을 구성할 수 있습니다.
지엑스피1
LiveKit을 통한 WebRTC 지원
이 서버에는 이제 LiveKit 통합을 통한 WebRTC 지원이 포함되어 다음을 사용할 수 있습니다.
저지연 실시간 화면 공유
향상된 성능 및 반응성
기존 VNC에 비해 네트워크 효율성이 더 우수합니다.
네트워크 상황에 따른 자동 품질 조정
WebRTC 기능을 사용하려면 다음이 필요합니다.
LiveKit 서버를 설정하거나 LiveKit Cloud를 사용하세요
위의 구성 예제에 표시된 대로 LiveKit 환경 변수를 구성합니다.
개발자 지침
저장소를 복제합니다
# Clone the repository
git clone https://github.com/yourusername/mcp-remote-macos-use.git
cd mcp-remote-macos-useDocker 이미지 빌드
# Build the Docker image
docker build -t mcp-remote-macos-use .크로스 플랫폼 퍼블리싱
여러 플랫폼에 Docker 이미지를 게시하려면 docker buildx 명령을 사용할 수 있습니다. 다음 단계를 따르세요.
새로운 빌더 인스턴스를 만듭니다 (아직 만들지 않았다면):
docker buildx create --use여러 플랫폼에 대한 이미지를 빌드하고 푸시합니다 .
docker buildx build --platform linux/amd64,linux/arm64 -t buryhuang/mcp-remote-macos-use:latest --push .지정된 플랫폼에서 이미지를 사용할 수 있는지 확인하세요 .
docker buildx imagetools inspect buryhuang/mcp-remote-macos-use:latest
용법
이 서버는 MCP 도구를 통해 원격 MacOs 기능을 제공합니다.
도구 사양
이 서버는 원격 macOS 제어를 위해 다음과 같은 도구를 제공합니다.
원격_macos_get_screen
원격 macOS 컴퓨터에 연결하여 원격 데스크톱의 스크린샷을 가져옵니다. 연결 세부 정보에 환경 변수를 사용합니다.
원격_macos_키_전송
키보드 입력을 원격 macOS 컴퓨터로 전송합니다. 연결 세부 정보에 환경 변수를 사용합니다.
원격_macos_마우스_움직임
원격 macOS 컴퓨터에서 마우스 커서를 지정된 좌표로 이동합니다. 좌표는 자동으로 조정됩니다. 연결 세부 정보는 환경 변수를 사용합니다.
원격_macos_마우스_클릭
원격 macOS 컴퓨터에서 지정된 좌표에 마우스 클릭을 수행합니다. 좌표는 자동으로 조정됩니다. 연결 세부 정보는 환경 변수를 사용합니다.
리모트_맥_마우스_더블_클릭
원격 macOS 컴퓨터에서 지정된 좌표를 마우스로 두 번 클릭합니다. 좌표는 자동으로 조정됩니다. 연결 세부 정보는 환경 변수를 사용합니다.
원격_macos_마우스_스크롤
원격 macOS 컴퓨터에서 지정된 좌표로 마우스 스크롤을 수행합니다. 좌표는 자동으로 조정됩니다. 연결 세부 정보는 환경 변수를 사용합니다.
원격_macos_오픈_애플리케이션
애플리케이션을 열고 활성화하고, 추가적인 상호작용을 위해 PID를 반환합니다.
원격_macos_마우스_드래그_앤_드롭
자동 좌표 크기 조정을 사용하여 원격 macOS 컴퓨터에서 시작 지점에서 마우스 드래그 작업을 수행하고 끝 지점으로 놓습니다.
모든 도구는 연결 매개변수를 요구하는 대신 설치 중에 구성된 환경 변수를 사용합니다.
제한 사항
인증 지원 :
Apple 인증(프로토콜 30)만 지원됩니다.
보안 참고 사항
https://support.apple.com/guide/remote-desktop/encrypt-network-data-apdfe8e386b/mac https://cafbit.com/post/apple\_remote\_desktop\_quirks/
512비트 소수를 사용하는 Diffie-Hellman 키 합의 프로토콜을 사용하는 프로토콜 30만 지원합니다. 이 프로토콜은 macOS 11부터 macOS 12까지 OS X 10.11 이하 클라이언트와 통신할 때 사용됩니다.
마크다운 표로 변환된 정보는 다음과 같습니다.
원격 데스크톱을 실행하는 macOS 버전 | macOS 클라이언트 버전 | 입증 | 제어하고 관찰하다 | 항목 복사 또는 패키지 설치 | 다른 모든 작업 | 프로토콜 버전 |
맥OS 13 | 맥OS 13 | 2048비트 RSA 호스트 키 | 2048비트 RSA 호스트 키 | 인증을 위한 2048비트 RSA 호스트 키, 그 다음 128비트 AES | 2048비트 RSA 호스트 키 | 36 |
맥OS 13 | 맥OS 10.12 | 로컬 전용 보안 원격 암호(SRP) 프로토콜입니다. LDAP 또는 macOS 서버에 바인딩된 경우 Diffie-Hellman(DH)이 10.11 이하 버전입니다. | SRP 또는 DH, 128비트 AES | 인증을 위해 SRP 또는 DH를 사용한 후 128비트 AES를 사용합니다. | 2048비트 RSA 호스트 키 | 35 |
macOS 11에서 macOS 12로 | macOS 10.12에서 macOS 13으로 | 로컬 전용 SRP(Secure Remote Password) 프로토콜, LDAP에 바인딩된 경우 Diffie-Hellman | SRP 또는 DH 1024비트, 128비트 AES | 2048비트 RSA 호스트 키 macOS 13 ~ macOS 10.13 | 2048비트 RSA 호스트 키 macOS 10.13 이상 | 33 |
macOS 11에서 macOS 12로 | OS X 10.11 또는 이전 버전 | DH 1024비트 | DH 1024비트, 128비트 AES | 512비트 소수를 사용한 Diffie-Hellman 키 합의 프로토콜 | 512비트 소수를 사용한 Diffie-Hellman 키 합의 프로토콜 | 30 |
원격 MacOS 시스템에 접속할 때는 항상 안전하고 인증된 연결을 사용하세요. 이 도구는 신뢰할 수 있고 접속 권한이 있는 서버에서만 사용해야 합니다.
특허
자세한 내용은 LICENSE 파일을 참조하세요.
Available Tools
8 toolsremote_macos_get_screenB
Connect to a remote MacOs machine and get a screenshot of the remote desktop. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool connects remotely and uses environment variables, but doesn't mention behavioral traits like authentication needs, potential latency, error handling, or what the output looks like (e.g., image format). For a remote operation tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, consisting of two sentences that directly state the purpose and connection method. There's no wasted text, and it efficiently communicates key information, though it could be slightly more structured (e.g., separating prerequisites).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (remote operation) and lack of annotations and output schema, the description is moderately complete. It covers the basic action and connection method but misses details like output format, error cases, or dependencies. For a tool with no structured support, it's adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameters need documentation. The description adds context about environment variables for connection details, which compensates for the lack of parameters. This provides useful semantic information beyond the schema, earning a high score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Connect to a remote MacOs machine and get a screenshot of the remote desktop.' It specifies the verb ('get a screenshot') and resource ('remote desktop'), but doesn't explicitly distinguish it from sibling tools (e.g., mouse actions, application opening). This makes it clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal usage guidance: it mentions using environment variables for connection details, which implies prerequisites. However, it doesn't specify when to use this tool versus alternatives (e.g., other remote tools for different actions) or any exclusions. This lack of explicit context guidance limits its helpfulness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_mouse_clickC
Perform a mouse click at specified coordinates on a remote MacOs machine, with automatic coordinate scaling. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate for mouse position (in source dimensions) | |
| y | Yes | Y coordinate for mouse position (in source dimensions) | |
| source_width | No | Width of the reference screen for coordinate scaling | |
| source_height | No | Height of the reference screen for coordinate scaling | |
| button | No | Mouse button (1=left, 2=middle, 3=right) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It mentions 'automatic coordinate scaling' and 'environment variables for connection details', which adds some behavioral context, but fails to disclose critical aspects like required permissions, network dependencies, error handling, or what happens if the remote machine is unavailable. For a remote control tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and includes key behavioral details (coordinate scaling, environment variables) without unnecessary elaboration. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a remote control tool with no annotations and no output schema, the description is incomplete. It lacks information on error conditions, return values, security implications, or performance characteristics, leaving significant gaps for an AI agent to understand tool behavior fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing full parameter documentation. The description adds minimal value beyond the schema by mentioning 'automatic coordinate scaling', which relates to source_width and source_height parameters, but doesn't explain scaling mechanics or environmental variable usage. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('perform a mouse click') and target ('on a remote MacOS machine'), distinguishing it from siblings like mouse_move or mouse_drag_n_drop by specifying click behavior. However, it doesn't explicitly differentiate from mouse_double_click, which is a similar click action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'automatic coordinate scaling' and 'environment variables for connection details', which provides some context, but offers no explicit guidance on when to use this tool versus alternatives like mouse_double_click or mouse_drag_n_drop, nor any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_mouse_double_clickB
Perform a mouse double-click at specified coordinates on a remote MacOs machine, with automatic coordinate scaling. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate for mouse position (in source dimensions) | |
| y | Yes | Y coordinate for mouse position (in source dimensions) | |
| source_width | No | Width of the reference screen for coordinate scaling | |
| source_height | No | Height of the reference screen for coordinate scaling | |
| button | No | Mouse button (1=left, 2=middle, 3=right) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It mentions 'automatic coordinate scaling' and 'Uses environment variables for connection details', which are useful behavioral disclosures. However, it doesn't cover important aspects like authentication requirements, error conditions, network dependencies, or what happens if the remote machine is unavailable - significant gaps for a remote control tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences that each add value: first stating the core action with key features, second explaining implementation details. No redundant information, though it could be slightly more front-loaded with the most critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a remote control tool with 5 parameters, no annotations, and no output schema, the description provides basic context but lacks completeness. It covers the what and how of coordinate scaling but misses important operational context like error handling, performance characteristics, or what the tool returns upon execution.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing complete parameter documentation. The description adds value by explaining the overall purpose of coordinate scaling and environment variable usage, but doesn't provide additional semantic context beyond what's already in the schema descriptions. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Perform a mouse double-click'), target resource ('on a remote MacOs machine'), and key behavior ('with automatic coordinate scaling'). It distinguishes from sibling tools like 'remote_macos_mouse_click' by specifying the double-click action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through 'remote MacOs machine' and 'automatic coordinate scaling', but doesn't explicitly state when to use this tool versus alternatives like 'remote_macos_mouse_click' or 'remote_macos_mouse_drag_n_drop'. No explicit when-not-to-use guidance or prerequisites are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_mouse_drag_n_dropB
Perform a mouse drag operation from start point and drop to end point on a remote MacOs machine, with automatic coordinate scaling.
| Name | Required | Description | Default |
|---|---|---|---|
| start_x | Yes | Starting X coordinate (in source dimensions) | |
| start_y | Yes | Starting Y coordinate (in source dimensions) | |
| end_x | Yes | Ending X coordinate (in source dimensions) | |
| end_y | Yes | Ending Y coordinate (in source dimensions) | |
| source_width | No | Width of the reference screen for coordinate scaling | |
| source_height | No | Height of the reference screen for coordinate scaling | |
| button | No | Mouse button (1=left, 2=middle, 3=right) | |
| steps | No | Number of intermediate points for smooth dragging | |
| delay_ms | No | Delay between steps in milliseconds |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but lacks critical behavioral details. It doesn't disclose whether this requires specific permissions, what happens if coordinates are out of bounds, whether it's synchronous/asynchronous, or error conditions. The mention of 'automatic coordinate scaling' is helpful but insufficient for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence efficiently conveys core functionality without redundancy. Every word earns its place by specifying the operation, target environment, and key technical feature.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no annotations and no output schema, the description is inadequate. It doesn't explain what happens after the drag operation, what success/failure looks like, or important behavioral constraints. The agent lacks sufficient context to use this tool safely and effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds minimal value beyond the schema by mentioning 'automatic coordinate scaling' which relates to source_width/source_height parameters, but doesn't provide additional context about how scaling works or parameter interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Perform a mouse drag operation'), target resource ('on a remote MacOs machine'), and key capability ('with automatic coordinate scaling'). It distinguishes from siblings like mouse_click and mouse_move by specifying drag-and-drop functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing remote connection established), nor does it differentiate from similar tools like mouse_move or when drag-and-drop is appropriate versus separate click operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_mouse_moveC
Move the mouse cursor to specified coordinates on a remote MacOs machine, with automatic coordinate scaling. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate for mouse position (in source dimensions) | |
| y | Yes | Y coordinate for mouse position (in source dimensions) | |
| source_width | No | Width of the reference screen for coordinate scaling | |
| source_height | No | Height of the reference screen for coordinate scaling |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'automatic coordinate scaling' and 'environment variables for connection details', which add some behavioral context, but lacks details on permissions, error handling, or what happens if the remote machine is unavailable. For a remote control tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the core functionality and key details (coordinate scaling, environment variables). No wasted words, though it could be slightly more structured for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a remote control tool with 4 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain return values, error conditions, or important behavioral aspects like how coordinates are mapped or what 'automatic scaling' entails in practice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description adds context about 'automatic coordinate scaling', which relates to the source_width and source_height parameters, but doesn't provide additional syntax or format details beyond what the schema specifies. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Move the mouse cursor') and target ('on a remote MacOs machine'), with additional context about coordinate scaling. It distinguishes from siblings like 'mouse_click' or 'mouse_drag_n_drop' by focusing on cursor positioning, but doesn't explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like 'mouse_drag_n_drop' for dragging or 'mouse_click' for clicking after moving. It mentions environment variables for connection, but doesn't specify prerequisites or exclusions for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_mouse_scrollA
Perform a mouse scroll at specified coordinates on a remote MacOs machine, with automatic coordinate scaling. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate for mouse position (in source dimensions) | |
| y | Yes | Y coordinate for mouse position (in source dimensions) | |
| source_width | No | Width of the reference screen for coordinate scaling | |
| source_height | No | Height of the reference screen for coordinate scaling | |
| direction | No | Scroll direction | down |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions 'automatic coordinate scaling' and 'environment variables for connection details', which adds some context about how the tool works. However, it doesn't disclose critical behavioral traits like whether this requires specific permissions, potential side effects, error conditions, or what happens if the remote machine is unavailable—significant gaps for a remote control tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the core functionality. Every element earns its place: the action (mouse scroll), target (remote MacOS machine), key feature (automatic coordinate scaling), and implementation detail (environment variables). There's no wasted verbiage or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (remote control with coordinate scaling), no annotations, and no output schema, the description is somewhat complete but has gaps. It covers the what and how at a high level but lacks details about behavioral expectations, error handling, or return values. For a tool that performs remote actions, more context about reliability and failure modes would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all 5 parameters. The description adds minimal value beyond the schema—it mentions 'automatic coordinate scaling' which relates to source_width/source_height parameters, but doesn't provide additional semantic context about parameter interactions or usage nuances. This meets the baseline 3 when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('perform a mouse scroll'), target resource ('on a remote MacOS machine'), and key functionality ('with automatic coordinate scaling'). It distinguishes itself from sibling tools like mouse_click, mouse_move, and mouse_drag_n_drop by specifying the scroll action rather than click, move, or drag operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through 'remote MacOS machine' and 'environment variables for connection details', suggesting this tool is for remote control scenarios. However, it doesn't explicitly state when to use this versus alternatives like mouse_move or other mouse actions, nor does it provide exclusion criteria or prerequisites beyond the implied remote connection setup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_open_applicationC
Opens/activates an application and returns its PID for further interactions.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | REQUIRED. App name, path, or bundle ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It mentions the action ('Opens/activates') and return value ('PID for further interactions'), but lacks details on behavioral traits such as error handling (e.g., if the app isn't found), permissions required, or whether it brings the app to foreground. This leaves significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action and outcome with zero wasted words. Every part ('Opens/activates', 'returns its PID', 'for further interactions') adds value, making it appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a mutation tool (opening/activating apps) with no annotations and no output schema, the description is incomplete. It doesn't cover error cases, side effects, or the format of the PID return, which are critical for safe and effective use in a remote macOS context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the single required parameter 'identifier' with its types. The description adds no additional parameter semantics beyond what the schema provides, such as examples or format details, so it meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Opens/activates an application') and the resource ('an application'), with a specific outcome ('returns its PID for further interactions'). However, it doesn't explicitly differentiate from sibling tools like 'remote_macos_send_keys' which might also interact with applications, leaving room for minor ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing the application installed), exclusions (e.g., not suitable for system apps), or how it relates to siblings like 'remote_macos_send_keys' for keyboard input after opening.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remote_macos_send_keysC
Send keyboard input to a remote MacOs machine. Uses environment variables for connection details.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Text to send as keystrokes | |
| special_key | No | Special key to send (e.g., 'enter', 'backspace', 'tab', 'escape', etc.) | |
| key_combination | No | Key combination to send (e.g., 'ctrl+c', 'cmd+q', 'ctrl+alt+delete', etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions using environment variables for connection details, which adds some context about authentication/configuration needs. However, it lacks critical behavioral information such as error handling, latency considerations, whether it's read-only or destructive, or any rate limits—essential for a remote control tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with two sentences that directly convey the core functionality and a key implementation detail. Every word earns its place, and it is front-loaded with the primary purpose. There is no unnecessary elaboration or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a remote control tool with no annotations and no output schema, the description is incomplete. It fails to explain return values, error conditions, or behavioral traits like whether it's safe or destructive. The environment variable mention is helpful but insufficient for full contextual understanding, especially compared to richer sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear descriptions for each parameter (text, special_key, key_combination). The description adds no additional parameter semantics beyond what the schema provides, such as examples of valid inputs or interactions between parameters. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('send keyboard input') and resource ('to a remote MacOs machine'), making the purpose immediately understandable. It distinguishes itself from sibling tools like mouse operations or screen capture by focusing on keyboard input. However, it doesn't explicitly differentiate from potential keyboard-related siblings that might exist elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions environment variables for connection details, which hints at prerequisites, but offers no explicit usage context, exclusion criteria, or comparison with sibling tools like remote_macos_open_application that might involve keyboard input indirectly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- Added
remote_macos_mouse_drag_n_drop - Added
remote_macos_open_application
6 tool updates
- First observed
remote_macos_get_screen - First observed
remote_macos_mouse_click - First observed
remote_macos_mouse_double_click - First observed
remote_macos_mouse_move - First observed
remote_macos_mouse_scroll - First observed
remote_macos_send_keys
TDQS
Scored across 8 tools
Every tool has a clearly distinct purpose focused on specific remote macOS interactions: screen capture, various mouse operations (click, double-click, drag-and-drop, move, scroll), application opening, and keyboard input. There is no overlap in functionality, making tool selection unambiguous.
All tools follow a consistent 'remote_macos_verb_noun' pattern using snake_case throughout (e.g., remote_macos_mouse_click, remote_macos_send_keys). This predictable naming scheme enhances readability and usability for agents.
With 8 tools, the server is well-scoped for remote macOS control, covering essential input/output operations (mouse, keyboard, screen, apps). Each tool earns its place without redundancy, and the count is typical for this domain.
The toolset provides strong coverage for basic remote desktop control, including visual feedback, mouse/keyboard input, and app management. A minor gap exists in lacking tools for file system access or system information retrieval, but core workflows are well-supported.
Maintenance
Related MCP Connectors
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Remote MCP server for AI.TV creators — delegate account operations to your AI agent over MCP.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn open-source MCP server for macOS that lets AI control desktop apps in the background without moving the cursor or stealing focus.28Apache 2.0
- AlicenseBqualityDmaintenanceMCP server that enables AI to fully control macOS — mouse, keyboard, terminal, screenshots, window management, UI element detection, and provides AI-optimized information reporting.3623 npmMIT
- FlicenseNot gradedqualityDmaintenanceA self-hosted, VNC-backed MCP server that enables AI agents to control a dedicated macOS user session remotely, without disturbing the console user.-
- AlicenseNot gradedqualityAmaintenanceAn MCP server that gives AI agents real OS-level control of macOS, enabling them to click real buttons, type real keys, and observe rendered screens just like a human would.1MIT