Peggle AI MCP Server
# Peggle AI MCP Server
This MCP server allows an AI to play Peggle by capturing the screen and controlling the mouse.
## Tools
- `capture_screen`: Captures the primary monitor and returns the image as base64.
- `click_at(x, y)`: Moves the mouse to (x, y) and performs a left click.
## How to use with LM Studio
1. **Build the server**:
```bash
npm install && npm run build
```
2. **Configure LM Studio**:
- Open LM Studio and go to the **MCP** tab.
- Click **Add Server**.
- Set the command to `node` (ensure node is in your PATH).
- Set the arguments to the absolute path of the built `index.js`, for example:
`C:\Users\{youruser}\peggle-ai-mcp\dist\index.js`
- If needed, you can use the mcp.json already included here and paste in the json contents to your mcp.json for it to work
- Alternatively, use `npx`:
- Command: `npx`
- Arguments: `-y C:\Users\maxwe\Desktop\peggle-ai-mcp`
3. **Start Playing**:
- Open Peggle (make sure it's on your primary monitor).
- In LM Studio, select a model that supports vision and tools (e.g., Ministral 3B, or any other vision-capable model).
- Ask the AI: "Take a screenshot of Peggle, analyze where the best shot is, and click there."
## Implementation Details
- **Screen Capture**: Uses `screenshot-desktop`.
- **Mouse Control**: Uses PowerShell commands via `child_process` for cross-platform compatibility on Windows without needing native build tools.
- **Protocol**: Model Context Protocol (MCP).
## Note on Windows 11
Ensure that PowerShell execution policy allows running the commands if you encounter issues. The server uses standard PowerShell calls that should work in most default configurations
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: one captures visual data from the screen, while the other performs a mouse interaction at specific coordinates. There is no overlap in functionality, making it impossible to confuse them.
Both tools follow a consistent verb_noun pattern with snake_case naming: capture_screen and click_at. The naming is clear, predictable, and adheres to the same convention throughout.
With only two tools, the server feels thin for a general-purpose automation or interaction domain like screen capture and mouse control. Key operations such as keyboard input, drag-and-drop, or region-specific screenshots are missing, limiting the scope significantly.
For a server implied to handle desktop automation or interaction (based on screen capture and mouse click), there are significant gaps. Missing tools include keyboard input, mouse movement without clicking, drag operations, and more advanced screen interactions, which will likely cause agent failures in broader automation tasks.