AIsistent
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| RDP_HOST | No | RDP server hostname/IP (for headless mode) | |
| RDP_PASS | No | RDP password | |
| RDP_USER | No | RDP username | |
| RDP_DOMAIN | No | RDP domain (optional) | |
| AISISTENT_TEMP_DIR | No | Screenshot temp directory | temp_captures/ |
| AISISTENT_YOLO_WEIGHTS | No | Path to YOLO weights file | models/weights/best.pt |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| benchmarkA | Run performance benchmark on OCR + YOLO. Args: image_base64: Optional base64-encoded PNG. If empty, generates a synthetic image. |
| capture_rdp_screenA | Capture the RDP window (or full screen as fallback). macOS: captures Microsoft Remote Desktop window via Quartz WindowID. Windows/Linux: captures primary monitor via MSS (pip install aistent[all]). |
| run_apple_ocrA | Run OCR on a base64-encoded image. macOS: Apple Vision OCR on Neural Engine (~0.05s). Other: EasyOCR on CPU or CUDA GPU (pip install aistent[all]). |
| detect_rdp_buttons | Detect UI buttons/icons via YOLOv8. Auto-selects GPU backend: CUDA (NVIDIA), MPS (Apple Silicon), or CPU. |
| inject_rdp_clickB | Inject a mouse click at percentage-based screen coordinates. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
All four tools have clearly distinct purposes: benchmark tests performance, capture grabs a screen, OCR extracts text, and inject_click interacts via mouse. No overlap in functionality.
Naming is mixed: 'benchmark' is a noun while others use verb_noun (capture_rdp_screen, run_apple_ocr, inject_rdp_click). Also, prefixes are inconsistent—some include 'rdp' while others do not.
Four tools is a reasonable number for a server focused on RDP automation and OCR. It covers the basic workflow without being overwhelming, though a few more would be welcome.
The toolset covers the core capture-OCR-click pipeline and includes a benchmark for testing. However, common automation actions like keyboard input or scrolling are missing, leaving notable gaps.