SlimSnap MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| get_latest_captureA | Returns the newest SlimSnap capture as structured JSON: every OCR'd text element with its normalised bounding box and colour, plus any annotations the user drew (arrows, boxes, highlights, callouts) and which element each one points at. Use this whenever the user refers to their screen, their last screenshot, 'this', 'here', or something they just marked. Returns no image. The text and coordinates answer most questions about a capture on their own; the pixels are available separately when a question is genuinely visual. |
| list_capturesA | Lists the user's recent captures newest first, with a short text preview of each. This is the cheap summary view; get_capture returns a single capture in full. Useful for finding an earlier capture when the user does not mean the newest one, or for working through several marked screens. |
| get_captureA | Returns one specific capture as structured JSON, same shape as get_latest_capture. Get ids from list_captures. |
| get_capture_imageA | Returns the actual pixels of one frame, already downscaled to a size the model can read. Useful for visual questions that the text and coordinates cannot settle, such as spacing, visual style, colour balance, or imagery with no text in it. Costs substantially more tokens than the structured form of the same capture. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool has a clearly distinct role: list_captures gives an overview, get_latest_capture and get_capture return structured data for different selections, and get_capture_image provides the raw pixels. The descriptions reinforce these boundaries, so an agent should not confuse them.
All tool names follow a consistent verb_noun pattern: list_captures, get_capture, get_latest_capture, and get_capture_image. The variation in pluralization and the additional 'latest' qualifier are natural and predictable.
Four tools is a well-scoped set for a capture retrieval server. There are no redundant tools and each one contributes a distinct function, covering both listing and accessing individual captures.
The tool surface fully covers the core lifecycle of reading captures: listing available captures, retrieving the latest capture, retrieving a specific capture by ID, and fetching the raw image when needed. There are no obvious missing operations for the stated purpose.