mcp-see
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GEMINI_API_KEY | No | API key from Google AI Studio. Enables all tools. | |
| OPENAI_API_KEY | No | OpenAI API key for GPT-4o vision. Alternative provider for describe and describe_region. | |
| ANTHROPIC_API_KEY | No | Anthropic API key for Claude vision. Alternative provider for describe and describe_region. | |
| GOOGLE_CLOUD_PROJECT | No | GCP project ID for Vertex AI instead of Gemini API. Requires ADC setup. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| describeC | Get an AI-generated description of an image. Supports multiple providers (Gemini, OpenAI, Claude). |
| detectA | Detect objects in an image and return bounding boxes. Uses Gemini for native bounding box support. Coordinates are normalized 0-1000 as [ymin, xmin, ymax, xmax]. |
| describe_regionA | Crop an image to a bounding box and describe that region in detail. Use this after detect() to zoom in on specific objects. |
| analyze_colorsA | Extract dominant colors from an image region using K-Means clustering in LAB color space. Returns colors sorted by frequency with human-readable names from color.pizza. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a distinct image analysis operation: full-image description, object detection, region-level description, and color extraction. The only slight overlap is between describe and describe_region, but the latter's description clarifies its use after detect().
All names are snake_case and start with a verb, but the pattern is mixed: describe and detect are bare verbs while describe_region and analyze_colors follow verb_noun. This minor inconsistency is still readable.
Four tools is a well-scoped set for an image analysis server, with each tool fulfilling a distinct role. The count is neither bloated nor thin.
The surface covers core image understanding: description, object detection, region zoom, and color analysis. However, notable gaps exist such as OCR/text extraction or image classification, which some users might expect from an 'mcp-see' server.