low-hallucination-vision
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | Yes | The model identifier (e.g., mimo-vl-2.5) | |
| VISION_API_KEY | Yes | API key for the vision model | |
| VISION_API_BASE | Yes | The base URL for the OpenAI-compatible API endpoint | |
| VISION_TEMPERATURE | No | Temperature for model responses (lower = less hallucination) | 0.2 |
| VISION_CONFIDENCE_THRESHOLD | No | Confidence threshold below which claims are flagged as dubious | 0.6 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageA | Analyze an image with anti-hallucination safeguards. |
| ocr_extractA | Extract visible text from an image (OCR only, no scene description). |
| detect_elementsA | Detect objects in an image with mandatory bounding boxes. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The analyze_image tool includes modes for OCR and detection, directly overlapping with ocr_extract and detect_elements. While descriptions reference the dedicated tools, the redundancy creates potential confusion about which tool to choose. The general-purpose nature of analyze_image vs. the specialized tools provides some clarity, but boundaries are not crisp.
Two tools follow the verb_noun pattern (analyze_image, detect_elements), while ocr_extract inverts the order. All are snake_case and descriptive, so the inconsistency is minor and does not impede readability.
Three tools is a well-scoped count for a focused vision server, each targeting a distinct primary task: general analysis, OCR, and object detection. This falls squarely within the ideal 3-15 range.
The set covers core vision workflows: general scene description, text extraction, and object detection. Minor gaps exist (e.g., no dedicated UI screenshot tool despite analyze_image's mode), but the surface is functional and sufficient for typical use cases.