Skip to main content
Glama
1orZero

z_ai_vision_mcp_server_clone

by 1orZero
README.md
# z_ai_vision_mcp_server_clone

OpenAI-compatible MCP server for running image analysis tools against your own vision model endpoint.

## Tools

- `ui_to_artifact`
- `extract_text_from_screenshot`
- `diagnose_error_screenshot`
- `understand_technical_diagram`
- `analyze_data_visualization`
- `ui_diff_check`
- `analyze_image`

## Configuration

Set either `VISION_ENDPOINT` or `VISION_BASE_URL`.

| Variable | Required | Description |
| --- | --- | --- |
| `VISION_ENDPOINT` | Yes, unless `VISION_BASE_URL` is set | Full chat completions endpoint. |
| `VISION_BASE_URL` | Yes, unless `VISION_ENDPOINT` is set | Base URL; `/chat/completions` is appended. |
| `VISION_MODEL` | Yes | Vision model name sent in the request body. |
| `VISION_API_KEY` | No | Bearer token. Omit for local endpoints that do not require auth. |
| `VISION_PROVIDER` | No | Label for your provider. Defaults to `custom`. |
| `VISION_MAX_IMAGE_MB` | No | Local image size limit. Defaults to `5`. |
| `VISION_TIMEOUT_MS` | No | Request timeout. Defaults to `300000`. |
| `VISION_TEMPERATURE` | No | Optional model temperature. |
| `VISION_TOP_P` | No | Optional model top_p. |
| `VISION_MAX_TOKENS` | No | Optional max_tokens. |

You can also place these values in a local `.env` file in the working directory where the server starts. Real environment variables override `.env` values.

## Run

```bash
npm install
npm run build
VISION_ENDPOINT=http://localhost:11434/v1/chat/completions VISION_MODEL=llava npm start
```

Or with `.env`:

```bash
npm start
```

## MCP Client Example

```json
{
  "mcpServers": {
    "z-ai-vision-clone": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "z_ai_vision_mcp_server_clone"],
      "env": {
        "VISION_ENDPOINT": "https://your-provider.com/v1/chat/completions",
        "VISION_MODEL": "your-vision-model",
        "VISION_API_KEY": "your-api-key"
      }
    }
  }
}
```

TDQS

B3/5.0

Scored across 7 tools

Disambiguation4/5

Each tool has a distinct purpose (e.g., data visualization vs. error screenshot vs. text extraction), but 'analyze_image' as a catch-all creates slight ambiguity since it overlaps with the specialized tools.

Naming Consistency4/5

Most names follow a verb_noun pattern (analyze_*, extract_text_*, understand_*), but 'ui_diff_check' and 'ui_to_artifact' deviate slightly from the typical structure.

Tool Count5/5

With 7 tools focused on image and UI analysis, the count is well-scoped—no unnecessary duplication and full coverage of the domain.

Completeness5/5

The tool set covers major visual analysis tasks (charts, screenshots, text extraction, UI comparison, diagram understanding) with no obvious gaps for the intended use case.

Maintenance

ActivityStale
ResponsivenessNo issues