Skip to main content
Glama
README.md
# MCP HydroCoder Vision

[English Installation](./INSTALL_.md) | [中文安装](./INSTALL_CN.md) | [中文 README](./README_CN.md)

A vision-language MCP server that enables Claude Code to analyze images using **Qwen3 VL 4B** model running locally via LM Studio.

## Features

- 🔍 **Image Analysis** - Describe images in detail
- 📝 **Text Extraction (OCR)** - Extract text from images in multiple languages
- 💻 **UI to Code** - Generate HTML/CSS/JS code from UI/design screenshots
- 🏠 **100% Local** - All processing happens on your machine, no cloud API needed
- ⚡ **Fast** - Qwen3 VL 4B runs efficiently on 8GB VRAM

## Prerequisites

1. **LM Studio** installed and running
2. **Qwen3 VL 4B** model loaded in LM Studio
3. **Node.js 18+**

## Installation

### 1. Clone the repository

```bash
git clone https://github.com/hydroCoderClaud/mcp-hydrocoder-vision.git
cd mcp-hydrocoder-vision
```

### 2. Install dependencies

```bash
npm install
```

### 3. Build the project

```bash
npm run build
```

## Configuration

### 1. Start LM Studio

1. Open LM Studio
2. Download and load `Qwen3-VL-4B-Instruct` model
3. Start the local server (default: `http://localhost:1234`)

### 2. Configure Claude Code

Add to your `~/.claude/settings.json`:

```json
{
  "mcpServers": {
    "hydrocoder-vision": {
      "command": "npx",
      "args": ["-y", "mcp-hydrocoder-vision"],
      "env": {
        "LM_STUDIO_URL": "http://localhost:1234/v1/chat/completions",
        "VISION_MODEL": "Qwen3-VL-4B-Instruct"
      }
    }
  }
}
```

## Usage

### Available Tools

#### `analyzeImage`

Analyze an image and get a detailed description.

```
/analyzeImage imagePath: "C:/path/to/image.png" prompt: "What's in this image?"
```

#### `extractText`

Extract text from an image (OCR).

```
/extractText imagePath: "C:/path/to/document.png" language: "English"
```

#### `describeForCode`

Generate code from a UI/design screenshot.

```
/describeForCode imagePath: "C:/path/to/design.png" framework: "Vue"
```

## Environment Variables

| Variable | Default | Description |
|----------|---------|-------------|
| `LM_STUDIO_URL` | `http://localhost:1234/v1/chat/completions` | LM Studio API endpoint |
| `VISION_MODEL` | `Qwen3-VL-4B-Instruct` | Model name to use |

## Development

```bash
# Run in development mode (watch mode)
npm run dev

# Build for production
npm run build

# Start the built server
npm start
```

## Troubleshooting

### "Request failed: ECONNREFUSED"

- Make sure LM Studio is running
- Check that the local server is enabled
- Verify the `LM_STUDIO_URL` is correct

### "No response from model"

- Ensure Qwen3 VL 4B model is loaded in LM Studio
- Check LM Studio logs for errors
- Try a simpler prompt first

### Image not found

- Use absolute paths
- Ensure the file exists and is accessible
- Check file permissions

## License

MIT

TDQS

B3.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: analyzeImage provides general image description, describeForCode generates code from UI images, and extractText performs OCR. There is no overlap in functionality, making tool selection straightforward for an agent.

Naming Consistency4/5

The tools follow a consistent verb-based naming pattern (analyze, describe, extract) with clear objects (Image, ForCode, Text). While describeForCode uses a prepositional phrase, it remains readable and maintains a logical structure across the set.

Tool Count4/5

Three tools is a minimal but reasonable count for a vision-focused server. It covers core image analysis tasks without being overly sparse, though additional tools like image editing or format conversion could enhance completeness.

Completeness3/5

The tools cover key vision tasks (description, code generation, OCR), but there are notable gaps such as image manipulation, format conversion, or batch processing. The surface is functional but not fully comprehensive for a vision domain.

Maintenance

ActivityInactive
ResponsivenessNo issues