Skip to main content
Glama
README.md
# llm-vision-mcp

A TypeScript MCP server that gives text-only LLMs image understanding through StepFun vision models.

This is useful when the primary model, such as GLM5.2 or DeepSeek V4, does not support image input. The model can call these MCP tools, receive text or structured visual analysis, then continue reasoning with the result.

## Tools

- `analyze_image`: general image understanding
- `extract_text_from_image`: OCR for screenshots, logs, documents, code, and UI text
- `diagnose_error_screenshot`: error screenshot and stack trace diagnosis
- `understand_technical_diagram`: architecture, flowchart, UML, ER, sequence, and network diagrams
- `analyze_data_visualization`: charts, tables, dashboards, and metrics screenshots
- `ui_to_artifact`: UI screenshot to implementation notes or design specs
- `ui_diff_check`: expected vs actual UI screenshot comparison

## Setup

```bash
npm install
cp .env.example .env
```

Set `STEPFUN_API_KEY` in the MCP client environment. Injecting env vars through the MCP client config is usually the most explicit and reliable setup.

Required:

```bash
STEPFUN_API_KEY=your_stepfun_api_key
```

Optional:

```bash
# standard | step_plan
STEPFUN_API_MODE=standard
STEPFUN_BASE_URL=https://api.stepfun.com/v1
STEPFUN_VISION_MODEL=step-1o-turbo-vision
STEPFUN_DEFAULT_DETAIL=high
STEPFUN_TIMEOUT_MS=120000
```

## Step Plan

Step Plan uses the same API key style but a different Base URL:

```bash
STEPFUN_API_MODE=step_plan
```

When `STEPFUN_API_MODE=step_plan` is set, defaults change to:

```bash
STEPFUN_BASE_URL=https://api.stepfun.com/step_plan/v1
STEPFUN_VISION_MODEL=step-3.7-flash
```

You can still override either value explicitly:

```bash
STEPFUN_API_MODE=step_plan
STEPFUN_BASE_URL=https://api.stepfun.com/step_plan/v1
STEPFUN_VISION_MODEL=step-3.7-flash
```

For backward compatibility, `STEPFUN_USE_STEP_PLAN=true` also enables Step Plan mode when `STEPFUN_API_MODE` is not set.

## Run

```bash
npm run build
npm run start
```

## Run With npx

After this package is published to npm:

```bash
npx -y llm-vision-mcp
```

The published package runs on Node.js and does not require Bun on the user's machine.

## MCP Client Config

Example:

```json
{
  "mcpServers": {
    "llm-vision-mcp": {
      "command": "node",
      "args": ["/Users/shaoyun/workdir/llm-vision-mcp/dist/index.js"],
      "env": {
        "STEPFUN_API_KEY": "your_stepfun_api_key",
        "STEPFUN_API_MODE": "step_plan",
        "STEPFUN_DEFAULT_DETAIL": "high"
      }
    }
  }
}
```

npm package example after publishing:

```json
{
  "mcpServers": {
    "llm-vision-mcp": {
      "command": "npx",
      "args": ["-y", "llm-vision-mcp"],
      "env": {
        "STEPFUN_API_KEY": "your_stepfun_api_key",
        "STEPFUN_API_MODE": "step_plan",
        "STEPFUN_DEFAULT_DETAIL": "high"
      }
    }
  }
}
```

## Image Inputs

Every single-image tool accepts:

```json
{
  "image": "/absolute/path/to/screenshot.png",
  "question": "What does this error mean?",
  "detail": "high"
}
```

The `image` field supports:

- local file path
- `file://` path
- `http://` or `https://` URL
- `data:image/...;base64,...` Data URL

`ui_diff_check` accepts two images:

```json
{
  "expected_image": "/absolute/path/to/expected.png",
  "actual_image": "/absolute/path/to/actual.png",
  "question": "Focus on layout and missing buttons.",
  "detail": "high"
}
```

## Notes

- Use `detail: "high"` for OCR, UI, diagrams, charts, and screenshots.
- Use `detail: "low"` for faster, cheaper coarse image understanding.
- StepFun supports JPG/JPEG, PNG, WebP, and static GIF image inputs.

TDQS

A3.8/5.0

Scored across 7 tools

Disambiguation5/5

Each tool targets a distinct image understanding task: general analysis, OCR, error diagnosis, diagram explanation, data visualization, UI-to-code conversion, and UI diffing. The specialized scopes prevent confusion, even though analyze_image is broad, it serves as a catch-all rather than overlapping with specific tools.

Naming Consistency4/5

Most tools follow a verb-first naming pattern (analyze_image, extract_text_from_image, diagnose_error_screenshot, understand_technical_diagram, analyze_data_visualization), but two tools (ui_to_artifact, ui_diff_check) are noun-first. Despite this minor deviation, all names are readable and use consistent snake_case.

Tool Count5/5

With 7 tools, the server is well-scoped for image understanding. Each tool covers a distinct sub-domain (general analysis, OCR, errors, diagrams, data viz, UI artifacts, UI comparison), and no tool feels redundant or unnecessary.

Completeness5/5

The tool surface is comprehensive for an image understanding server, covering general understanding, text extraction, error diagnosis, diagram interpretation, data visualization analysis, UI-to-code conversion, and UI regression checking. There are no obvious missing operations within the stated domain.

Maintenance

ActivityInactive
ResponsivenessNo issues