jl-image-vison
by liugangnhm
README.md
# jl-image-vison
MCP server for image understanding — structured description of UI designs, HTML pages, and visual layouts, powered by [Agnes 2.0 Flash](https://agnes-ai.com/zh-Hans/docs/agnes-20-flash).
## Tool: `describe_image`
Analyse an image and return a richly structured breakdown — layout hierarchy, UI components, readable text, interactions, design notes, and accessibility observations. Optimised for UI mockups, Figma exports, HTML/CSS screenshots, dashboards, and app interfaces.
### Input
| Parameter | Type | Required | Description |
|---|---|---|---|
| `image` | string | yes | The image to analyse. Accepts three formats: (1) a public URL — `https://example.com/screenshot.png`; (2) a local file path — `~/Desktop/mockup.png` or `C:\Users\you\shot.png`; (3) a base64 data URL — `data:image/png;base64,...`. |
| `detail_level` | enum: `standard` \| `detailed` \| `comprehensive` | no (default `detailed`) | How thorough the description should be. |
| `focus` | string | no | Steering hint, e.g. `"layout only"`, `"extract all text"`, `"focus on accessibility"`, `"identify design system"`. |
### Output
- `content[0].text` — the full structured description (markdown).
- `structuredContent` — machine-readable fields: `summary`, `description`, `model`, `tokens`.
## Setup
```bash
npm install
npm run build
```
## Configuration
Environment variables (all optional except `AGNES_API_KEY` at runtime):
| Variable | Default | Description |
|---|---|---|
| `AGNES_API_KEY` | — | Your Agnes AI API key (required to call the tool). |
| `AGNES_BASE_URL` | `https://apihub.agnes-ai.com/v1` | Override the API base URL. |
| `AGNES_MODEL` | `agnes-2.0-flash` | Override the model name. |
Get an API key from the [Agnes AI developer console](https://agnes-ai.com).
## Connect to a host
### Claude Code / Claude Desktop (stdio)
Add to your MCP config:
```json
{
"mcpServers": {
"image-vison": {
"command": "node",
"args": ["D:/my/oss/jl_image-vison-mcp/dist/index.js"],
"env": {
"AGNES_API_KEY": "sk-your-key-here"
}
}
}
}
```
### Any MCP client via npx
```bash
AGNES_API_KEY=sk-your-key npx tsx src/index.ts
```
## Develop
```bash
npm run dev # run with tsx (no build step)
npm run typecheck # type-check only
npm run build # compile to dist/
npm run inspect # launch MCP Inspector for interactive testing
```
## Limitations (v0.2)
- **Single image** — one image per call. Multi-image comparison is not yet supported.
- **Agnes 2.0 Flash only** — first release supports only this model.
- **Local file size limit** — 10 MB per image (Agnes payload limit).
- **Claude Code pasted images** — when you paste an image into Claude Code, it may pass a file path or a data URL; both are now supported. If you get a "Cannot read file" error, the path Claude Code passed doesn't exist on the server's filesystem — upload to a public URL instead.
TDQS
A4.3/5.0
Scored across 1 tool
Disambiguation5/5
With only one tool, there is no possibility of confusion between tools. The tool's purpose is clear and distinct by default.
Naming Consistency5/5
The single tool 'describe_image' follows a clear verb_noun pattern, and consistency is trivially maintained with only one tool.
Tool Count3/5
A single tool feels thin for a server dedicated to image vision, but the tool is comprehensive and not trivial, making this borderline.
Completeness4/5
The tool provides a rich analysis covering layout, text, interactions, and accessibility, which covers the primary use case. However, other vision operations like comparison or classification are not available, leaving minor gaps.
Maintenance
ActivityStale
ResponsivenessNo issues