Skip to main content
Glama

mcp-vision-image-fadhli

MCP server untuk analisis gambar lewat 9router — dengan penemuan otomatis model vision yang benar-benar bekerja di mesin Anda.

Tidak perlu daftar model hardcoded. Tidak perlu API key Google. Tidak ada quota free tier.


Kenapa perlu penemuan otomatis?

9router mengekspos ratusan model, dan katalognya punya flag capabilities.vision. Flag itu tidak bisa dipercaya.

Diukur pada satu mesin (24 Sep 2026):

Model di katalog

824

Mengklaim vision: true

468

Benar-benar bisa memproses gambar

segelintir

Sampel 50 model yang mengklaim vision: hanya 4 yang menjawab benar. Sisanya gagal karena:

  • 401/403 — API key provider mati

  • 402 — kredit habis

  • 5xx — error upstream (dibungkus 9router jadi 503)

  • 200 dengan content kosong — paling menyesatkan: HTTP sukses, tapi tidak ada jawaban

  • timeout

Yang penting: penyebabnya hampir selalu provider, bukan model. Dan karena setiap instalasi 9router punya provider & akun yang berbeda, daftar model yang bekerja tidak bisa di-hardcode — harus ditemukan di mesin tempat server berjalan.


Related MCP server: vision-mcp

Cara kerja

analyze_image
   │
   ├─ 1. Coba model eksplisit / env ROUTER9_MODEL / cache / default
   │     └─ berhasil? selesai. (kasus umum, cepat)
   │
   └─ 2. Semua gagal → pindai mesin ini
         ├─ sapu SATU model per provider (gelombang 1)
         ├─ provider hidup? perdalam di gelombang berikutnya
         ├─ uji pakai gambar 3 kotak, nilai jawabannya
         └─ simpan hasil → pakai model tercepat

Mengapa menyapu provider dulu: kegagalan itu soal provider. Menguji 5 model dari provider yang sama itu sia-sia — kalau providernya mati, kelimanya mati. Gelombang 1 mengambil satu model per provider, jadi berapa pun jumlah providernya, sekali jalan langsung ketahuan mana yang hidup.

Model dinilai secara objektif. Gambar uji berisi 3 kotak (merah, hijau, biru). Model yang benar-benar bisa melihat akan menyebut ketiga warna; model yang mengarang tidak akan cocok.


Instalasi

npm install -g mcp-vision-image-fadhli

Butuh 9router berjalan di mesin yang sama (default http://127.0.0.1:20128).

Konfigurasi MCP

Tambahkan ke ~/.claude.json (atau config MCP klien Anda):

{
  "mcpServers": {
    "mcp-vision-image": {
      "type": "stdio",
      "command": "node",
      "args": ["C:\\path\\ke\\mcp-vision-image\\src\\server.js"],
      "env": {
        "ROUTER9_API_KEY": "sk-xxxxxxxxxxxx"
      }
    }
  }
}

API key diambil dari dashboard 9router → Endpoint & Key.

Cuma satu key yang dibutuhkan. Key 9router membuka seluruh pool model Anda — tidak perlu key Antigravity, B.AI, atau provider lain satu per satu.

Variabel lingkungan

Variabel

Default

Keterangan

ROUTER9_API_KEY

—

Wajib. API key 9router.

VISION_PROVIDERS

ag,oc

Provider yang dipakai, dipisah koma. Isi * untuk semua.

ROUTER9_BASE_URL

http://127.0.0.1:20128

Alamat 9router.

ROUTER9_MODEL

—

Paksa satu model, lewati pemilihan otomatis.

ROUTER9_MAX_TOKENS

2000

Batas token jawaban.

ROUTER9_TIMEOUT_MS

120000

Timeout per model.

Provider yang dipakai

Default ag,oc — Antigravity dan OpenCode Free. Keduanya dipakai karena terukur, bukan dipilih sembarangan:

Provider

Model

Verifikasi

ag (Antigravity)

gemini-3.8-flash-medium, -high, 3.7-flash-high

✅ 4,1 dtk

oc (OpenCode Free)

mimo-v2.5-free, muse-spark-1.3-contributor-free, 1.2-contributor-free

✅ 3,8–10 dtk

Semua diuji 25 Sep 2026 lewat endpoint Anthropic 9router memakai gambar uji (3 kotak merah/hijau/biru) — keempatnya menjawab dengan benar.

Kenapa dibatasi: memindai semua provider menemukan 4 model bekerja dari 30 diuji; dibatasi ke ag,oc menemukan 8 dari 8. Provider lain menghabiskan waktu dan kuota untuk model yang kreditnya kosong atau key-nya mati.

Ganti lewat env kalau instalasi Anda berbeda:

"env": {
  "ROUTER9_API_KEY": "sk-xxxxxxxxxxxx",
  "VISION_PROVIDERS": "ag"          // Antigravity saja
}

Provider yang diminta tapi tidak ada di katalog dilaporkan sebagai peringatan — tidak diabaikan diam-diam.

Catatan: kenapa endpoint Anthropic

Modul ini memanggil /v1/messages, bukan /v1/chat/completions. Dua alasan, keduanya terukur:

  1. Input gambar di endpoint OpenAI mengembalikan content kosong — HTTP-nya 200, tapi jawabannya tidak ada, untuk semua model yang diuji.

  2. Provider oc kena gate tambahan di endpoint OpenAI. OpenCode Free hanya menjawab kalau tool quartet {bash, glob, grep, read} ikut dikirim; tanpanya, HTTP 200 dengan content kosong. Endpoint Anthropic tidak menerapkan gate itu, jadi oc bekerja tanpa trik tambahan.

Bentuk responsnya juga tidak seragam — Antigravity menjawab SSE, oc menjawab JSON. Parser menangani keduanya.


Tools

analyze_image

Menganalisis gambar dari file lokal atau URL.

Parameter

Keterangan

image_path

Path file lokal

image_url

URL gambar (diunduh otomatis)

prompt

Instruksi spesifik (opsional)

model

Paksa model tertentu (opsional)

Jawabannya menyertakan model mana yang dipakai dan berapa lama.

discover_vision_models

Memindai 9router untuk menemukan model yang benar-benar bisa memproses gambar di mesin ini. Hasilnya di-cache dan dipakai otomatis oleh analyze_image.

Jalankan ulang kalau daftar terasa basi (provider berganti, kredit berubah).

Parameter

Default

Keterangan

target

5

Berhenti setelah sekian model bekerja

max_probe

48

Batas atas model diuji (pengaman kuota)

include_non_vision

false

Ikut uji yang tidak mengklaim vision — untuk memeriksa akurasi metadata

Batas selalu dilaporkan. Kalau ada model yang tidak diuji karena kena batas, jumlahnya disebutkan — supaya hasilnya tidak terbaca "menyeluruh" padahal tidak.

get_usage_stats

Statistik pemakaian per model, riwayat harian, status backend, dan ringkasan hasil pemindaian terakhir.


Catatan teknis

Kenapa endpoint Anthropic, bukan OpenAI?

Untuk input gambar, /v1/chat/completions di 9router mengembalikan HTTP 200 dengan content kosong — untuk semua model yang diuji. /v1/messages dengan model yang sama menjawab benar. Jadi ini soal jalur translasi gambar di 9router, bukan soal model.

Endpoint

Hasil untuk input gambar

/v1/chat/completions

HTTP 200, content kosong

/v1/messages

jawaban benar

Blok thinking dibuang

Model thinking mengirim rantai penalaran di thinking_delta (sering dalam bahasa Mandarin). Parser hanya mengambil text_delta, dan max_tokens dijaga cukup tinggi supaya jatah tidak habis di thinking lalu menyisakan content kosong.


Test

ROUTER9_API_KEY=xxx npm test

13 pemeriksaan, termasuk memanggil tool lewat protokol MCP sungguhan (stdio) dan memverifikasi jawaban terhadap gambar uji secara objektif.

Untuk memindai lebih luas di luar test:

node test/probe-vision.js --per-provider 2    # sampel stratifikasi
node test/probe-vision.js --all               # sapu bersih
node test/probe-vision.js --only ag/gemini-3.8-flash-high
node test/probe-vision.js --control           # uji akurasi metadata katalog

Lisensi

MIT

Available Tools

2 tools
analyze_imageA

Menganalisis dan mendeskripsikan gambar (file path lokal atau URL) menggunakan Gemini Vision AI.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel Gemini yang digunakan. Default: "gemini-3.7-flash". Bisa di-override via env GEMINI_MODEL.
promptNoPertanyaan atau instruksi spesifik seputar gambar (misal: "Bacakan teks di gambar ini", "Apakah ada kucing?"). Default: deskripsi lengkap.
image_urlNoURL gambar online yang di-copy dari browser atau clipboard (misal: "https://example.com/foto.jpg" atau URL berakhiran .png/.jpg/.webp/.gif). Server akan mengunduh gambar dari URL tersebut.
image_pathNoPath file gambar lokal di sistem (misal: "C:\path\to\image.png" atau "./foto.jpg")

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the tool uses Gemini Vision AI and accepts local paths or URLs, but it does not disclose the return format, network dependency, file size limits, or privacy implications. An agent needs more context to anticipate side effects and output shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no unnecessary words. Every element—verb, resource, input types, and model provider—is useful and helps the agent quickly understand and select the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a moderately complex image analysis tool with no annotations and no output schema, the description covers the core capability and input forms but omits output format, supported file types, size limits, and network behavior. It is adequate for basic invocation but not fully complete for an agent that needs to anticipate results and constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for all four parameters, so the schema already explains the model, prompt, image_url, and image_path in detail. The description's mention of 'file path lokal atau URL' loosely maps to the image_path and image_url parameters, but it adds no additional semantic meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Menganalisis dan mendeskripsikan') and identifies the resource ('gambar' / image), explicitly naming the supported input types (local file path or URL) and the underlying model (Gemini Vision AI). This clearly distinguishes it from the sibling tool get_usage_stats, which has a completely different purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for image analysis and description by naming the input types, but it does not provide explicit when-to-use or when-not-to-use guidance, nor does it reference the sibling tool as an alternative. The sibling is distinct enough that confusion is unlikely, but the description itself offers no usage rules beyond the obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_usage_statsA

Menampilkan statistik pemakaian API Gemini Vision (jumlah panggilan hari ini, total, sisa limit harian, per model, dan riwayat harian).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The verb 'Menampilkan' (displays) indicates a read-only operation, and the description clearly lists the types of data returned. It does not mention side effects, auth requirements, or rate limits, but for a simple stats display tool, this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the main purpose and then lists specific statistics. It is concise with no unnecessary words, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters and no output schema, the description fully specifies what the tool does and what data it returns. The sibling tool is clearly different, and the description is complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, so the schema is non-informative. The description adds no parameter details because there are none to add. Baseline for 0 parameters is 4, and the description does not need to compensate for any missing schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: displaying usage statistics for the Gemini Vision API, listing specific metrics such as today's calls, total calls, remaining daily limit, per-model breakdown, and daily history. This distinguishes it from the sibling tool analyze_image, which is for image analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied by the description: use when needing API usage stats. It does not explicitly exclude other tools or provide alternatives, but the clear focus on usage statistics makes the use case obvious. No explicit 'when not to use' is given, but it is not necessary given the simplicity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.3.0
    • First observedanalyze_image
    • First observedget_usage_stats

TDQS

A4/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: analyze_image handles image analysis, while get_usage_stats reports API usage metrics. No overlap or ambiguity exists.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern: analyze_image and get_usage_stats. The naming is uniform and predictable.

Tool Count3/5

With only 2 tools, the server feels thin but is reasonable for a focused single-purpose image analysis service. It sits at the borderline of being too minimal.

Completeness5/5

For its stated purpose of image analysis via Gemini Vision, the server covers the core operation (analyze_image) and adds useful monitoring (get_usage_stats). No obvious missing operations within this narrow domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers