mcp-vision-image-fadhli
Allows analyzing images using Google's Gemini API, supporting local file paths or URLs (PNG/JPG/WebP/GIF/BMP/SVG) and retrieving usage statistics.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-vision-image-fadhliCan you analyze this image and describe what's in it? https://example.com/cat.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-vision-image-fadhli
MCP server untuk analisis gambar (image vision) menggunakan Gemini API — berjalan via stdio.
Fitur
analyze_image— analisis gambar dari path lokal atau URL online (PNG/JPG/WebP/GIF/BMP/SVG)get_usage_stats— statistik pemakaian API (harian, per model, sisa limit)
Related MCP server: vision-mcp
Instalasi & Penggunaan
Jalankan langsung via npx (tanpa install):
GEMINI_API_KEY=xxx npx -y mcp-vision-image-fadhliSetup di Claude Code
claude mcp add mcp-vision-image -s user -t stdio -e GEMINI_API_KEY=xxx -- npx -y mcp-vision-image-fadhliAtau otomatis via claudecode-setup — wizard akan mendaftarkan server ini beserta MCP lain.
Catatan:
GEMINI_API_KEYwajib di-set (env atau parameterapiKey). Dapatkan key gratis di https://aistudio.google.com/apikey
Penggunaan di Claude Code
Setelah terdaftar, minta bantuan dengan menyebut file/URL gambar:
analisis gambar ini: C:\path\ke\foto.pngDevelopment
npm install
GEMINI_API_KEY=xxx npm test # smoke test
npm start # jalankan sebagai stdio serverLisensi
MIT
Available Tools
2 toolsanalyze_imageA
Menganalisis dan mendeskripsikan gambar (file path lokal atau URL) menggunakan Gemini Vision AI.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model Gemini yang digunakan. Default: "gemini-3.7-flash". Bisa di-override via env GEMINI_MODEL. | |
| prompt | No | Pertanyaan atau instruksi spesifik seputar gambar (misal: "Bacakan teks di gambar ini", "Apakah ada kucing?"). Default: deskripsi lengkap. | |
| image_url | No | URL gambar online yang di-copy dari browser atau clipboard (misal: "https://example.com/foto.jpg" atau URL berakhiran .png/.jpg/.webp/.gif). Server akan mengunduh gambar dari URL tersebut. | |
| image_path | No | Path file gambar lokal di sistem (misal: "C:\path\to\image.png" atau "./foto.jpg") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the tool uses Gemini Vision AI and accepts local paths or URLs, but it does not disclose the return format, network dependency, file size limits, or privacy implications. An agent needs more context to anticipate side effects and output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no unnecessary words. Every element—verb, resource, input types, and model provider—is useful and helps the agent quickly understand and select the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderately complex image analysis tool with no annotations and no output schema, the description covers the core capability and input forms but omits output format, supported file types, size limits, and network behavior. It is adequate for basic invocation but not fully complete for an agent that needs to anticipate results and constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for all four parameters, so the schema already explains the model, prompt, image_url, and image_path in detail. The description's mention of 'file path lokal atau URL' loosely maps to the image_path and image_url parameters, but it adds no additional semantic meaning beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Menganalisis dan mendeskripsikan') and identifies the resource ('gambar' / image), explicitly naming the supported input types (local file path or URL) and the underlying model (Gemini Vision AI). This clearly distinguishes it from the sibling tool get_usage_stats, which has a completely different purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for image analysis and description by naming the input types, but it does not provide explicit when-to-use or when-not-to-use guidance, nor does it reference the sibling tool as an alternative. The sibling is distinct enough that confusion is unlikely, but the description itself offers no usage rules beyond the obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usage_statsA
Menampilkan statistik pemakaian API Gemini Vision (jumlah panggilan hari ini, total, sisa limit harian, per model, dan riwayat harian).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The verb 'Menampilkan' (displays) indicates a read-only operation, and the description clearly lists the types of data returned. It does not mention side effects, auth requirements, or rate limits, but for a simple stats display tool, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the main purpose and then lists specific statistics. It is concise with no unnecessary words, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and no output schema, the description fully specifies what the tool does and what data it returns. The sibling tool is clearly different, and the description is complete for its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, so the schema is non-informative. The description adds no parameter details because there are none to add. Baseline for 0 parameters is 4, and the description does not need to compensate for any missing schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: displaying usage statistics for the Gemini Vision API, listing specific metrics such as today's calls, total calls, remaining daily limit, per-model breakdown, and daily history. This distinguishes it from the sibling tool analyze_image, which is for image analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by the description: use when needing API usage stats. It does not explicitly exclude other tools or provide alternatives, but the clear focus on usage statistics makes the use case obvious. No explicit 'when not to use' is given, but it is not necessary given the simplicity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.3.0- First observed
analyze_image - First observed
get_usage_stats
TDQS
The two tools have clearly distinct purposes: analyze_image handles image analysis, while get_usage_stats reports API usage metrics. No overlap or ambiguity exists.
Both tools follow a consistent verb_noun pattern: analyze_image and get_usage_stats. The naming is uniform and predictable.
With only 2 tools, the server feels thin but is reasonable for a focused single-purpose image analysis service. It sits at the borderline of being too minimal.
For its stated purpose of image analysis via Gemini Vision, the server covers the core operation (analyze_image) and adds useful monitoring (get_usage_stats). No obvious missing operations within this narrow domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
MCP server for NanoBanana AI image generation and editing
Related MCP Servers
- AlicenseBqualityCmaintenanceAn MCP server for image generation using the Gemini API.1382MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that analyzes images using OpenRouter's Gemini Flash model, supporting local file paths and URLs.32MIT
- FlicenseAqualityDmaintenanceMCP server for generating images using Gemini, supporting multiple aspect ratios and both AI Studio and Vertex AI backends.1-
- AlicenseAqualityCmaintenanceMCP server for generating and editing images using Google Gemini API. Supports text-to-image generation, image editing, and image description.320MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/FadhliRajwaaRahmana/mcp-vision-image'
If you have feedback or need assistance with the MCP directory API, please join our Discord server