Image Analysis MCP Server
image-mcp-server
An MCP server that receives image URLs or local file paths and analyzes image content using the GPT-4o-mini model.
Features
Receives image URLs or local file paths as input and provides detailed analysis of the image content
High-precision image recognition and description using the GPT-4o-mini model
Image URL validity checking
Image loading from local files and Base64 encoding
Related MCP server: Letz AI MCP
Installation
Installing via Smithery
To install Image Analysis Server for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @champierre/image-mcp-server --client claudeManual Installation
# Clone the repository
git clone https://github.com/champierre/image-mcp-server.git # or your forked repository
cd image-mcp-server
# Install dependencies
npm install
# Compile TypeScript
npm run buildConfiguration
To use this server, you need an OpenAI API key. Set the following environment variable:
OPENAI_API_KEY=your_openai_api_keyMCP Server Configuration
To use with tools like Cline, add the following settings to your MCP server configuration file:
For Cline
Add the following to cline_mcp_settings.json:
{
"mcpServers": {
"image-analysis": {
"command": "node",
"args": ["/path/to/image-mcp-server/dist/index.js"],
"env": {
"OPENAI_API_KEY": "your_openai_api_key"
}
}
}
}For Claude Desktop App
Add the following to claude_desktop_config.json:
{
"mcpServers": {
"image-analysis": {
"command": "node",
"args": ["/path/to/image-mcp-server/dist/index.js"],
"env": {
"OPENAI_API_KEY": "your_openai_api_key"
}
}
}
}Usage
Once the MCP server is configured, the following tools become available:
analyze_image: Receives an image URL and analyzes its content.analyze_image_from_path: Receives a local file path and analyzes its content.
Usage Examples
Analyzing from URL:
Please analyze this image URL: https://example.com/image.jpgAnalyzing from local file path:
Please analyze this image: /path/to/your/image.jpgNote: Specifying Local File Paths
When using the analyze_image_from_path tool, the AI assistant (client) must specify a valid file path in the environment where this server is running.
If the server is running on WSL:
If the AI assistant has a Windows path (e.g.,
C:\...), it needs to convert it to a WSL path (e.g.,/mnt/c/...) before passing it to the tool.If the AI assistant has a WSL path, it can pass it as is.
If the server is running on Windows:
If the AI assistant has a WSL path (e.g.,
/home/user/...), it needs to convert it to a UNC path (e.g.,\\wsl$\Distro\...) before passing it to the tool.If the AI assistant has a Windows path, it can pass it as is.
Path conversion is the responsibility of the AI assistant (or its execution environment). The server will try to interpret the received path as is.
Note: Type Errors During Build
When running npm run build, you may see an error (TS7016) about missing TypeScript type definitions for the mime-types module.
src/index.ts:16:23 - error TS7016: Could not find a declaration file for module 'mime-types'. ...This is a type checking error, and since the JavaScript compilation itself succeeds, it does not affect the server's execution. If you want to resolve this error, install the type definition file as a development dependency.
npm install --save-dev @types/mime-types
# or
yarn add --dev @types/mime-typesDevelopment
# Run in development mode
npm run devLicense
MIT
Available Tools
2 toolsanalyze_imageC
Receives an image URL and analyzes the image content using GPT-4o-mini
| Name | Required | Description | Default |
|---|---|---|---|
| imageUrl | Yes | URL of the image to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the analysis uses 'GPT-4o-mini', which hints at AI-based processing, but doesn't disclose critical traits like rate limits, authentication needs, response format, error handling, or whether it's a read-only or mutating operation. For a tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function. It's appropriately sized and front-loaded with the core action. However, it could be slightly more structured by explicitly separating purpose from technical details, but overall it avoids waste and is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of an AI-based image analysis tool with no annotations and no output schema, the description is incomplete. It doesn't explain what the analysis returns (e.g., text description, labels, confidence scores), any limitations (e.g., image size, content restrictions), or error conditions. For a tool that likely produces rich output, this lack of context makes it inadequate for an agent to use effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the single parameter 'imageUrl' documented as 'URL of the image to analyze'. The description adds no additional meaning beyond this, such as URL format requirements, supported image types, or size limits. Since the schema does the heavy lifting, the baseline score of 3 is appropriate, but no extra credit is earned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'analyzes the image content using GPT-4o-mini' with a specific verb ('analyzes') and resource ('image content'). It distinguishes from the sibling tool 'analyze_image_from_path' by specifying it works with URLs rather than file paths. However, it doesn't explicitly mention what kind of analysis is performed (e.g., object detection, description generation, etc.), keeping it at a 4 rather than a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention the sibling tool 'analyze_image_from_path' or explain the difference between URL-based and path-based image analysis. There's no context about prerequisites, limitations, or appropriate use cases, leaving the agent with minimal usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_image_from_pathA
Loads an image from a local file path and analyzes its content using GPT-4o-mini. AI assistants need to provide a valid path for the server execution environment (e.g., Linux path if the server is running on WSL).
| Name | Required | Description | Default |
|---|---|---|---|
| imagePath | Yes | Local file path of the image to analyze (must be accessible from the server execution environment) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that the tool 'analyzes its content using GPT-4o-mini,' which adds useful context about the analysis method. However, it doesn't disclose important behavioral traits like whether this is a read-only operation, potential rate limits, authentication needs, or what kind of analysis results to expect. The description doesn't contradict annotations (since none exist), but it leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized at two sentences. The first sentence states the core purpose, and the second adds important usage guidance about path requirements. There's no wasted text, and information is front-loaded. It could be slightly more concise by combining ideas, but it's efficient overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no annotations, no output schema, and a simple single parameter, the description provides adequate basic information about what the tool does and path requirements. However, for an analysis tool, it should ideally mention what kind of analysis is performed (e.g., object detection, captioning, etc.) or what format the results take. The description is complete enough for basic understanding but lacks depth about the analysis output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single parameter 'imagePath' with its description. The description adds some context about path requirements ('valid path for the server execution environment') but doesn't provide additional semantic meaning beyond what's in the schema. With high schema coverage, the baseline is 3 even without extra parameter information in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Loads an image from a local file path and analyzes its content using GPT-4o-mini.' This specifies both the action (loads and analyzes) and the resource (image from local file path). However, it doesn't explicitly differentiate from its sibling 'analyze_image' tool, which likely handles images differently (e.g., from URLs or base64).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool: when you have a local file path accessible from the server execution environment. It mentions the need for 'valid path for the server execution environment (e.g., Linux path if the server is running on WSL),' which helps guide usage. However, it doesn't explicitly state when NOT to use it or mention the sibling tool as an alternative for different image sources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools are nearly identical in purpose—both analyze image content using GPT-4o-mini—differing only in input source (URL vs. local path). This creates high ambiguity, as an agent might easily misselect between them based on minor context cues, and they essentially duplicate functionality with no distinct operational domains.
Tool names follow a perfectly consistent verb_noun pattern with clear, descriptive suffixes ('analyze_image' and 'analyze_image_from_path'). Both use snake_case and maintain a uniform naming style, making them predictable and easy to interpret within the set.
With only 2 tools, the server feels under-scoped for an 'Image Analysis' domain, as it lacks basic operations like object detection, text extraction, or image comparison. The count is too low to support comprehensive analysis workflows, limiting agent capabilities to a single, narrowly defined task.
The toolset is severely incomplete for image analysis, missing essential functions such as image preprocessing, metadata extraction, or batch processing. It offers only one core action (analysis) in two input variants, leaving obvious gaps that will hinder agents from performing varied or advanced tasks in this domain.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Holiday photo MCP server: list and fetch personal holiday photos inline in Claude chat.
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
Generate and edit images and create short videos inside Claude. Prepaid credits, no subscription.
MCP server for NanoBanana AI image generation and editing
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA powerful server that integrates the Moondream vision model to enable advanced image analysis, including captioning, object detection, and visual question answering, through the Model Context Protocol, compatible with AI assistants like Claude and Cline.19Apache 2.0

Letz AI MCPofficial
FlicenseBqualityDmaintenanceA Model Context Protocol server that enables Claude to generate and upscale images through the Letz AI API, allowing users to create images directly within Claude conversations.22- FlicenseBqualityDmaintenanceAI-powered image generation and editing server for Claude Code, supporting multiple Gemini models, batch generation, shared context, and conversion to PDF/PPTX.51
- AlicenseNot gradedqualityDmaintenanceA powerful MCP server that brings AI vision capabilities to Claude Desktop. Analyze images and videos using OpenAI GPT-4o, Claude, or any compatible vision API.22MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/champierre/image-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server