Skip to main content
Glama

Извлекатель изображений MCP

Сервер MCP для извлечения и преобразования изображений в base64 для анализа LLM.

Этот сервер MCP предоставляет инструменты для помощников на основе искусственного интеллекта, позволяющие:

  • Извлечение изображений из локальных файлов

  • Извлечение изображений из URL-адресов

  • Обработка изображений в кодировке base64

Как это выглядит в Курсоре:

Подходящие случаи:

  • анализ результатов теста драматурга: скриншоты

Установка

Рекомендуется: использование npx в mcp.json (самый простой способ)

Рекомендуемый способ установки этого сервера MCP — использование npx непосредственно в файле .cursor/mcp.json :

{
  "mcpServers": {
    "image-extractor": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-image-extractor"
      ]
    }
  }
}

Этот подход:

  • Автоматически устанавливает последнюю версию

  • Не требует глобальной установки

  • Надежно работает в различных средах

Альтернатива: установка по локальному пути

Если вы предпочитаете использовать локальную установку пакета, вы можете клонировать репозиторий и указать на собранные файлы:

{
  "mcpServers": {
    "image-extractor": {
      "command": "node",
      "args": ["/full/path/to/mcp-image-extractor/dist/index.js"],
      "disabled": false
    }
  }
}

Ручная установка

# Clone and install 
git clone https://github.com/ifmelate/mcp-image-extractor.git
cd mcp-image-extractor
npm install
npm run build
npm link

Это сделает команду mcp-image-extractor доступной глобально.

Затем настройте в .cursor/mcp.json :

{
  "mcpServers": {
    "image-extractor": {
      "command": "mcp-image-extractor",
      "disabled": false
    }
  }
}

Устранение неполадок для пользователей курсора : если вы видите ошибку «Не удалось создать клиент», попробуйте метод установки по локальному пути, описанный выше, или убедитесь, что вы используете правильный путь к исполняемому файлу.

Related MCP server: MCP URL Fetcher

Доступные инструменты

извлечь_изображение_из_файла

Извлекает изображение из локального файла и преобразует его в base64.

Параметры:

  • file_path (обязательно): Путь к локальному файлу изображения.

Примечание: Все изображения автоматически изменяются до оптимальных размеров (макс. 512x512) для анализа LLM, чтобы ограничить размер выходных данных base64 и оптимизировать использование контекстного окна.

извлечь_изображение_из_url

Извлекает изображение из URL-адреса и преобразует его в base64.

Параметры:

  • url (обязательно): URL-адрес изображения для извлечения

Примечание: Все изображения автоматически изменяются до оптимальных размеров (макс. 512x512) для анализа LLM, чтобы ограничить размер выходных данных base64 и оптимизировать использование контекстного окна.

извлечь_изображение_из_base64

Обрабатывает изображение в кодировке base64 для анализа LLM.

Параметры:

  • base64 (обязательно): данные изображения, закодированные в Base64

  • mime_type (необязательно, по умолчанию: "image/png"): MIME-тип изображения

Примечание: Все изображения автоматически изменяются до оптимальных размеров (макс. 512x512) для анализа LLM, чтобы ограничить размер выходных данных base64 и оптимизировать использование контекстного окна.

Пример использования

Вот пример того, как использовать инструменты от Клода:

Please extract the image from this local file: images/photo.jpg

Клод автоматически использует инструмент extract_image_from_file для загрузки и анализа содержимого изображения.

Please extract the image from this URL: https://example.com/image.jpg

Клод автоматически воспользуется инструментом extract_image_from_url для извлечения и анализа содержимого изображения.

Докер

Сборка и запуск с помощью Docker:

docker build -t mcp-image-extractor .
docker run -p 8000:8000 mcp-image-extractor

Лицензия

Массачусетский технологический институт

Available Tools

3 tools
extract_image_from_base64A

Extract and analyze images from base64-encoded data. Ideal for processing screenshots from clipboard, dynamically generated images, or images embedded in applications without requiring file system access.

ParametersJSON Schema
NameRequiredDescriptionDefault
base64YesBase64-encoded image data to analyze (useful for screenshots, images from clipboard, or dynamically generated visuals)
max_heightNoFor backward compatibility only. Default maximum height is now 512px
max_widthNoFor backward compatibility only. Default maximum width is now 512px
mime_typeNoMIME type of the image (e.g., image/png, image/jpeg)image/png
resizeNoFor backward compatibility only. Images are always automatically resized to optimal dimensions (max 512x512) for LLM analysis

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the tool 'extract[s] and analyze[s]' images, implying both extraction and analysis functions, but doesn't detail what analysis entails, potential limitations, or error handling. The description adds some context about use cases but lacks behavioral specifics like performance characteristics or output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured in two sentences: the first states the purpose, and the second provides usage context. Every sentence adds value without redundancy, making it appropriately sized and front-loaded with essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters with 100% schema coverage and no output schema, the description is moderately complete. It covers purpose and usage context but lacks details on what 'analyze' means in terms of output, which is a gap since there's no output schema to compensate. For a tool with analysis functionality, more behavioral transparency would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 5 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema, such as explaining the 'base64' parameter's format or the 'analyze' aspect mentioned in the purpose. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Extract and analyze images from base64-encoded data.' It specifies the verb (extract and analyze) and resource (images from base64 data). However, it doesn't explicitly differentiate from sibling tools like extract_image_from_file or extract_image_from_url, which handle different input sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: 'Ideal for processing screenshots from clipboard, dynamically generated images, or images embedded in applications without requiring file system access.' This gives practical scenarios, but it doesn't explicitly state when NOT to use it or directly compare it to the sibling tools that handle files or URLs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_image_from_fileA

Extract and analyze images from local file paths. Supports visual content understanding, OCR text extraction, and object recognition for screenshots, photos, diagrams, and documents.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_pathYesPath to the image file to analyze (supports screenshots, photos, diagrams, and documents in PNG, JPG, GIF, WebP formats)
max_heightNoFor backward compatibility only. Default maximum height is now 512px
max_widthNoFor backward compatibility only. Default maximum width is now 512px
resizeNoFor backward compatibility only. Images are always automatically resized to optimal dimensions (max 512x512) for LLM analysis

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses analysis capabilities (visual understanding, OCR, object recognition) and supported file types, but doesn't mention performance characteristics, rate limits, authentication needs, error conditions, or output format. It provides basic behavioral context but lacks operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. The first sentence states the core purpose and scope. The second sentence elaborates on capabilities and supported content types. Every word serves a purpose, and the description is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, 100% schema coverage, but no annotations or output schema, the description provides good purpose and usage context. However, it lacks information about what the tool returns (output format), error handling, or operational constraints. Given the absence of output schema, more detail about return values would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing complete parameter documentation. The description adds value by mentioning supported file types (PNG, JPG, GIF, WebP) and analysis capabilities, which helps contextualize the file_path parameter. However, it doesn't provide additional semantic context beyond what the schema already documents well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('extract and analyze images'), the resource ('from local file paths'), and distinguishes from siblings by specifying 'local file paths' (vs. base64 or URL sources). It lists supported analysis types (visual content understanding, OCR, object recognition) and file types, providing comprehensive purpose differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly indicates when to use this tool vs. alternatives by specifying 'from local file paths' and listing supported file types/formats. This clearly distinguishes it from sibling tools extract_image_from_base64 and extract_image_from_url, providing perfect contextual guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_image_from_urlA

Extract and analyze images from web URLs. Perfect for analyzing web screenshots, online photos, diagrams, or any image accessible via HTTP/HTTPS for visual content analysis and text extraction.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_heightNoFor backward compatibility only. Default maximum height is now 512px
max_widthNoFor backward compatibility only. Default maximum width is now 512px
resizeNoFor backward compatibility only. Images are always automatically resized to optimal dimensions (max 512x512) for LLM analysis
urlYesURL of the image to analyze for visual content, text extraction, or object recognition (supports web screenshots, photos, diagrams)

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions analysis purposes ('visual content analysis and text extraction') and that images are 'accessible via HTTP/HTTPS', but lacks details on permissions, rate limits, error handling, or output format. It adds some context but leaves significant behavioral traits unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by specific use cases. Every sentence earns its place by clarifying scope and applications without redundancy, making it efficiently structured and appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters with high schema coverage but no annotations and no output schema, the description is moderately complete. It covers the purpose and usage context well, but as a tool with potential behavioral complexities (e.g., network access, analysis output), it lacks details on permissions, errors, or result format, leaving gaps in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by implying the 'url' parameter is for 'web screenshots, photos, diagrams', but does not provide additional syntax, format, or usage details for parameters. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('extract and analyze images'), resource ('from web URLs'), and scope ('for visual content analysis and text extraction'). It distinguishes from sibling tools by specifying 'from web URLs' versus 'from_base64' or 'from_file', making the purpose unambiguous and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use this tool ('for analyzing web screenshots, online photos, diagrams, or any image accessible via HTTP/HTTPS'), but does not explicitly state when not to use it or name alternatives like the sibling tools. It implies usage scenarios without explicit exclusions or comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedextract_image_from_base64
    • First observedextract_image_from_file
    • First observedextract_image_from_url

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose based on the source of the image: base64-encoded data, local file paths, and web URLs. The descriptions reinforce this by specifying different use cases (e.g., clipboard screenshots, local files, online images), leaving no ambiguity for an agent to misselect.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with 'extract_image_from_' as a prefix, followed by the source type (base64, file, url). This predictable naming scheme makes it easy for agents to understand and navigate the tool set without confusion.

Tool Count5/5

With 3 tools, the server is well-scoped for its purpose of extracting images from different sources. Each tool earns its place by covering a distinct input method (base64, file, URL), providing a complete set for the domain without being overly sparse or bloated.

Completeness5/5

The tool surface is complete for the domain of image extraction, covering all major input sources: base64 data, local files, and web URLs. There are no obvious gaps, as these three methods encompass the typical ways images are accessed in applications, ensuring agents can handle various scenarios without dead ends.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers