Skip to main content
Glama
jkawamoto

Florence-2 MCP Server

by jkawamoto

Florence-2 MCP Server

uv Python Application pre-commit GitHub License

An MCP server for processing images using Florence-2.

You can process images or PDF files stored on a local or web server to extract text using OCR (Optical Character Recognition) or generate descriptive captions summarizing the content of the images.

Installation

Claude

Download the latest MCP bundle mcp-florence2.mcpb from the Releases page, then open the downloaded .mcpb file or drag it into the Claude Desktop's Settings window.

You can also manually configure this server for Claude Desktop. Edit the claude_desktop_config.json file by adding the following entry under mcpServers:

{
  "mcpServers": {
    "florence-2": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/jkawamoto/mcp-florence2",
        "mcp-florence2"
      ]
    }
  }
}

After editing, restart the application.

For more information, see: Connect to local MCP servers - Model Context Protocol.

goose

Open this link

goose://extension?cmd=uvx&arg=--from&arg=git%2Bhttps%3A%2F%2Fgithub.com%2Fjkawamoto%2Fmcp-florence2&arg=mcp-florence2&id=florence2&name=Florence-2&description=An%20MCP%20server%20for%20processing%20images%20using%20Florence-2

to launch the installer, then click "Yes" to confirm the installation.

You can also directly edit the config file (~/.config/goose/config.yaml) to include the following entry:

extensions:
  florence2:
    name: Florence-2
    cmd: uvx
    args: [ --from, git+https://github.com/jkawamoto/mcp-florence2, mcp-florence2 ]
    enabled: true
    type: stdio

For more details on configuring MCP servers in Goose, refer to the documentation: Using Extensions | goose.

LM Studio

To configure this server for LM Studio, click the button below.

Add MCP Server florence-2 to LM Studio

Related MCP server: 🪄 ImageSorcery MCP

Tools

ocr

Process an image file or URL using OCR to extract text.

Arguments:

  • src: A file path or URL to the image file that needs to be processed.

caption

Processes an image file and generates captions for the image.

Arguments:

  • src: A file path or URL to the image file that needs to be processed.

process

Processes an image file with a custom prompt using the Florence-2 model.

Arguments:

  • src: A file path or URL to the image file that needs to be processed.

  • prompt: A custom prompt for the Florence-2 model.

License

This application is licensed under the MIT License. See the LICENSE file for more details.

Available Tools

3 tools
captionA

Processes an image file and generates captions for the image.

ParametersJSON Schema
NameRequiredDescriptionDefault
srcYesA file path or URL to the image file that needs to be processed.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description discloses basic behavior (process image, generate caption) but lacks details on performance, file size limits, or output format. With no annotations, it carries full burden and does adequately for a simple tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, clear sentence with no redundancy. Every word contributes meaning, and it is front-loaded with the core function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, output schema exists), the description is complete enough for an agent to understand purpose. Slightly more context on output format would elevate to 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers parameter 'src' with description 'A file path or URL to the image file'. The tool description adds no additional meaning beyond what the schema provides. Schema coverage is 100%, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Processes' with a clear resource 'image file' and outcome 'generates captions'. It effectively distinguishes from the sibling tool 'ocr' which likely extracts text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool vs alternatives. Usage is implied from context (generating captions vs OCR), but no exclusions or prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ocrB

Process an image file or URL using OCR to extract text.

ParametersJSON Schema
NameRequiredDescriptionDefault
srcYesA file path or URL to the image file that needs to be processed.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided. The description does not disclose side effects, required permissions, rate limits, or any behavioral traits beyond the basic action. It is too minimal for a tool without annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, front-loaded, with no unnecessary words. It efficiently conveys the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity and presence of an output schema (not detailed here), the description covers the essential input and action. It could mention output format, but overall adequate for a straightforward OCR tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the lone parameter, which already describes 'src' as a file path or URL. The tool description adds no extra meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (OCR) and the resource (image file or URL) and the output (extract text). It distinguishes from sibling tool 'caption' which likely does a different task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives. The description does not mention prerequisites, when not to use, or compare with sibling 'caption'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

processB

Processes an image file with a custom prompt using the Florence-2 model.

ParametersJSON Schema
NameRequiredDescriptionDefault
srcYesA file path or URL to the image file that needs to be processed.
promptYesA custom prompt for the Florence-2 model.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must carry the full behavioral burden. It only says the tool 'processes' the image and names the model, but does not disclose side effects, failure modes, or notable runtime behavior. The term 'processes' is opaque and adds little transparency beyond the obvious operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to identifying what the tool does and with which model.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with a full input schema and an output schema, this is minimally sufficient. The main missing piece is explicit usage guidance relative to the sibling tools, but the description does state the core action, target file, and model.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters adequately. The description adds the model context ('Florence-2') and clarifies that the prompt is custom, but it does not significantly expand on the parameter meaning already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Processes an image file') and a resource plus tool ('custom prompt using the Florence-2 model'). It is not a tautology and gives enough context to understand the tool's basic role, though 'processes' is somewhat generic. The mention of a custom prompt weakly distinguishes it from the sibling tools 'caption' and 'ocr'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'custom prompt' implies this tool is for flexible, user-defined vision tasks rather than the specialized sibling operations 'caption' and 'ocr'. However, the description never explicitly states when to prefer this tool over those alternatives, nor does it give exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.5.0
    • Addedprocess
  2. 2 tool updatesv0.3.14
    • Changedcaption1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "result": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Result",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "result"
        +  ],
        +  "title": "captionOutput",
        +  "type": "object"
        +}
    • Changedocr1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "properties": {
        +    "result": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Result",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "result"
        +  ],
        +  "title": "ocrOutput",
        +  "type": "object"
        +}
  3. 2 tool updates
    • First observedcaption
    • First observedocr

TDQS

B3.4/5.0

Scored across 3 tools

Disambiguation3/5

The 'process' tool is generic and overlaps with 'caption', as captioning is a specific use case that could be handled by process. 'ocr' is distinct. The descriptions help clarify intended use, but the boundary between process and caption is not fully clear.

Naming Consistency4/5

All tools use a single lowercase verb (process, caption, ocr), which is consistent in style. However, the lack of noun objects (e.g., 'process_image' vs 'process') makes them slightly less predictable, but the pattern is uniform.

Tool Count4/5

With 3 tools, the server is on the lighter side but within a reasonable range for a focused vision model server. It covers the core capabilities without being bloated, though it could benefit from a few more specialized tools.

Completeness3/5

The tools cover generic processing, captioning, and OCR, but Florence-2 supports many other vision tasks (e.g., object detection, segmentation, grounding) that are not exposed. The generic 'process' tool mitigates some gaps, but the surface feels incomplete for the model's full potential.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers