Skip to main content
Glama
jkawamoto

Florence-2 MCP Server

by jkawamoto

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
ocrB

Process an image file or URL using OCR to extract text.

captionA

Processes an image file and generates captions for the image.

processB

Processes an image file with a custom prompt using the Florence-2 model.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.4/5.0

Scored across 3 tools

Disambiguation3/5

The 'process' tool is generic and overlaps with 'caption', as captioning is a specific use case that could be handled by process. 'ocr' is distinct. The descriptions help clarify intended use, but the boundary between process and caption is not fully clear.

Naming Consistency4/5

All tools use a single lowercase verb (process, caption, ocr), which is consistent in style. However, the lack of noun objects (e.g., 'process_image' vs 'process') makes them slightly less predictable, but the pattern is uniform.

Tool Count4/5

With 3 tools, the server is on the lighter side but within a reasonable range for a focused vision model server. It covers the core capabilities without being bloated, though it could benefit from a few more specialized tools.

Completeness3/5

The tools cover generic processing, captioning, and OCR, but Florence-2 supports many other vision tasks (e.g., object detection, segmentation, grounding) that are not exposed. The generic 'process' tool mitigates some gaps, but the surface feels incomplete for the model's full potential.

Maintenance

ActivityMaintained
ResponsivenessSlow