Supervision Draw
Server Details
Draw detections onto an image with Roboflow supervision's own annotators, over HTTP and MCP....
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 1 tool
With only a single tool, there is no possibility of confusing it with another operation. Its purpose (annotating an image with detections) is unambiguous and clearly stated.
'annotate' is a clean, descriptive verb that fits the tool's action. A single name cannot demonstrate a repeatable convention, so consistency can't be fully credited, but nothing is inconsistent or confusing.
A one-tool server is very thin for an MCP surface, even if the tool is a broad multi-annotator operation. The count is borderline rather than clearly mismatched, since annotation is a single conceptual action.
The annotator covers the core rendering step (image in, annotated image out) but bundles everything into one monolithic call with no discovery of annotator options, no model-inference step, and no companion operations. It works for the stated purpose yet leaves notable gaps around the broader annotation lifecycle.
Available Tools
1 toolannotateAnnotate an image with detections using supervision's annotators (box, round_boxAInspect
Annotate an image with detections using supervision's annotators (box, round_box, box_corner, circle, dot, ellipse, triangle, label, color, mask, polygon, halo, background_overlay, blur, pixelate, percentage_bar), applied in order. Input: base64 image up to 4 megapixels, up to 500 detections with xyxy pixel boxes and optional class_id, confidence, ASCII label, polygon. Returns the annotated image (base64 PNG or JPEG), size and counts per class. Price: $0.003 a call.
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | Base64 of a PNG, JPEG, BMP or WebP file (a data: URI also works), up to 4 million pixels and about 2 MB | |
| format | No | Output format (default png) | |
| quality | No | JPEG quality 1-100 (default 90) | |
| annotators | No | Annotators applied in order, each a type name or an object {"type": ..., options}. Default [box, label]. Put blur, pixelate, color, mask or background_overlay before box and label. Types: box, round_box, box_corner, circle, dot, ellipse, triangle, label, color, mask, polygon, halo, background_overlay, blur, pixelate, percentage_bar. Options use supervision's names: color, color_lookup, thickness, opacity, roundness, corner_length, radius, position, start_angle, end_angle, base, height, width, kernel_size, pixel_size, outline_thickness, outline_color, border_color, border_thickness, force_box, text_color, text_scale, text_thickness, text_padding, text_position, border_radius, smart_position. color: A hex colour like "#ff8800" or a name: black, blue, green, grey, red, roboflow, white, yellow (default: supervision's class palette). color_lookup: What picks the colour: "class" (default when every detection has a class_id), "index" (position in the list) or "track" (tracker_id). position and text_position: Anchor point: center, center_left, center_right, top_center, top_left, top_right, bottom_left, bottom_center, bottom_right, center_of_mass | |
| detections | Yes | Up to 500 detections, drawn in this order | |
| class_names | No | Names indexed by class_id, used in generated labels and in counts.byClass (printable ASCII) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does well: it discloses input caps (4 megapixels, 500 detections), label constraints (ASCII), the conditional requirement that polygon be supplied on every detection when mask/polygon/halo are used, the default annotator set, the return payload (base64 PNG/JPEG, size, counts per class), and the price. It does not mention auth, rate limits, or failure behavior, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and the two sentences are information-dense with no filler. The long parenthetical annotator enumeration is bulky but is the only place a caller sees the full vocabulary, so it earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return value (annotated image as base64 PNG/JPEG, plus size and per-class counts). Combined with the input limits, ordering constraint, and pricing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description largely restates schema content (annotator type list, default [box, label], ordering rule) rather than adding new semantics, so it does not earn above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Annotate an image with detections') and enumerates the exact annotator vocabulary the tool understands, so an agent immediately knows the operation and its scope. There are no siblings, but the definition is self-distinguishing from any generic image tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete ordering rule ('applied in order'; 'Put blur, pixelate, color, mask or background_overlay before box and label') that tells the agent how to construct a valid call. There are no sibling tools, so there is no alternative to route against, but no explicit when-not-to-use guidance is offered either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- First observed
annotate
Related MCP Connectors
Roboflow computer vision for AI agents: datasets, annotation, versioning, workflows, inference.
Convert images to PNG, JPEG, WebP, or AVIF through one public remote MCP tool.
Tesseract-compatible OCR over HTTP: send an image, get its text and word boxes in Tesseract's...
Create images & video from any MCP agent — 17 models, spend limits, one URL.
Related MCP Servers
- AlicenseBqualityBmaintenanceMCP server for image annotation, supporting bounding boxes, arrows, highlights, callouts, text, and circles, plus barcode and text detection with OCR.81MIT
- AlicenseNot gradedqualityDmaintenanceFastAPI-based MCP server integrating YOLOv8 object detection and embedding services, enabling AI agents to analyze images via the Model Context Protocol.MIT
- AlicenseBqualityBmaintenanceLocal STDIO MCP server for object-detection annotation workflows. Manages datasets, categories, images, bounding boxes, and exports in JSONL format without modifying source images.23Apache 2.0
- AlicenseAqualityBmaintenanceMCP server for instruckt visual annotations. Enables AI agents to retrieve pending annotations, view screenshots, and resolve annotations.345 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.