Skip to main content
Glama

VisionFlow Match

Locate a pattern image inside a larger image even when it is scaled, rotated or

homography

Locate a pattern image inside a larger image even when it is scaled, rotated or viewed at an angle: ORB or SIFT feature matching plus cv2.findHomography with RANSAC. Returns whether it was found, the 3x3 homography (pattern pixels to image pixels), the four projected corners, the bounding box, and the inlier count as confidence. Use it to find a UI element, logo or object in a screenshot or photo. Price: $0.004 a call.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
imageYesBase64 of a PNG, JPEG, BMP or WebP file (a data: URI also works). At most 4 million pixels and about 2 MB. Alpha is dropped.
ratioNoLowe's ratio test threshold, 0.5-0.95 (default 0.75): a match is kept when its distance is below ratio times the second-best distance
detectorNoKeypoint detector and descriptor: orb (default, fast, binary descriptors, Hamming distance) or sift (slower, better with scale changes, L2 distance)
templateYesBase64 of the pattern to find, same formats. At most 1 million pixels and about 1 MB. Needs visible texture or corners; a flat-coloured pattern has no features.
min_inliersNoFewest RANSAC inliers for found to be true, 4-200 (default 10)
max_featuresNoMost keypoints kept per image, 100-5000 (default 2000)
reproj_thresholdNoRANSAC reprojection error in pixels below which a match is an inlier, 0.5-20 (default 3)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does well: it discloses the algorithm, the complete return payload (found flag, 3x3 homography, projected corners, bounding box, inlier count as confidence), and the per-call cost of $0.004. It stops short of describing failure modes (e.g. what a found=false result implies) or error conditions on oversized/invalid inputs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with what the tool does and what it returns, followed by the use case and price. No filler, no restatement of the title, and every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description compensates well by enumerating all returned values, stating the algorithm, and quoting cost. The remaining gaps are minor: no explicit note about behavior on failure to find the pattern, and no alternative-tool routing guidance for a tool with two closely named siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 and the schema does the heavy lifting for all seven parameters. The description still adds meaning beyond the schema by tying 'inlier count as confidence' to the min_inliers threshold and by framing the detector choice and RANSAC homography as the core pipeline, which helps an agent reason about the knobs rather than just fill them in.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Locate a pattern image inside a larger image') plus the technical mechanism (ORB/SIFT + cv2.findHomography with RANSAC), and the qualifier 'even when it is scaled, rotated or viewed at an angle' functionally differentiates it from a plain template matcher. However, it never names the sibling tools (feature-match, template-match), so the routing decision is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear application context with 'Use it to find a UI element, logo or object in a screenshot or photo', and the 'scaled, rotated or viewed at an angle' condition implicitly states when this is preferred over a rigid template matcher. There is no explicit when-not guidance and no mention of the sibling tools as alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources