Skip to main content
Glama

page_match_image

Locate a target image inside a webpage screenshot via pixel-accurate template matching. Returns match positions and confidence scores for captcha gaps, glyphs, and logos.

Instructions

REAL TEMPLATE MATCHING: multi-scale normalized cross-correlation of a needle against a haystack, computed IN-PAGE at full grayscale resolution (no grid loss). Returns top matches [{x, y, score, scale}] — needle-center positions in haystack-image pixels. needle_rect optionally crops the needle (e.g. an instruction glyph band from the same image — GeeTest icon-click pattern). Images are cached per URL: single-use challenge URLs fetch exactly once. THE tool for: captcha piece->gap, glyph->character, logo->page.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
hay_refYesRef of the HAYSTACK element to search in.
page_idYes
hay_rectNoOptional search-region crop of the haystack (x, y, w, h in haystack-image px). Omit for the full image. USE THIS to exclude the needle's own area (self-match guard).
needle_refYesRef of the NEEDLE element (canvas / img / background-image).
session_idYes
needle_rectNoOptional crop of the needle image (x, y, w, h in needle-image px). Omit to use the whole image. Use this to match a sub-region (e.g. an instruction glyph band inside the same image).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.7.3

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does substantial work: it discloses in-page computation, full-grayscale/no-grid-loss behavior, return format, coordinate semantics, caching behavior, and single-fetch behavior for challenge URLs. It does not explicitly assert read-only/side-effect-free status, but the matching operation itself strongly implies a non-destructive read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than minimal, but nearly every sentence carries distinct useful information: algorithm, output shape, crop behavior, caching, and use cases. The marketing-style caps ('REAL', 'THE tool') add noise, but the structure is front-loaded and information-dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no output schema and no annotations, the description is notably complete: it provides the return structure, coordinate meaning, algorithm, cache semantics, and intended use cases. It does not specify how many top matches are returned or answer edge cases like no-match, but those are minor gaps for an image-matching tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description enriches parameter meaning beyond the schema by giving a concrete needle_rect use case (instruction glyph band in a GeeTest pattern) and explaining output coordinates in haystack-image pixels. Not every parameter is expanded, but session_id/page_id are standard context IDs and the schema already documents the crop parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific operation: multi-scale normalized cross-correlation of a needle against a haystack, and specifies the output as matched needle-center positions. It also names concrete use cases (captcha piece->gap, glyph->character, logo->page), which clearly differentiates it from generic vision or OCR siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states 'THE tool for: captcha piece->gap, glyph->character, logo->page,' giving clear when-to-use context. It does not explicitly name excluded alternatives, but the use-case framing is strong enough to route an agent toward this tool for template-matching-style problems.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.