Skip to main content
Glama

page_geetest_click

Solves GeeTest v3 word-click captchas from the live page by extracting the challenge image, running local character detection and order matching, and returning page-coordinate click targets for automation.

Instructions

GEETEST ICON-CLICK SOLVER: solve the GeeTest v3 word-click captcha (characters on a photo, click in the strip's order) from the live page. Fetches the challenge image off the element's background, runs the trained ONNX pair (YOLOv8s char detection + siamese order matching, 100% local CPU, ~0.5s), and returns click targets in page coordinates plus the raw boxes. Call AFTER the challenge popup is open. Then click each 'click' point via page_drag and press the geetest confirm button. Returns JSON {clicks: [{x, y}, ...] in page px, boxes_raw: [[x, y], ...] in image px, rect: {...}, img: {w, h}}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
refYesRef of the .geetest_item_wrap element (carries the challenge background image).
page_idYes
session_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.7.3

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses that the tool runs a local ONNX model ('100% local CPU, ~0.5s'), fetches the image from the element background, and only returns coordinates rather than clicking itself ('Then click each... via page_drag'). It does not explicitly state that it never submits or confirms the captcha, which costs it full marks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but mostly earns its length: it front-loads the purpose, then adds meaningful operation details, prerequisites, follow-up actions, and the return JSON structure. Slight redundancy exists between the all-caps title and the first clause, and the model-architecture detail could be trimmed without losing operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is unusually complete: it states the prerequisite, the processing model, the page-flow steps, and the exact JSON response shape including coordinate spaces. It still leaves session_id/page_id semantics unexplained and does not cover failure cases, but nothing essential to invoking it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, and the description does nothing to compensate: it never mentions session_id or page_id, which are both required. Only 'ref' is documented in the schema itself, so an agent must guess what the other two identifiers mean or how they relate to the live page.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'GEETEST ICON-CLICK SOLVER' and then states exactly what it does: 'solve the GeeTest v3 word-click captcha... from the live page.' It names a specific captcha variant, a specific element, and a specific output (click targets), making it clearly distinguishable from siblings like page_geetest_slide or page_captcha_rotate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear sequencing: 'Call AFTER the challenge popup is open' and then 'click each point via page_drag and press the geetest confirm button.' This tells the agent when to invoke it and what to do next, though it does not explicitly contrast it with alternative captcha tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.