Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
GHOSTFOX_HOMEYesPath to the unpacked Ghostfox engine directory, e.g. /opt/ghostfox.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
page_match_imageA

REAL TEMPLATE MATCHING: multi-scale normalized cross-correlation of a needle against a haystack, computed IN-PAGE at full grayscale resolution (no grid loss). Returns top matches [{x, y, score, scale}] — needle-center positions in haystack-image pixels. needle_rect optionally crops the needle (e.g. an instruction glyph band from the same image — GeeTest icon-click pattern). Images are cached per URL: single-use challenge URLs fetch exactly once. THE tool for: captcha piece->gap, glyph->character, logo->page.

page_clickA

Click an element by CSS selector. For form controls (buttons, inputs), uses a JS click; for links and other elements, dispatches real mouse events at coordinates. Prefer page_click_ref when you have a page_a11y ref — it scrolls into view first and is more reliable on web-component UIs. Returns 'ok' on success.

page_snapshotA

Extract the visible text content of a page as plain text (token-friendly). Returns: url, title, and content (all visible text, no HTML). For semantic element data with refs and values, use page_a11y instead — it gives you interactive elements with roles and names. Use this when you just need to READ page content without needing to interact with elements.

page_consoleA

DEBUG CORTEX: read the page's console output (log/warning/error) captured at the PROTOCOL level — the page cannot hide or patch it. Invisible debugger for AI agents: reproduce the bug, read what the page logged. Optionally filter noise by clearing after read. Returns JSON [{kind, text, url, line, ts}]. Auto-starts capture on first call.

page_wait_forA

Wait until a CSS selector becomes visible on the page (replaces manual sleeps). Returns true if found, false on timeout.

page_commentA

Submit a comment on the current page. Automatically: 1) detects your username via session_me, 2) checks if you already commented (skips if duplicate), 3) finds the comment editor (Reply or Join conversation), 4) opens it, 5) types your text, 6) clicks submit, 7) verifies. Returns result JSON.

page_typeA

Type text character-by-character into an element by CSS selector (human-like key events). Prefer page_type_ref when you have a page_a11y ref — it handles rich editors (Lexical/Draft/ProseMirror) and returns a verified receipt. Use this only when you only have a CSS selector and don't need rich editor support. Returns 'ok' on success.

page_read_refA

Read the FULL value of an element by its page_a11y ref — no truncation. Use when the snapshot's 200-char preview isn't enough (body text, long input fields).

page_network_startA

PROTOCOL-LEVEL NETWORK CAPTURE: start recording every HTTP response for this page, BELOW the page (invisible to page JS, unpatchable). The answer data of any captcha/API travels here. Start before triggering the flow you want to see.

page_geetest_slideA

GEETEST SLIDE SOLVER: solve the GeeTest v3-style slide puzzle from the live page, then perform the human drag itself. Extracts the three canvases (bg, puzzle slice, full reference bg) via toDataURL, finds the hole with |bg-fullbg| diff + largest-blob + morphological closing (the JPEG-noise trap), measures the piece's solid-alpha left edge, and drags the slider by hole_x0 - piece_x0 with the engine's humanized mouse. Returns JSON {drag_x, hole: [x0, x1], piece_x0}. Call AFTER the challenge popup is open. No vision model, pure pixel math.

page_hcaptchaA

HCAPTCHA SOLVER: solve hCaptcha image challenges from the live page with the host vision model (Cloudflare Workers AI GLM). Handles ALL challenge variants — tile grids, pattern-break icon fields, any image-pick — by asking the model for the exact pixel coordinates of every element to click, then clicking with the humanized mouse. Multi-round: loops until the page's h-captcha-response field receives a token (up to max_rounds, default 4). Credentials come from CLOUDFLARE_API_KEY / CLOUDFLARE_ACCOUNT_ID env or the params. Returns JSON {success, rounds, response_len}.

captcha_solveA

Solve a CAPTCHA through the configured provider (env GHOSTFOX_CAPTCHA_PROVIDER=2captcha + GHOSTFOX_CAPTCHA_KEY). Turnstile/hcaptcha: pass sitekey + pageurl; image captchas: pass image_base64. Stealth-first: prefer not being challenged at all.

page_captcha_ocrA

CAPTCHA OCR: classify the text of a normal image captcha (the distorted-text family) with the ddddocr model (CRNN+LSTM trained specifically on captcha text, 8210-char charset). 100% local — runs on onnxruntime inside the runtime. Takes the ref of the captcha element (or the whole viewport) and returns the recognized text. Feed it into the answer field and submit. Far stronger than generic OCR on captcha fonts.

page_ocrA

TIER 2 VISION — LOCAL OCR: extract text from an element (or the whole viewport if no ref). 100% local (pure-Rust ML models auto-download once to GHOSTFOX_HOME/models). Answers 'what text is written there' — for image captchas, canvas text, scanned UI. Models download on first call (~12MB, once).

page_type_refA

Type text into the element a page_a11y ref points at (inputs and rich editors). Returns landed chars as a receipt.

session_meA

Detect the currently logged-in username on the page (Reddit, X, GitHub, HN). Returns username or 'unknown'. Use before commenting to avoid duplicates.

page_network_bodyC

Fetch the BODY of a captured response by requestId — Network.getResponseBody BY PROTOCOL. The JSONP/API answer data that page JS can't see, served from the engine. THIS is the 0s and 1s layer.

page_dismiss_modalA

Detect and dismiss common modals: cookie banners, consent dialogs, popups, overlays. Returns what was dismissed.

page_pressA

Press a named key (Enter, Tab, Escape, ArrowDown, ...) — e.g. Enter to submit a search box.

page_evalA

Evaluate a JavaScript expression in the page's main frame and return its JSON value. Read-only introspection is safest; treat results of mutations with care.

identity_generateA

Generate a new coherent browser identity, returned as TOML.

session_createA

Create a new browsing session: launches the engine with a fresh coherent identity. Returns session_id.

page_openA

Navigate to a URL in an existing session. Returns a page_id (string) that must be passed to all subsequent page tools. Waits for the page to load. If the page has iframes or shadow DOM, use page_a11y instead of guessing CSS selectors. Example: page_open(session_id, 'https://example.com') returns a page_id like 'abc123'.

page_pixelsA

SUPERMAN GLASSES: render an element (canvas / img / background-image) as a compact luminance GRID of digits 0-9 the agent READS directly — see shapes, holes, object orientation, image layout WITHOUT a vision model. 0=black, 9=white. Captcha gaps appear as darker cells, upright skies are bright rows on top. Pass the ref from page_a11y. Returns JSON {w, h, grid:[rows of digits]}.

page_a11yA

Semantic snapshot of the page: every visible interactive element with a stable ref, role (button/link/textbox/...), accessible name and CURRENT value — pierces shadow DOM, so web-component UIs (Reddit, modern frameworks) are fully visible. Use this instead of guessing CSS selectors.

page_network_readA

List captured HTTP responses [{url, requestId}] since capture start. Optionally filter by URL substring (e.g. 'geetest' or 'api.'). Pair with page_network_body to read the response content BY PROTOCOL.

page_geetest_clickA

GEETEST ICON-CLICK SOLVER: solve the GeeTest v3 word-click captcha (characters on a photo, click in the strip's order) from the live page. Fetches the challenge image off the element's background, runs the trained ONNX pair (YOLOv8s char detection + siamese order matching, 100% local CPU, ~0.5s), and returns click targets in page coordinates plus the raw boxes. Call AFTER the challenge popup is open. Then click each 'click' point via page_drag and press the geetest confirm button. Returns JSON {clicks: [{x, y}, ...] in page px, boxes_raw: [[x, y], ...] in image px, rect: {...}, img: {w, h}}.

confirm_actionA

SAFETY GATE: confirm a dangerous action before executing it. Call this BEFORE clicking refs on pages flagged with danger_zone (financial/medical/legal/auth).

session_evidenceA

READ-ONLY: Retrieve the complete audit trail for a session. No side effects — does not modify the session or browser. Returns JSON with: session_id, dir (evidence directory path), identity_toml (path to the identity file used), event_count (total tool calls recorded), events (array of all events with timestamps), and snapshots (list of page snapshot file paths). The session_id must come from a previous session_create call. For a nonexistent session, returns an error. Evidence persists after the session ends and includes: append-only events.jsonl, page snapshots (markdown), screenshots (PNG), and identity.toml. Use session_create first if you need a session.

page_init_scriptA

INIT SCRIPTS: run your JS at DOCUMENT START on every navigation — BEFORE any page script loads. The deepest interception layer in Ghostfox: hooks installed here are what page libraries (gt.js, analytics, frameworks) capture, so their stored references are YOURS. Use for: API response interception (JSONP callback swaps), state capture, anti-detection instrumentation, environment patches. Applies to subsequent navigations.

identity_auditA

Audit an identity TOML for coherence violations (contradictory signals a detector would flag).

page_fillA

Set an input's value directly (form fill). Works where key-event typing hits engine bugs; fires input/change events like real edits.

page_screenshotA

Capture a PNG screenshot of a page (viewport by default, full page with full_page=true). Saved under the session recordings dir; returns the file path. Feeds the live view when enabled.

page_contrastA

HIGH-PASS VISION: local-contrast grid of an element's image (|gray - gaussian_blur|). Makes ANYTHING blended into a background VISIBLE — captcha characters on photos, watermarks, hidden strokes. 0 = flat area, 9 = strong edge. This is the native tool born from solving GeeTest icon-click with pure math. Works on canvas / img / background-image (one fetch per challenge).

session_pagesA

POPUP/TAB VISION: list EVERY browser target — pages you opened AND popups the site opened (OAuth windows, payment flows). Popups are auto-attached and given a page_id usable with every page_* tool immediately. Returns JSON [{page_id, target_id, url}].

page_move_toA

Move the mouse onto an element (hover) along a HUMAN-LIKE path: bezier arc, ease-in-out velocity, sub-pixel tremor, overshoot+correction. Triggers hover menus and tooltips. Pass the ref from page_a11y.

page_captcha_rotateA

ROTATE CAPTCHA SOLVER: brute-force sweep + human replay. Rotate challenges ask to turn an image upright; standard demos rotate 15deg per button click and answer with a check button. The tool first sweeps every angle programmatically (instant JS clicks, ~30s), reads the visible feedback after each check (success keywords: pass/通过/正确/成功/verif/succeeded), then REPLAYS the winning rotation with real human mouse drags and the final check. Returns JSON {clicks, angle, status}. For fixed-image demos the sweep converges every time.

page_dragA

Drag with a HUMAN-LIKE movement profile: approaches the source, presses, drags along a bezier arc with ease-in-out velocity and micro-pauses, settles, releases. Pass from_ref + to_ref to drag element onto element, or from_ref + offset_x/offset_y to drag by pixels (slider captchas, resize handles). All from page_a11y refs.

page_upload_fileA

Upload a file to an input[type=file] by CSS selector. The file must exist on the machine running the engine.

page_click_refA

Click an element by its ref from page_a11y. Scrolls it into view first. No selectors needed.

page_errorsA

DEBUG CORTEX: read the page's UNCAUGHT JS EXCEPTIONS with stack traces, captured at the protocol level. The error console a developer opens in DevTools — as a tool. Returns JSON [{text, url, line, stack: [frames], ts}]. Auto-starts capture on first call.

page_visionA

TIER 4 VISION — VISUAL CORTEX: detect word bounding boxes in an element (or viewport) via the built-in text-detection ML model. Returns JSON [{x, y, w, h}, ...] in image pixels. The first shipped visual-cortex model — see where text lives in ANY image (click-order captchas, canvas text, scanned layout). 100% local.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.4/5.0

Scored across 42 tools

Disambiguation3/5

Some overlap exists among the many vision, input, and captcha tools (e.g., page_snapshot vs page_a11y vs page_ocr, or page_type vs page_type_ref vs page_fill), though the descriptions do a good job of explaining when to prefer one over another. The highly specialized captcha solvers are each tied to a specific challenge type, but the sheer number of adjacent tools still creates selection ambiguity.

Naming Consistency4/5

The vast majority of tools follow a clear page_<verb_or_noun> snake_case convention, with session_ and identity_ prefixes for lifecycle and identity concerns. Minor deviations like captcha_solve, confirm_action, and the noun-style page_console/page_a11y keep it from being perfectly uniform, but the overall pattern is predictable.

Tool Count2/5

At 42 tools, this is well above the 25+ threshold and feels too heavy even for a broad stealth-browser and captcha-solving purpose. Many vision and captcha variants could likely be consolidated or surfaced through a smaller capability-oriented set. The count is not chaotic enough for a 1, but it will tax an agent's ability to choose efficiently.

Completeness4/5

The toolset covers the core lifecycle thoroughly: session creation, identity generation/audit, page navigation, interaction, reading, network capture, console/error introspection, captcha solving, and evidence retrieval. Minor gaps like explicit session teardown, page reload/back-forward navigation, or cookie management are missing, but agents can work around those with page_eval and navigation to a URL.

Maintenance

ActivityMaintained
ResponsivenessNo issues