Skip to main content
Glama

act_on_page

Automates browser tasks in PageBolt from a URL and plain-English goal: run an observe-plan-act-verify loop until the outcome is reached, then return actions and success status.

Instructions

Give PageBolt a URL and a plain-English GOAL; it runs an observe→plan→act→verify loop server-side until the goal is met, then returns a structured trace of every action it took plus a success/failure status. This is the "hands" on top of observe_page (the "eyes") — you do NOT author selectors or a step list yourself. Use act_on_page when you only know the OUTCOME you want (e.g. "log in and open billing", "accept the cookie banner and start a trial"); use run_sequence when you already know the exact deterministic steps/selectors (cheaper). Available on Starter+ plans. Cost is metered: 2 requests base + 1 per step taken. SECURITY: page text is treated as untrusted — the agent pursues only your goal and ignores instructions embedded in the page. Scope allowedDomains tightly and avoid destructive flows.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesRequired. The page to start on.
goalYesRequired. Plain-English description of the outcome you want (e.g. "Log in and go to the billing page").
maxStepsNoCap on planning iterations (default 8). Clamped to your plan ceiling (Starter 10, Growth 15, Scale 20).
session_idNoRun inside an existing persistent session (Starter+; create with create_session) to reuse cookies/login. Otherwise an ephemeral browser is used and discarded.
credentialsNoLogin credentials. The agent references them as {{username}}/{{password}} and they appear in the returned trace as <redacted>.
allowedDomainsNoHosts the agent may navigate to (e.g. ["app.example.com"]). Defaults to the start URL host only; navigation elsewhere is rejected.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.17.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: server-side execution model, returned structured trace plus success/failure status, plan gating (Starter+), metered cost model (2 base + 1 per step), session reuse vs ephemeral browser, credential redaction, and a security posture (untrusted page text, prompt-injection resistance, tighten allowedDomains, avoid destructive flows). This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core mechanism, then routing, then cost/plan, then security. Despite being dense, every sentence carries distinct actionable information (what it does, when to prefer it, pricing, safety). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 6-param, nested-object, no-output-schema tool, the description covers everything an agent needs: execution model, return shape (trace + status), cost, plan eligibility, session semantics, and security constraints. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters including the credentials sub-object and its redaction behavior. The description reinforces the goal concept with examples and advises scoping allowedDomains tightly, but adds little parameter syntax or meaning beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Give PageBolt a URL and a plain-English GOAL') and precisely delineates its role — the 'hands' on top of observe_page (the 'eyes') — with a clear description of the observe→plan→act→verify loop. It is immediately distinguishable from siblings like observe_page and run_sequence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative ('use run_sequence when you already know the exact deterministic steps/selectors (cheaper)') and gives the selecting condition ('when you only know the OUTCOME you want'), backed by two concrete goal examples. This is exactly the when-to-use vs alternative guidance the dimension asks for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.