Skip to main content
Glama

Do the task on the live websites and return the result

run
Destructive

Executes the task on the real websites (the search, the price check, the availability lookup, the configurator, the booking flow) and returns what came back. Runs a script you authored against the get_library vocabulary, on the live sites, and returns { ok, result, logs, error, ms }. Call get_library FIRST — it gives the exact function names, argument shapes, and return types; this description is the LANGUAGE + how-to (get_library is just the vocabulary).

THE LANGUAGE — plain async JavaScript: • bowmark is a ready global (no import). Call capabilities off it — await bowmark.<capability>.<method>(...) — always await, they're async. • Individual sites are callable too, at await bowmark.providers.<provider>.<fn>(...). Use one when you specifically want THAT site; otherwise prefer the capability, which fans out across sites and routes around failures. • Real control flow: await, if, loops, array methods (map/filter/sort/slice), and Promise.all for fan-out. • return a value to get it back (JSON-serialized). log(...) for progress lines. • Standard JavaScript built-ins are there (JSON, Math, Date, RegExp, Intl, Promise), plus URL and URLSearchParams — use them to resolve a relative link against the page it came from and to build query strings. Nothing else from the Web platform exists: no fetch, setTimeout, TextEncoder or crypto. • bowmark is the ONLY I/O — no fetch, process, filesystem, or import/require. Write a plain async body, not a wrapping function. • Keep scripts small and deterministic — no infinite loops. Runs in a hard sandbox with CPU + memory + wall-clock limits.

SENDING IT: pass the script text as run({ script })script is the only argument (there is no site argument; the library exposes every capability under bowmark). result is whatever you returned; logs are your log() lines in order; on a throw/timeout ok:false and error is set.

CHECK status BEFORE ok. It is ok | error | partial | needs_user. • partial means the script RAN and result is real and usable, but some of what it called never answered — so the result is narrower than what you asked for. ok is still true; this is not a failure. incomplete.summary says what happened in one sentence, incomplete.failures names each call that threw and what the site said, and incomplete.degraded names each call that answered while reporting its OWN results thin. You MUST say so when you present the result: name what was missed, and do not describe it as complete, exhaustive, or 'all' of anything. A partial you report as whole is a wrong answer, not a slightly smaller right one. • Before you conclude a partial is final, check incomplete.failures[].fixable. fixable: true means YOUR ARGUMENT was rejected, not the site — the error text names what that function actually takes, so re-read it in get_library, fix the argument and run again; that recovers the whole answer. For any other failure re-running usually returns the same thing. • needs_user means a site needs the USER signed in — it is NOT a failure and NOT something you can fix by editing the script. needs lists the sites; meta.handoff.url is a single-use link that expires (meta.handoff.expiresAt). Give the user that URL, say which sites it covers, and WAIT. When they tell you they're done, send the SAME script again unchanged. Do NOT retry before then — it will stop at the same place and cost another run. Do NOT try to log in yourself, ask them for a password, or work around it with a different site. • Logged-in runs need a Bowmark API key on the connection; if you get needs_user saying so, tell the user to add one rather than retrying.

trace is the execution trace — every capability you called and the providers it fanned out to under the hood: [{ kind:'capability', capability:'flights', method:'search', ms }, { kind:'provider', capability:'flights', provider:'google_flights', fn:'search', results, status, ms }, …]. The script never visits websites — it calls capabilities that route to providers, and the trace is the receipt.

Composition is the point — call a method MULTIPLE times and combine results. To sweep a date range, call the search per date inside Promise.all and sort/filter the merged array (each flight result carries its date, so you can tell the runs apart). See the get_library examples for the exact shape.

SOME capabilities return their rows alongside a warnings array — { flights, warnings }, { hotels, warnings }, { cars, warnings }. Others return a bare array. The signature in get_library tells you which; go by it rather than assuming. Where there IS a warnings array it names any site dropped from the fan-out, and the rows themselves look identical with or without it. Read it, and pass on anything it says rather than quoting a 'cheapest' that only ranks the sites that happened to answer. Dropping warnings from what you return does not hide it — the run comes back status: 'partial' regardless, because the runtime counts what your script CALLED, not what it chose to report.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
scriptYesThe JavaScript script body to execute (async, against the `bowmark` global). e.g. `const { flights, warnings } = await bowmark.flights.search({from:'SFO',to:'JFK',depart:'2026-09-01'}); return { best: flights.sort((a,b)=>a.price-b.price)[0], warnings };`

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
msNo
okYes
logsNo
metaNo
errorNo
needsNo
runIdNo
resultNo
statusNo
incompleteNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing the `status` values (`ok`, `error`, `partial`, `needs_user`), explaining `incomplete.failures[].fixable`, describing `trace` as the execution receipt, and warning about the `warnings` array in capability results. It also clearly states sandbox limits: no `fetch`, `process`, filesystem, `import`, or other Web platform APIs. These disclosures align with the `openWorldHint`/`destructiveHint` annotations rather than contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is behaviorally complex and the length is justified. It is organized with headings, bullet lists, inline code examples, and a clear progression from the language, to sending the script, to interpreting results. Front-loading the most critical instructions ('Call `get_library` FIRST') helps the agent prioritize. Minor repetition around sandbox restrictions is acceptable because those are safety-critical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity, the description is remarkably complete. It covers the execution model, the JavaScript environment, the return envelope, status semantics, `needs_user` handling, partial-result reporting obligations, trace structure, warnings handling, and composition patterns. The presence of an output schema and annotations further reduces the burden, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already fully documents the single `script` parameter, the description adds substantial meaning: the script must be a plain async body using `await bowmark.<capability>.<method>(...)`, must `return` JSON-serializable values, may use `log(...)`, and must not wrap itself in a function. It also explicitly states `script` is the only argument and there is no `site` argument, preventing a common misuse the schema alone would not prevent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it executes an authored JavaScript script against the real websites and returns `{ ok, result, logs, error, ms }`. It clearly distinguishes itself from siblings by instructing the agent to call `get_library` FIRST for the vocabulary, so the agent knows `run` is the execution step, not the reference tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: call `get_library` first, prefer capability methods over individual providers, and use a specific provider only when that exact site is needed. It also gives detailed instructions for handling `partial` and `needs_user` results, including when to rerun, when to wait, and what to tell the user.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.8/5.0
Disambiguation5/5

Each tool has a clearly distinct job: get_library returns capability documentation, run executes scripts against live sites, register creates credentials, and report captures feedback. There is no meaningful overlap or plausible misselection between them.

Naming Consistency4/5

All tool names are lowercase imperative verbs and follow a simple, readable style. The pattern is slightly inconsistent because get_library includes an object while register, report, and run are bare verbs, but the convention is still predictable enough.

Tool Count5/5

Four tools is well-scoped for this server's purpose: discover, execute, authenticate, and give feedback. Each tool earns its place and the count is within the ideal range for a focused MCP server.

Completeness5/5

The core workflow is fully covered: get_library and run form a complete discover-and-execute loop, register handles access and quotas, and report provides a path for missing capabilities. There are no obvious dead ends or missing lifecycle steps for the domain.