Bowmark
Server Details
Do things on live websites: prices, availability, quotes, bookings, anything behind a form or login.
- Status
- Healthy
- Uptime
- 99.9% over 43 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- bowmark-ai/skill
- GitHub Stars
- 0
TDQS
Scored across 9 tools
Each tool has a distinct job: get_library supplies vocabulary, run executes, report files feedback, and the connection/secret tools split cleanly by resource. The one real hazard is delete_connection vs logout_connection, two similar actions on the same saved login, but the descriptions explicitly contrast them ('delete forgets the entry; logout ends the session'), which largely defuses it.
Seven of nine tools follow a clean snake_case verb_noun pattern (list_connections, delete_connection, request_secret, get_secret_link, etc.). The two bare-verb names, 'report' and 'run', break the pattern, but both are short, readable, and unambiguous in context.
Nine tools is well-scoped for a platform that manages saved logins, secrets, a script vocabulary, execution, and feedback. Each tool occupies a distinct slot with no filler.
The read/execute lifecycle is well covered: list and delete connections, list and request secrets, fetch secret links, discover the library, run scripts, report gaps. The notable omission is any secret-removal or revocation tool (delete_connection has no secret-side counterpart), forcing users to the dashboard, but connection/secret creation is also deliberately dashboard-side, so this is a consistent, workable boundary.
Available Tools
9 toolsdelete_connectionForget one of the user's saved loginsADestructiveIdempotentInspect
Forget a saved login by the id list_connections returned. This deletes Bowmark's own record and cookies for it — it does NOT sign the account out on the site, and it cannot be undone; the user signs in again to get a working connection back.
Only call this when the user asked to remove a saved login. To sign out of a site, use logout_connection instead. Never call it to "fix" a connection that is merely stale — needs_reauth/expired recover with a new sign-in, not a delete.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The connection `id` from `list_connections`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although annotations already mark this as destructive and non-read-only, the description adds critical behavioral nuance: it deletes Bowmark's record and cookies, does NOT sign the account out on the site, and cannot be undone. This meaningfully goes beyond the annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized and appropriately sized for a destructive operation. It front-loads the core action, then covers consequences, irreversibility, and usage boundaries — every sentence carries distinct value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description is complete: it covers what happens, what does not happen, irreversibility, when to call it, and when not to. Combined with strong annotations, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter `id` is already documented as 'The connection `id` from `list_connections`'. The description repeats that provenance without adding new parameter-level detail, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('forget') and resource ('saved login'), and anchors the operation to the id returned by `list_connections`. It also disambiguates from the sibling `logout_connection` by explaining what the operation is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ('Only call this when the user asked to remove a saved login'), names the exact alternative for signing out (`logout_connection`), and prohibits use on stale connections where re-authentication is the correct path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_libraryCheck what Bowmark can do for this taskARead-onlyInspect
Use this whenever a task touches a live website. It answers, definitively and cheaply, whether Bowmark can already DO the thing: look up current prices, check real availability or stock, search a site, get a quote or a fare, drive a configurator, start a booking, or pull anything that only exists behind a form or a filter. Behind a LOGIN is narrower: it works once a purpose-built provider exists for that site, or the caller has a saved connection there — a generic caller with neither gets only a form-field reader, which cannot pass an identity check or obtain a vendor API key on its own.
Use this for Bowmark questions too. Before answering how to write or run a Bowmark script, what the sandbox supports, or how to start when no site or task is named, call it with the caller's words. The matching entry gives the platform guidance; guessing from general programming knowledge does not.
Checking is cheap, so check. One read-only call, no site is touched, and an unrecognized query returns a one-line index instead of an error, so the check never dead-ends and never costs you an attempt. If nothing fits, you have lost one cheap call and can use your normal approach.
What comes back is the callable function library you write against: the runtime globals (log) PLUS, for each capability your query named, its namespace, TypeScript types, functions, and worked examples. Everything listed is real and callable. The language rules and how to run a script are on the run tool description.
Pass query — what you want to DO ("flights", "price a GPU") or, if you have one in mind, the COMPANY or site ("Kayak", "newegg.com"). If the caller gave only a URL, use its hostname as the query — the query parameters identify a page state, not the company. A phrase in the user's own words is fine; it is matched against the whole library. You get what you asked about and nothing else. If nothing matches — or you send no query — you get instead a one-line index: pick whichever entry fits and CALL AGAIN with its name to get the types and examples you need to write a script.
Every response is bounded, and it says so when it is a slice. A broad query can match more than one response carries; when that happens the answer opens with a partial-answer line naming what it left out. Read it before concluding anything — absence from a sliced list means nothing, and the fix is one narrower query (a single task, or a single company by name), which always returns that entry in full. Only an answer that does NOT say it is a slice supports the conclusion that a task is uncovered.
THREE tiers come back, and the third is not written by Bowmark at all. SITE TOOLS (bowmark.sites["skims.com"].search_catalog(...)) are the tools a SITE publishes about ITSELF — its own MCP server, or the tools its pages register — read live at the moment you name the host and callable like anything else. Nothing has to have been built for that site in advance, so naming a site is worth doing even when you expect nothing to exist for it. The descriptions are the site's own words, every tool it exposes is callable (including ones that change something there), and site tools are free.
The other two tiers come back too. CAPABILITIES (bowmark.flights.search(...)) are the default and usually what you want: one call fans out across several sites, dedupes, ranks, and routes around a site that's failing. PROVIDERS (bowmark.providers.kayak.search(...)) are the individual sites, callable directly — they appear only when your query NAMED a company, or when the capability has just one provider behind it. A direct provider call gets that site's own raw shape and no failover, so prefer the capability unless you specifically want that site.
Loop: call get_library → write a JS script against the bowmark global → send it to run.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional, but pass one — it decides how much detail comes back. What the user wants to do, in their words ("flights", "price a GPU", "book a table", "check stock"), or a company/site if they named one ("Kayak", "newegg.com"). If they supplied only a URL, pass its hostname rather than the full URL or its query parameters — and a HOSTNAME is worth passing even for a site you expect nothing to exist for, because that is what makes Bowmark look for the site's own tools. A rough guess is always safe: a value that matches nothing returns the one-line index rather than an error, and so does omitting it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the annotations by disclosing that the call is read-only, touches no site, returns a one-line index on no match, can return partial/sliced results, and has three tiers including live third-party site tools. No contradiction with annotations is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but it is well-structured with bolded sections and a clear front-loaded usage statement. Some redundancy exists around the cheapness and one-line-index behavior, but nearly every paragraph carries actionable guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description fully explains return shapes, the three tiers, bounded/sliced responses, and the complete loop from `get_library` to `run`. Nothing essential is missing for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the single parameter well, and the description adds significant extra meaning: pass the user's words, use a hostname from a URL rather than query parameters, a rough guess is safe, and omitting the query still returns the index. This materially improves correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear, specific purpose: determine whether Bowmark can already perform a task on a live website and return the callable function library for it. The title and description align, and the behavior is concretely distinguished from generic knowledge lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this whenever a task touches a live website and for Bowmark questions, and it frames the check as cheap and never dead-ending. It also explains the follow-up workflow with `run` and what to do when nothing matches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_secret_linkGet the user a link to one of their own stored loginsARead-onlyInspect
Return the dashboard link for one stored secret this account holds, so you can answer "where is my Acme password?" with somewhere to go. It returns a URL, never a value — opening it requires the user to be signed in and to confirm who they are, and only their own browser can read what comes back.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The stored secret's name, as `list_secrets` shows it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, and the description adds meaningful behavioral detail: it returns only a URL, opening it requires the user to be signed in and confirm identity, and only the user's own browser can read the resolved value. This is security-relevant context beyond what readOnlyHint provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the function and a concrete use case; the second adds the critical return-type and security behavior. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only tool with strong annotations, the description fully covers what the tool returns, the authentication requirement, and the user-specific access model. Since there is no output schema, the explicit statement that it returns a URL is valuable and sufficient for an agent to invoke and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single `name` parameter is already documented as the secret's name as shown by list_secrets. The description adds only a mild constraint ('one stored secret this account holds') but no new format or validation detail, which matches the baseline of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: return the dashboard link for one stored secret. It also distinguishes itself from value-returning siblings by explicitly promising 'a URL, never a value,' so an agent can tell it apart from tools like request_secret or list_secrets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this when the user wants a navigable link to one of their stored secrets, such as answering 'where is my Acme password?' with somewhere to go. It does not explicitly name alternatives or say 'not when the secret value itself is needed,' so it stops short of a 5, but the URL-versus-value contrast provides strong guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_connectionsSee which sites Bowmark is already signed in toARead-onlyInspect
List the saved logins this account holds — the id to name in a later signed-in call as { connection }, which site, which account, whether it is still good, and when it was last used. A live one means a script can reach that site's signed-in pages with no sign-in step and usually no browser.
Useful before a task that needs an account: a connection marked needs_reauth or expired is why a run would pause and ask the user to sign in again. Adding a NEW login is still the user's to do, at https://bowmark.ai/dashboard/connections — delete_connection only forgets one this account already holds.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds behavioral context: that live connections allow scripts to reach signed-in pages without sign-in steps, and that needing reauth explains pauses. However, it doesn't detail the exact return format, which could be a gap, but the extra context is valuable beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two paragraphs, each with a clear purpose: first defines the output, second explains usage and exclusions. It is front-loaded with the core function. Slightly verbose with the URL and extra clarifications, but each sentence adds value, so no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool, the description covers what is returned, the meaning of statuses, and how to interpret results for scripts. It also clarifies limitations (adding logins is user-side). Without an output schema, it explains return fields adequately. Small gap: doesn't mention any pagination or limits, but for a likely small list this is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to add. The description focuses on the output and usage, which is appropriate. Since there are no params, the baseline is 4, and the description meaningfully explains the return fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists saved logins and enumerates the information included (id, site, account, status, last use), with a specific verb and resource. It distinguishes itself from siblings by explicitly mentioning it lists existing connections, not creating or deleting them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use (before a task needing an account) and what the statuses mean ('needs_reauth' or 'expired' causes pauses). It also clarifies that adding a new login is the user's responsibility and points to the dashboard, preventing confusion with delete_connection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_secretsSee which of the user's logins Bowmark already holdsARead-onlyInspect
List the stored secrets this account holds, by NAME. Returns name, type, which hosts each may be used against, when it expires and when it was last used — never a value, and no Bowmark endpoint returns one.
Call this BEFORE request_secret. A secret the user has already set is ready to use; asking for it again sends them to a page for nothing. In a script, refer to one by name: bowmark.secret('acme_pw').
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds the critical behavioral trait that no value is ever returned and no endpoint returns one, which is essential for an agent to avoid expecting secrets. It also specifies exactly what fields are returned. This adds significant transparency beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short paragraphs. The first sentence states the purpose and returns, the second adds usage guidance. Every sentence is informative with no redundancy. It is front-loaded with the core action and constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no parameters, the description is complete. It describes the output fields, the constraint on never returning values, and usage guidance with an example. No output schema exists, so the description carries the burden, which it does effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already covers all parameters (none). The description adds no parameter details because none exist. Baseline 4 is appropriate for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List the stored secrets this account holds, by NAME' and enumerates the returned fields, explicitly noting it never returns a value. It also differentiates from sibling request_secret by ordering. This is a specific verb+resource that distinguishes the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs 'Call this BEFORE request_secret' and explains the rationale (avoiding sending user to a page for nothing). It also gives script usage example. This clearly guides when to use this tool versus the key alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
logout_connectionSign one of the user's saved logins outADestructiveIdempotentInspect
Sign a saved login out by the id list_connections returned. Bowmark drops the cookies it held, and where the site supports it (siteLogout: "supported" in the reply) the session is ended ON THE SITE too — siteSignedOut: true means Bowmark checked the site no longer accepts it. The saved login is KEPT, status logged_out, so the user can sign back in to the same entry later.
Call this when the user asks to sign out of a site. delete_connection is the other action: it forgets the entry entirely and leaves the site session alive.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The connection `id` from `list_connections`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint and idempotentHint, but the description adds essential detail: Bowmark drops cookies, ends the site session if supported, and keeps the login with status logged_out. It also explains the meaning of siteSignedOut: true, going well beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: the first sentence states the action and id source, followed by behavioral details and usage guidance. Every sentence earns its place with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-id tool with no output schema, the description explains side effects, the site logout capability, the meaning of the response fields, and the distinction from delete_connection. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single id parameter at 100% with a description that already states it comes from list_connections. The tool description repeats this source but adds no new semantics beyond that, so the baseline 3 is appropriate for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Sign a saved login out') and resource (saved login/connection), and explicitly contrasts it with delete_connection, distinguishing it from siblings. The first sentence names the action and the identifier source, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to call it ('Call this when the user asks to sign out of a site') and identifies the alternative tool (delete_connection) with its different outcome (forgets the entry, leaves the site session alive). This gives clear guidance on selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reportReport what Bowmark could not do for this taskAInspect
Record what was missing, wrong, or incomplete so Bowmark can build it. Pass the runId returned by run when this report is about a run; omit it when get_library did not cover the task. This records feedback only and does not retry a run.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | No | Optional `runId` returned by `run` for the result you are reporting. | |
| report | Yes | What Bowmark could not do or returned incorrectly. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no informative annotations, the description usefully discloses that this records feedback only and has no retry behavior. It could add detail about side effects or permissions, but it provides materially more transparency than the raw schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences front-load the purpose, then provide the runId conditional and a no-retry warning. There is no filler or repetition of schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity 2-parameter tool with an output schema, the description covers purpose, parameter conditions, and behavioral boundaries. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so a baseline of 3 applies. The description adds value by explaining when runId should be present or omitted, which goes beyond the schema's generic 'Optional runId returned by run'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Record what was missing, wrong, or incomplete') and a clear resource (feedback for Bowmark to build). The title and description together distinguish this from run and get_library by framing it as a feedback-only operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditions: pass runId when reporting on a run, omit it when get_library did not cover the task, and do not use this to retry a run. This directly routes an agent among the sibling tools and states a when-not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_secretAsk the user to set a login Bowmark will needAIdempotentInspect
Create a named, empty slot for a secret and get back a link the user opens to fill it in. They type the value on a Bowmark page and choose how long it lives, from 5 minutes to never. It is encrypted in their own browser before it leaves, so nothing on the way — this connection included — ever sees it.
Give the user the link. Never ask them to type a password, API key or one-time code to you. A secret in this conversation is in your context, in the transcript and in the logs.
Name it for the person, not for your script. The name you pass is the heading on the page they open and the row they see in their secret list months later, so make it <site>_<what it is>: letterboxd_password, stripe_api_key, acme_totp_seed. A run id, a timestamp, a uuid or a bare password is refused.
Call list_secrets first: if the name already exists and is set, use it instead. name is lowercase letters, digits, _, . and -. hosts narrows where the value may be used and is worth passing — a secret is refused against any other site.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Lowercase name, `<site>_<what it is>` — `letterboxd_password`, `stripe_api_key`, `acme_totp_seed`. Letters, digits, _ . - only. THE USER READS THIS: it is the heading on the page they open and the row in their credential list months later, so use words, never a run id, a timestamp or a bare `password`. | |
| type | Yes | What kind of value the person will be asked for. | |
| hosts | No | Hosts the value may be used against, e.g. ["acme.com"]. Refused elsewhere. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses crucial behaviors beyond annotations: browser-side encryption, transcript/log exposure, refusal of non-human-readable names, and host-slot enforcement. None of these are visible in the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core mechanism, then uses bold directives for the two most safety-critical instructions. Despite length, every sentence earns its place by adding security, naming, or usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Provides outcome (a link), user journey, lifetime range, preconditions, validation rules, and failure behavior. No output schema exists, but the return value is described sufficiently for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already covers all three parameters, so the baseline is 3. The description adds meaning for name (user-facing heading, months-later list row, naming refusals) and hosts (worth passing, refused elsewhere), though type semantics remain only in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: create a named empty secret slot and return a user-facing fill-in link. This clearly distinguishes it from lookup/listing siblings like list_secrets and get_secret_link.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions: call list_secrets first and reuse an existing set secret, and instructs the agent to hand the link to the user rather than collect the secret. It does not explicitly compare with get_secret_link, so it loses a point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runDo the task on the live websites and return the resultADestructiveInspect
Executes the task on the real websites (the search, the price check, the availability lookup, the configurator, the booking flow) and returns what came back. Runs a script you authored against the get_library vocabulary, on the live sites, and returns { ok, result, logs, error, ms }. Call get_library FIRST — it gives the exact function names, argument shapes, and return types; this description is the LANGUAGE + how-to (get_library is just the vocabulary).
THE LANGUAGE — plain async JavaScript:
• bowmark is a ready global (no import). Call capabilities off it — await bowmark.<capability>.<method>(...) — always await, they're async.
• Individual sites are callable too, at await bowmark.providers.<provider>.<fn>(...). Use one when you specifically want THAT site; otherwise prefer the capability, which fans out across sites and routes around failures.
• Real control flow: await, if, loops, array methods (map/filter/sort/slice), and Promise.all for fan-out.
• return a value to get it back (JSON-serialized). log(...) for progress lines.
• Standard JavaScript built-ins are there (JSON, Math, Date, RegExp, Intl, Promise), plus URL and URLSearchParams — use them to resolve a relative link against the page it came from and to build query strings. Nothing else from the Web platform exists: no fetch, setTimeout, TextEncoder or crypto.
• bowmark is the ONLY I/O — no fetch, process, filesystem, or import/require. Write a plain async body, not a wrapping function.
• bowmark.files keeps a file the run produces (a CSV, an image, a transcript) in the caller's Bowmark account and gives you a link to it: save({ name, text | base64, contentType? }) → { id, url, expiresAt } (25 MB from a script), list(), get(id), url(id, { expiresIn }), makePublic(id) → { publicUrl }, makePrivate(id), delete(id). Return the link rather than inlining the bytes.
• Keep scripts small and deterministic — no infinite loops. Runs in a hard sandbox with CPU + memory + wall-clock limits.
Your own tool-call budget is tighter than you'd guess, and it decides how many calls fit in one script. Most MCP clients time a single tool call out at around 55 seconds, and ONE ordinary capability call already spends 30-55 seconds of that fanning out to live sites — see COMPOSITION below before calling a second capability in the same script.
SENDING IT: pass the script text as run({ script }) — script is the only argument (there is no site argument; the library exposes every capability under bowmark). result is whatever you returned; logs are your log() lines in order; on a throw/timeout ok:false and error is set.
CHECK status BEFORE ok. It is ok | error | partial | needs_user.
• partial means the script RAN and result is real and usable, but some of what it called never answered — so the result is narrower than what you asked for. ok is still true; this is not a failure. incomplete.summary says what happened in one sentence, incomplete.failures names each call that threw and what the site said, and incomplete.degraded names each call that answered while reporting its OWN results thin. You MUST say so when you present the result: name what was missed, and do not describe it as complete, exhaustive, or 'all' of anything. A partial you report as whole is a wrong answer, not a slightly smaller right one.
• Before you conclude a partial is final, check incomplete.failures[].fixable. fixable: true means YOUR ARGUMENT was rejected, not the site — the error text names what that function actually takes, so re-read it in get_library, fix the argument and run again; that recovers the whole answer. For any other failure re-running usually returns the same thing.
• When a fan-out capability (shipping, flights, hotels, cars, etc.) fails on EVERY provider, incomplete.failures[].dropped carries the per-provider breakdown as an array of { provider, fn, reason, detail } — branch on that instead of parsing the prose in .error.
• needs_user means a site needs the USER signed in — it is NOT a failure and NOT something you can fix by editing the script. needs lists the sites; meta.handoff.url is a single-use link that expires (meta.handoff.expiresAt). Give the user that URL, say which sites it covers, and WAIT. When they tell you they're done, send the SAME script again unchanged. Do NOT retry before then — it will stop at the same place and cost another run. Do NOT try to log in yourself, ask them for a password, or work around it with a different site.
• Logged-in runs need a Bowmark API key on the connection; if you get needs_user saying so, tell the user to add one rather than retrying.
ONE EXIT IP FOR THE WHOLE SCRIPT: await bowmark.egress.pin({ ttlMinutes: 120 }) before your first call, and every call the script makes afterwards leaves from ONE address. Reach for it when a site binds a cart, a token or a sign-in to the address that made the request, and the symptom is being logged out or asked to start over between steps. It is not a dedicated IP: the exit is a shared residential peer, shared: true says so, and the vendor can still replace it without telling us. Pin FIRST — a site you already called keeps the address it used, and the result says pinnedAfterCalls: true when you left it too late. Up to 7 days; anything longer comes back clamped, and the ttlMinutes you get back is the one that was granted.
A DEDICATED IP, IF pin ISN'T ENOUGH: pin is free and shared — a stranger's device, gone in at most 7 days. bowmark.egress.lease is the other product: a static IP nobody else is ever served, ordered for a real term and charged to the account, for a site whose session dies even on a pinned exit. Reach bowmark.egress.quote({ country }) first — it returns your own price for that country's real availability, read live; never assume a term or a number. lease({ country, idempotencyKey }) places the order (idempotencyKey is yours to choose, so a retried call cannot buy two); it is STATIC ISP ONLY today, real money, and there is no refund — confirm with your user before calling it, the way you would before any other purchase. list() and get(id) read what the account already holds; use(id) routes this run's later calls through one specific lease; release(id) stops auto-renewal (the IP keeps working for its already-paid term, then goes away for good).
SAVED SECRETS: savedSecrets lists every secret the run stored in the user's Bowmark account, each as { name, type, viewUrl }. An API key a key-making function returns is saved there automatically as <vendor>_api_key, and later runs that call that vendor use it with nothing to configure. When savedSecrets is present, tell the user what was saved and give them each viewUrl, where they can view or manage it.
trace is the execution trace — every capability you called and the providers it fanned out to under the hood: [{ kind:'capability', capability:'flights', method:'search', ms }, { kind:'provider', capability:'flights', provider:'google_flights', fn:'search', results, status, ms }, …]. The script never visits websites — it calls capabilities that route to providers, and the trace is the receipt.
COMPOSITION MEANS PARALLEL, NOT SEQUENTIAL, AND THAT GOES FOR PROVIDER CALLS TOO — an await inside a for loop is the single slowest thing you can write here. Measured 2026-09-18 over 147 real runs: a script with one awaits a for-loop body ran a median of 49 seconds against 19 for the rest, and on YouTube transcripts specifically, three fetched one after another took 27.3s where six fetched together took 9.8s. Collect the ids first, then Promise.all over them; never walk a list awaiting each item. Default to ONE capability call per script — most already spend 30-55 seconds of your own ~55-second tool-call budget on their own, so a second call made AFTER the first routinely never returns before your client gives up, and the script errors with nothing to show for either call. If you genuinely need several, run them TOGETHER inside Promise.all — in parallel they cost about what one call costs, not the sum of them — and never call them one after another. To sweep a date range, call the search per date inside Promise.all and sort/filter the merged array (each flight result carries its date, so you can tell the runs apart). See the get_library examples for the exact shape. If even one call will not fit your budget, narrow the query (fewer dates, a single site instead of a fan-out) or split the work across separate turns — do not compose more into one script to make it fit. A FAN-OUT OF FIVE THAT LOSES ONE SHOULD STILL HAND BACK FOUR: Promise.all rejects the whole thing the moment any leg does, so one refused request turns a finished answer into nothing. Use Promise.allSettled and keep the fulfilled legs. That is safe HERE and nowhere else, because Bowmark reports the dropped leg for you — the runtime counts what your script CALLED at the dispatch point, not what it chose to report, so a swallowed failure still comes back as status: 'partial' with the call named in incomplete. Say so when you present the result.
SOME capabilities return their rows alongside a warnings array — { flights, warnings }, { hotels, warnings }, { cars, warnings }. Others return a bare array. The signature in get_library tells you which; go by it rather than assuming. Where there IS a warnings array it names any site dropped from the fan-out, and the rows themselves look identical with or without it. Read it, and pass on anything it says rather than quoting a 'cheapest' that only ranks the sites that happened to answer. Dropping warnings from what you return does not hide it — the run comes back status: 'partial' regardless, because the runtime counts what your script CALLED, not what it chose to report.
bowmark.delegate hands a task to a coding agent (Claude Code) in a fresh, throwaway Linux machine — for work that needs a real computer: write and run code, work in a git repo, install packages, process files. It is billed to the account for machine time and model tokens, under maxCostUsd (default $5), so tell your user it is running and that it costs money. start({ prompt, env?, repo? }) returns id at once; pass any token the task needs in env, never in the prompt. In later runs call status(id, { waitMs: 60000 }) until it is idle, then read result, files and diff, and stop(id). send(id, message) continues the same conversation. It is not for reading a website — use the library or browser_agent for that. get_library({ query: "coding agent" }) has full types and examples.
bowmark.browser_agent is a one-shot live browser for ONE task on a site nothing in the library covers (a lookup came back empty, or a function failed on it). It is billed per turn to the account, and it is not a watcher: it cannot poll on a schedule or run unattended, so a standing watch or daily job is not a fit — say so. When it fits and your user asked you to get the thing done, start({ task }); it returns id and watchUrl at once, and in the same reply tell them it is running, that it costs money, and give them watchUrl. A login only they hold does not stop it — they sign in themselves on the live browser through that link. If you hold off for any reason, still NAME it in your reply with its cost — never end on a manual workaround without mentioning it. In later runs call status(id, { waitMs: 60000 }). When status is needs_input, relay question to your user and send(id, their_answer). When idle, read result and stop(id). Always stop a session when done — idle browsers keep spending money. Your account holds max 3 concurrent sessions. One task = one goal: never bundle multiple queries into one task or the agent silently reuses results. Start separate sessions for each. A NORMAL task takes 1-4 poll cycles to reach idle — roughly 1-3 minutes total, not one. Each poll's steps shows the agent's current activity, and the SAME step name on back-to-back polls (or a run of DIFFERENT step names) is ordinary mid-task activity, not a stall — do not stop a run that is still inside that window just because progress looks repetitive. get_library({ query: "browser agent" }) has full types and examples.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | The JavaScript script body to execute (async, against the `bowmark` global). e.g. `const { flights, warnings } = await bowmark.flights.search({from:'SFO',to:'JFK',depart:'2026-09-01'}); return { best: flights.sort((a,b)=>a.price-b.price)[0], warnings };` |
Output Schema
| Name | Required | Description |
|---|---|---|
| ms | No | |
| ok | Yes | |
| logs | No | |
| meta | No | |
| error | No | |
| needs | No | |
| runId | No | |
| notice | No | |
| result | No | |
| status | No | |
| incomplete | No | |
| savedFiles | No | |
| emptyHanded | No | |
| secretNeeds | No | |
| savedSecrets | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-readonly, open-world, destructive, non-idempotent operation, and the description goes well beyond that: it documents the four `status` values and their meanings, how `partial` must be reported, the `fixable` flag, `needs_user` handoff semantics, the hard sandbox CPU/memory/wall-clock limits, and the ~55s per-call budget. It also discloses money/cost behavior for `delegate`/`browser_agent` and the shared-vs-dedicated nature of egress IPs. No statement contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core guidance is front-loaded and each major section is clearly headed, which is appropriate for a general code-execution tool. However the body is very long and packs in full documentation for adjacent capabilities (`delegate`, `browser_agent`, `egress.lease`) that arguably belong in those tools' own definitions, diluting focus on `run`. A meaningful share of the text is tangential rather than earning its place here.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, return values need no extra explaining, yet the description still covers status semantics, failure recovery, budget/composition limits, egress behavior, and saved secrets. For a high-complexity, open-world execution tool this is as complete a briefing as an agent could need before calling it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already carries an inline `script` example, so 3 is the baseline. The description adds genuine meaning beyond the schema: `script` is the ONLY argument (there is no `site` argument), it must be a plain async body rather than a wrapping function, `return` gets JSON-serialized back, and `log()` produces ordered log lines. That clarifies invocation in ways the schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific verb, resource, and scope: it executes an authored JavaScript script against the `bowmark` global on live sites and returns a defined shape (`{ ok, result, logs, error, ms }`). It positions itself precisely against its closest sibling by directing the agent to `get_library` FIRST as the vocabulary source while framing this tool as the execution/how-to layer. An agent can distinguish `run` from `get_library` or `report` without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use and when-not-to: call `get_library` first, default to ONE capability call per script, use `Promise.all`/`allSettled` rather than sequential awaits, and use `bowmark.delegate` or `browser_agent` only for work the library does not cover (with `delegate` explicitly excluded for simple site reads). It also names the conditions that select each alternative and spells out the `needs_user`/`partial` recovery branches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Added
delete_connection - Added
logout_connection
Related MCP Connectors
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Turn any webpage into a structured action manifest — clickable, fillable, submittable elements.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to browse websites, extract information, fill forms, interact with pages, and execute multi-step web tasks autonomously.-
- AlicenseBqualityBmaintenanceGives agents an autonomous browser that takes a single goal-level call and drives a real Chrome step by step, returning a verified result or a concrete question instead of guessing. It keeps passwords out of prompts, requires approval before buy/pay/delete actions, and detects login walls and bot challenges.8MIT

Cloak Agentofficial
AlicenseAqualityBmaintenanceEnables an agent to run whole browser tasks on live web pages from a plain-language goal — opening tabs, clicking, typing, scrolling, and returning the relevant page sections as markdown — with every step reported alongside its confidence. Tabs can be reused and closed across calls, so multi-step flows continue on the page the previous task ended on.21MIT- AlicenseAqualityCmaintenanceEnables AI agents to search hotels, check availability, manage reservations, and book rooms on Booking.com via browser automation.1526 npm5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.