Skip to main content
Glama

Auth Your Agent

An open-source MCP server that lets an AI agent act for a person on websites, with the person's approval on their phone.

The agent drives a sandboxed browser (the vault) that runs on the person's own machine. When a site asks for a password, a CAPTCHA or a 2FA code, the person takes over that browser from their phone, signs in, and hands it back. The agent never sees the password and never holds the cookies. When the task is done, the vault signs out of every site it used, then destroys the browser profile.

Website: authyouragent.com · License: MIT

What is in this repository

The MCP server, the browser vault and the SDKs:

Path

What it is

sdk/authyouragent/mcp_server.py

The MCP server (stdio, 13 tools). Entry point authyouragent-mcp.

vault/

The browser vault: broker.py (HTTP API the MCP server calls; drives Chromium, detects sign-in, signs out), egress.py (public-internet-only proxy), screen.py (phone screen stream), seccomp.json, chromium-policy.json, Dockerfile. Details: vault/README.md.

sdk/authyouragent/vault_cli.py

authyouragent vault up/down/status/env: runs the vault with every protection on.

sdk/authyouragent/

Python SDK: agent side (agent.py), website side (site.py), take-over helper for your own Playwright browser (takeover.py).

sdk-js/

JavaScript SDK, and npm-wrapper/ so Node clients can start the MCP server with npx authyouragent-mcp.

Dockerfile

Container for the MCP server alone.

The phone app and the approval service run at authyouragent.com. The agent signs its requests with its own key pair; the service never receives the person's passwords.

Related MCP server: agent-auth

MCP tools

Tool

What it does

navigate, click, type_text, read_page

Drive the vault's browser. Only public websites open. Clicks that submit a form, or whose button says create, send, save, delete, pay and the like, ask the owner's phone first.

list_secrets, fill_secret

Fill a username, password or authenticator code from the owner's Bitwarden or Vaultwarden into the page. The agent never sees the value; it fills only on the item's own site and only the right kind of field.

check_login_wall

Is the page blocked by a password, CAPTCHA, one-time code or sign-in approval? Reads the page's fields and text, so the model needs no vision.

request_takeover

Ask the owner to take over the browser from their phone. Returns done, cancelled, expired, incomplete, or waiting.

wait_for_takeover

Keep waiting after waiting.

request_approval

Ask the owner to approve an action. approved, denied or expired.

end_session

Sign out of every site used, then destroy the browser profile. Reports per site whether sign-out was confirmed.

check_agent_status

active or revoked.

report_site

Report a site where take over did not work.

Run it

From PyPI:

pip install "authyouragent[mcp]"
authyouragent vault up --agent-id ag_... --key /path/to/agent-key.pem

From this repository:

git clone https://github.com/kjames2001/authyouragent && cd authyouragent
pip install "./sdk[mcp]"
docker build -f vault/Dockerfile -t authyouragent/vault:local .
authyouragent vault up --agent-id ag_... --key /path/to/agent-key.pem --image authyouragent/vault:local --no-pull

The vault needs Docker. vault up prints the MCP settings to add to your client. (If you skip vault up, the MCP server starts the vault itself the first time the agent needs the browser; AYA_VAULT_AUTOSTART=0 turns that off.)

{
  "mcpServers": {
    "authyouragent": {
      "command": "authyouragent-mcp",
      "env": {
        "AYA_AGENT_ID": "ag_...",
        "AYA_KEY_FILE": "/path/to/agent-key.pem",
        "AYA_VAULT_URL": "http://127.0.0.1:7801",
        "AYA_VAULT_TOKEN_FILE": "~/.authyouragent/vault/token"
      }
    }
  }
}

Node clients can use npx authyouragent-mcp as the command. The agent ID and key come from adding the agent in the phone app (Agents > Add an agent). Without them the server still starts and lists its tools; each tool then says what is missing.

Security

  • Chromium's own sandbox is on (custom seccomp profile, all container capabilities dropped); gVisor is used automatically when Docker has the runsc runtime.

  • The browser can reach only the public internet: loopback, private networks, cloud metadata addresses and the vault's own ports are refused, checked on the resolved address.

  • The agent is disconnected while the owner is in control, and cannot read the browser's cookies.

  • Sign out first, wipe second: at the end of a session, when the agent stops sending heartbeats, or when it is revoked. Sites the vault could not sign out of are reported to the owner, also after a crash.

Known limits are listed in vault/README.md.

Documentation

Android app

Download from authyouragent.com/download/android.

License

MIT

Available Tools

13 tools
check_agent_statusA

Check whether your agent is still authorized by the owner. The owner can revoke your access at any time from their phone. Call this periodically (e.g. before starting a new task) to verify you are still active. Returns: active / revoked / error.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the owner can revoke access at any time and enumerates the possible return states (active / revoked / error). It does not explicitly state that the call is read-only or side-effect-free, but for a simple status check it is fairly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences plus a returns line, front-loaded with the core purpose. Every sentence earns its place by adding context, usage timing, or return values.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity (zero parameters) and the presence of an output schema, the description is nearly complete. It explains purpose, when to call, and return values. It omits what an agent should do after receiving a 'revoked' status, but that is arguably beyond the tool's scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to document. The baseline score for a parameterless tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific verb and resource: 'Check whether your agent is still authorized by the owner.' It is not a tautology and an agent can easily tell it apart from siblings like check_login_wall or request_takeover, though the description does not explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: 'Call this periodically (e.g. before starting a new task) to verify you are still active.' It does not mention when not to use it or name specific alternative tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_login_wallA

Check whether the page is stuck at a sign-in, one-time code (2FA) or sign-in approval step. Reads form fields and page text only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose a meaningful trait: it 'reads form fields and page text only,' signalling a passive, non-mutating check that does not click or submit. It does not state repeat-safety or side effects, but the read-only scope is a genuine disclosure beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the detection target front-loaded and the read-scope constraint following. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and with no parameters the surface is minimal. The description covers what the check detects and how, though it omits what the agent should do after a positive result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter surface for the description to explain; baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (login wall) and enumerates exactly the states it detects: sign-in, one-time code (2FA), sign-in approval. This distinguishes it from read_page and check_agent_status, though it does not name an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the purpose (call it to determine whether the page is blocked at an auth step), but there is no explicit when-to-use, when-not, or routing to siblings like request_takeover/request_approval that follow a positive detection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickA

Click an element on the current page and return the page's URL and title afterwards. If the click would commit something (it submits a form, or the button says create, send, save, delete, pay and the like), the vault first asks the owner to approve it on their phone and waits up to about five minutes; a denial returns an error and nothing is clicked. Sign-in and search forms are not interrupted. selector: a CSS or Playwright selector, e.g. "button[type=submit]" or "text=Add Server".

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so: it discloses the commit-detection heuristic (submit/create/send/save/delete/pay), the out-of-band phone approval and ~5-minute timeout, the failure semantics (denial returns an error and nothing is clicked), and the exemptions for sign-in/search forms.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place, covering action, return value, approval semantics, and parameter format. The selector guidance is appended at the end rather than front-loaded, which slightly weakens scanning order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema already exists, yet the description still notes what is returned (URL and title). Combined with the approval/denial semantics, an agent has everything needed to call this tool safely and predict delays.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must document the single parameter, and it does — 'a CSS or Playwright selector' with concrete examples like "button[type=submit]" and "text=Add Server". It stops short of stating ambiguity/unique-match behavior, so it is good but not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (an element on the current page), and specifies the return value (URL and title). An agent can distinguish it from siblings like type_text, fill_secret, and navigate without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear situational context: clicks that commit something trigger an owner approval flow with a ~5-minute wait, while sign-in and search forms are not interrupted. It doesn't name sibling alternatives (e.g., when to prefer fill_secret or type_text), so routing guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

end_sessionA

End the owner's session when you are finished with the web: the vault signs out of the sites used, then destroys the browser profile. Always call this when done. (The vault also does it by itself if you stop running or the owner revokes you.)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that sites are signed out and the browser profile is destroyed, and that the vault performs this automatically if the agent stops or is revoked. It does not state whether the destroyed profile means lost cookies/state or whether the operation is recoverable, leaving a modest gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and its effects before the operational instruction. The parenthetical about automatic invocation earns its place by explaining fallback behavior, though the phrasing is slightly loose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, zero-annotation tool with an output schema present, the description covers purpose, trigger, side effects, and automatic fallback. Nothing an agent needs in order to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the schema has nothing to document and there is no parameter semantics to explain. Baseline 4 applies per the rubric for parameterless tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (end) and resource (the owner's session), and goes further by describing the mechanism: signs out of sites, then destroys the browser profile. This clearly distinguishes it from every sibling, which are page-interaction tools (navigate, click, read_page).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Always call this when you are finished with the web" gives an explicit trigger condition, which is stronger than most tools. It lacks an explicit when-not, though the parenthetical clarifies the vault performs this automatically on agent stop or revocation, which implicitly sets expectations about necessity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fill_secretA

Fill a field from the owner's password manager. The vault types the value itself; you never see it. It only fills on a site saved with the item, and only the right kind of field: a password into a password field, a totp code into a one-time code field, a username into a text or email field. Then click the sign-in button as usual. selector: the field, e.g. "input[type=password]". name: the item name from list_secrets. field: username, password or totp.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
fieldNopassword
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well: it discloses that the vault types the value so the agent never sees it (a critical privacy trait), and that filling is constrained to saved sites and matching field kinds. It omits failure/error behavior and what happens when no matching item exists, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose, key constraint (agent never sees the value), and fill-matching rule are all front-loaded before the parameter notes. Sentences are tight and each carries information, though the trailing param block reads slightly like an appended glossary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be restated, and the description covers the fill semantics thoroughly for a 3-param tool with no annotations. The remaining gap is error/edge behavior (item not found, selector mismatch) for what is effectively a mutation of form state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: selector is explained as the target field with an example, name is tied to the sibling list_secrets as its source, and field enumerates the valid values (username, password, totp). This is meaningful semantics well beyond the schema's bare titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource (fill a field from the password manager) and immediately distinguishes itself from the sibling type_text by noting the vault enters the value rather than the agent. An agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditions for correct use (site must be saved with the item, field type must match) and describes the follow-up action (click the sign-in button as usual, i.e. use the click sibling). It does not explicitly name type_text as the alternative for non-secret text, so it stops short of full when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_secretsA

List the sign-ins the owner has made available to you from their password manager: each item's name, the sites it may be used on, and which fields it has (username, password, totp). Values are never shown to you.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well by disclosing that values are never shown and that only names, sites, and field types are exposed. It does not cover rate limits or permission requirements, but the critical privacy behavior is explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the resource, summarizes the returned metadata, and ends with the key constraint that values are never shown. Every clause adds useful information without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with an output schema already defining return structure, the description provides sufficient context: what is listed, what metadata is available, and that secret values are hidden. No annotations are needed for this safe read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The empty input schema is consistent with the description, which correctly implies no input is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (list) and resource (sign-ins/secrets from the owner's password manager), and clarifies what metadata is returned. It is clearly distinguishable from sibling tools like fill_secret, which would use a secret rather than enumerate available ones.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit: the agent can infer this is for discovering available sign-ins, likely before calling fill_secret. However, the description does not state when to use it versus alternatives or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

Read the current page: its URL, title and visible text, in reading order. Use it after navigate or click to see what is on screen (no screenshot needed). Text in form fields and hidden elements is not included. max_chars: how much text to return, 200 to 20000 (default 5000).

ParametersJSON Schema
NameRequiredDescriptionDefault
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so reasonably: it discloses that form-field text and hidden elements are excluded, that output follows reading order, and that the tool is a read operation requiring no screenshot. It omits things like permission or rate-limit considerations, but for a read-only page dump the key behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short, front-loaded sentences: purpose first, then when to use, then a behavioral exclusion, then the parameter. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be spelled out, yet the description still characterizes the content and its ordering. Combined with the parameter range/default and the exclusion rules, an agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it fully documents the single parameter: max_chars is explained as 'how much text to return, 200 to 20000 (default 5000)', including valid range and default. The stated default matches the schema, leaving no ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (the current page), and enumerates the exact payload: URL, title, and visible text in reading order. An agent knows precisely what this returns and how it differs from navigate and click, which only alter page state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use it after navigate or click to see what is on screen' and notes no screenshot is needed, which is clear context for invocation. It stops short of naming when *not* to use it or citing a concrete alternative sibling, so it doesn't reach the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_siteA

Report a site where take over did not work correctly, so the team can improve it. Call this whenever check_login_wall misses a login, takeover fails to complete, the handback does not trigger, or anything else goes wrong on a real site. site: the domain or URL, e.g. "reddit.com". kind: one of: shadow_dom, bot_detection, input_not_detected, takeover_failed, handback_failed, session_not_cleared, oauth_blocked, other. detail: one short line describing what happened (optional). url: the full page URL if known (optional).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
kindYes
siteYes
detailNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies an external side effect ('so the team can improve it') but does not state whether the report is sent off-box, whether it is throttled/deduplicated, whether it is anonymous, or what confirmation the caller gets. The 'real site' qualifier hints at a test-vs-production distinction but is not elaborated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage triggers, then a compact per-parameter glossary. Every line earns its place, and the parameter block is justified by the 0% schema coverage; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Combined with the trigger list and complete parameter glossary, the definition is nearly self-sufficient; only the side-effect/telemetry behavior of the report remains unstated for a tool that transmits failure data externally.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage and no enum constraints in the schema, the description fully compensates: it defines site (domain or URL with an example 'reddit.com'), enumerates all eight valid kind values that the schema does not constrain, and clarifies that detail and url are optional with format guidance ('one short line', 'full page URL'). This adds substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: report a site where takeover failed, with the framing 'so the team can improve it.' It distinguishes itself from siblings by naming check_login_wall, takeover, and handback as the failure sources, making it identifiable as the diagnostic-reporting tool among browser-automation siblings. Slightly odd phrasing ('report a site') keeps it short of a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use triggers are enumerated: 'Call this whenever check_login_wall misses a login, takeover fails to complete, the handback does not trigger, or anything else goes wrong on a real site.' It names the specific sibling tools whose failures should route here, plus a catch-all, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_approvalA

Ask your owner to approve a sensitive action before you perform it. The owner receives a push notification on their phone showing the site and action. They approve or deny. You should not perform the action until this returns approved. Use this before deleting data, changing settings, making purchases, modifying permissions, or any action that could have irreversible consequences. site: the domain or URL of the site where the action will be performed, e.g. "github.com". action: a short label for the action, e.g. "delete repository", "change password", "purchase subscription". wait_seconds: how long to wait for the owner's response (default 120, max 290). Returns: "approved" or "denied" or "expired" (owner did not respond in time).

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYes
actionYes
wait_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries full burden and does so thoroughly: it explains the push notification to the owner, the approve/deny/expired outcomes, the blocking wait behavior, and the wait_seconds limit (max 290). This gives the agent a complete mental model of the interaction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded and well-structured, moving from purpose to mechanics to usage to parameters. However, the final 'Returns: approved or denied or expired' line is partially redundant given that an output schema exists, slightly reducing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters, no annotations, and an existing output schema, the description covers everything an agent needs: what it does, when to use it, how it behaves, the meaning of each parameter, and what outcomes to expect. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it defines site (domain/URL, e.g. 'github.com'), action (short label, with examples like 'delete repository'), and wait_seconds (default 120, max 290). These examples and constraints add substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Ask your owner to approve a sensitive action before you perform it.' The scope is unambiguous and inherently distinguishes it from browser/navigation siblings like click or navigate. An agent can tell exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this before deleting data, changing settings, making purchases, modifying permissions, or any action that could have irreversible consequences.' Also provides the critical when-not: 'You should not perform the action until this returns approved.' Clear routing guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_takeoverA

Ask your owner to take over the browser from their phone when you are stuck at a login, 2FA code, CAPTCHA or anything only they can do. You are disconnected from the browser while they are in control, and you never see what they type. Returns result: done / cancelled / expired / incomplete, or result: waiting with a takeover_id for wait_for_takeover. reason: one short line shown to the owner, e.g. "Please sign in to example.com".

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
wait_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it delivers: it discloses that the agent is disconnected from the browser during control and never sees what the owner types, plus the possible terminal states (done/cancelled/expired/incomplete) and the 'waiting' intermediate state. It stops short of stating auth/permission requirements or timing constraints on the request itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then behavior, then return states, then the parameter note. Dense but every sentence carries information; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter, no-annotation tool the description is nearly complete, and it goes beyond an existing output schema by enumerating result values and the wait_for_takeover handoff. The only gap is the undocumented wait_seconds parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains 'reason' well (one short line shown to the owner, with an example), but 'wait_seconds' (default 240) is never mentioned, leaving one of two parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Ask your owner to take over the browser') and the exact triggering conditions (login, 2FA, CAPTCHA, anything only they can do). It is clearly separable from siblings like wait_for_takeover, which it names as the follow-up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use condition ('when you are stuck at a login, 2FA code, CAPTCHA') and routes the waiting case to wait_for_takeover. It does not contrast with other human-in-the-loop siblings such as request_approval, so no explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textA

Type into a form field (replaces its value). Never use this for the owner's passwords or codes: use fill_secret if the owner saved the sign-in, otherwise ask for a take over. selector: the field, e.g. "input[name=url]". text: what to type. submit: press Enter afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
submitNo
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the key behavioral traits: the write overwrites the field's current value, submit triggers Enter, and credential input is blocked with a defined fallback path. It does not cover wait/timing, error behavior, or rate limits, but the mutation and security semantics are clearly surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then the critical exclusion, then a compact per-parameter list. Every sentence earns its place and there is no padding, though the trailing parameter glossary is slightly terse rather than elegantly integrated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Purpose, exclusions, fallback routing, and all three parameters are covered, leaving only operational edge cases (page-load waits, failure modes) unaddressed, which is acceptable for this tool class.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all three parameters: selector with a concrete CSS example ('input[name=url]'), text as the value to type, and submit as 'press Enter afterwards'. The example selector syntax is genuinely additive beyond the bare schema types, though 'text: what to type' is thin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and target ('Type into a form field') and immediately qualifies scope with '(replaces its value)'. It also names a sibling (fill_secret) that must be used instead for credentials, so an agent can distinguish it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-not ('Never use this for the owner's passwords or codes') plus the correct alternative branching in both cases: fill_secret when the sign-in is saved, otherwise request a takeover. This is exactly the when/when-not/alternative guidance called for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_takeoverA

Keep waiting for a take over that request_takeover returned as "waiting". Returns the same results as request_takeover: done (continue on the signed-in page), cancelled or expired (stop and tell your user), incomplete (the page still asks for sign-in), or waiting again with the same takeover_id. takeover_id: the id request_takeover returned, e.g. "tk_...". wait_seconds: how long to wait this time (default as request_takeover).

ParametersJSON Schema
NameRequiredDescriptionDefault
takeover_idYes
wait_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses return states (done, cancelled, expired, incomplete, waiting), what each means for the agent's next action, and that wait_seconds defaults to request_takeover's value. It does not mention potential blocking behavior, rate limits, or invalid-id handling, leaving some operational context unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and then covers return values and parameters in a logical order. It is slightly repetitive by restating that results are the same as request_takeover, but every sentence adds useful guidance for invoking or interpreting the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (two parameters, polling follow-up) and the presence of an output schema, the description is largely complete: it explains the relationship to request_takeover, the possible outcomes, and both parameters. Minor gaps remain around the wait_seconds default and edge cases like invalid ids.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents both parameters: takeover_id is the id from request_takeover (e.g., 'tk_...'), and wait_seconds is how long to wait this time. However, wait_seconds says only 'default as request_takeover' without specifying the default value or units, so the explanation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('keep waiting') and resource ('takeover') and explicitly ties it to the result of 'request_takeover' returning 'waiting'. This differentiates it from the sibling request_takeover and other tools by clarifying it is a follow-up polling call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it: after request_takeover returns 'waiting'. It also provides outcome-based guidance (e.g., stop on cancelled/expired, continue on done). However, it does not explicitly state when not to use it or name alternatives beyond the implicit request_takeover relationship.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.3.8
    • Addedfill_secret
    • Addedlist_secrets
  2. 11 tool updatesv0.3.1
    • First observedcheck_agent_status
    • First observedcheck_login_wall
    • First observedclick
    • First observedend_session
    • First observednavigate
    • First observedread_page
    • First observedreport_site
    • First observedrequest_approval
    • First observedrequest_takeover
    • First observedtype_text
    • First observedwait_for_takeover

TDQS

A4.2/5.0

Scored across 13 tools

Disambiguation4/5

Most tools target clearly distinct operations (navigate vs click vs type_text vs read_page vs fill_secret). The takeover cluster (check_login_wall, request_takeover, wait_for_takeover) and the two check_* tools risk minor confusion, but the descriptions explicitly explain how they chain together.

Naming Consistency4/5

Strong snake_case, verb-first convention throughout (list_secrets, fill_secret, read_page, end_session, request_approval, request_takeover). Only minor deviations: bare-verb browser tools (navigate, click) and the verb_prep_noun form of wait_for_takeover, which still stay readable and predictable.

Tool Count5/5

13 tools is well within the ideal 3-15 range and each one earns its place: browser control, secret filling, the login/takeover lifecycle, approval, session teardown, reporting and status all cover distinct needs without redundancy.

Completeness4/5

The set covers the full auth-agent lifecycle: authorization check, browsing actions, secret filling, login-wall detection, owner takeover, approvals, session cleanup and issue reporting. Minor gaps exist (e.g. no explicit scroll, back-navigation or screenshot), but core agent workflows have no dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Lets AI assistants control your real Chrome browser to perform web tasks like reading pages, taking screenshots, clicking, and typing, using your existing logged-in sessions.
    133
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Zero-knowledge credential injection for AI agents. Your agent authenticates to websites and APIs without ever seeing a password, TOTP code, or API key.
    10 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to drive a privacy-first browser via MCP, using diff-based page observation to reduce token consumption, and offering secure autofill from a local vault.
    2,181 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Lets Claude Code drive a real Google Chrome while you watch a low-latency live view and take over the mouse and keyboard whenever a login, CAPTCHA or payment appears. Sensitive actions pause for phone-based approval, and the agent can be guided by tapping elements or sending short messages without exposing secrets.
    32 npm
    MIT