Skip to main content
Glama

Menso

Run a Menso test

run_test

Start a Menso test on a live website. An AI user with a realistic persona opens the site in a real cloud browser and tries to complete a task; Menso then scores the run (TRACES, 0-100) and records where the user got stuck, with a replay of every step. Templates: 'purchase' (default) reads the homepage and pricing and decides whether to buy, no account needed; 'signup' creates a new account with the email and password you pass and continues to the signed-in home. Each run spends Menso credits exactly like a run started on menso.io (speed 40 credits, quality 300 credits); the balance is checked before anything starts. The URL must be publicly reachable (use a preview deployment, not localhost). A run usually takes several minutes: poll get_status every 30-60 seconds, then call get_findings.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPublic address of the site to test, e.g. https://example.com.
tierNospeed: 40 credits, faster model. quality: 300 credits, strongest model.speed
emailNosignup only: email for the new test account. Use a dedicated test inbox.
passwordNosignup only: password for the new test account. Never reuse a real password.
templateNopurchase: decide whether to buy from the homepage and pricing. signup: create an account (needs email and password) and reach the signed-in home.purchase

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNo
tierNo
stateYes
creditsNo
test_idYes
templateNo
history_urlNo
queue_positionNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the generic flags (readOnly=false, idempotent=false, openWorld=true); the description adds the details an agent actually needs: credit cost per tier (40/300), that the balance is checked before anything starts, that runs take several minutes and must be polled, that the URL must be publicly reachable, and that the signup template creates a real account on the target site with the supplied credentials.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose first, then templates, then cost, then constraints, then the polling handoff. Every sentence earns its place except the template glosses, which partially restate the enum descriptions already present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-style, credit-consuming, multi-minute operation this covers cost, balance checking, environment requirements, side effects and the follow-up call sequence. An output schema exists, so the description correctly does not spend words on return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries the baseline. The description still adds cross-parameter meaning the schema does not: that signup consumes the email and password 'you pass' in order to create the account and land on the signed-in home, and that purchase is the default that short-circuits the need for credentials.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: 'Start a Menso test on a live website.' It goes further by defining what the test actually is (an AI persona on a real cloud browser attempting a task) and what the output is (a TRACES score plus a step replay), which cleanly separates it from the sibling polling tools get_status, get_findings and get_replay_link.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance for both templates ('purchase' needs no account; 'signup' requires email/password and continues to the signed-in home), plus explicit routing to alternatives after starting: 'poll get_status every 30-60 seconds, then call get_findings.' It also states an environment precondition (public URL, use a preview deployment, not localhost).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.