Skip to main content
Glama
Otha-Labs

Persuasion Taxonomy MCP

Plan an A/B test

plan_ab_test
Read-onlyIdempotent

Calculate the visitors and days an A/B test needs for marketing copy, using baseline conversion, minimum lift, and daily traffic to decide if it's worth running.

Instructions

Plan an A/B test for marketing copy. Give it the current conversion rate, the smallest lift worth detecting and the daily traffic, and it tells you how many visitors each version needs and how many days to run. It also frames the test as two different answers to the same reader question, so the result teaches you something you can reuse, and not just that "B won". Use it when someone asks how long to run a test, how much traffic they need, whether a test is worth running, or what to test next.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
powerNo
questionNoThe reader question both versions answer, each in its own way.
versionsNoHow many versions, counting the original.
version_aNoHow version A answers it.
version_bNoHow version B answers it.
confidenceNo
daily_visitorsYesHow many visitors a day enter the test, across all versions.
minimum_detectable_lift_percentYesThe smallest relative lift worth detecting, in percent. 20 means a 20% lift, like going from 3% to 3.6%.
baseline_conversion_rate_percentYesThe current conversion rate, in percent, so 3 means 3% and 0.8 means 0.8%.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and a closed world, so the safety profile is covered; the description adds the useful behavioral fact that this is a deterministic advisory calculation whose output includes a reusable framing of the test, not just a winner. It doesn't discuss edge behavior (e.g., how multi-version tests with versions > 2 are handled), which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with the core purpose front-loaded and then inputs, outputs and triggers in order; the middle sentence is a bit long and the 'not just that B won' aside is stylistic, but each sentence contributes distinct information rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the burden of describing returns: per-version visitor counts, days to run, and the framing output. Coverage is good for a 9-parameter tool, with only minor gaps around non-default power/confidence/versions semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78% (high), so the schema already documents most of the 9 parameters, including units for baseline rate and MDP lift, plus defaults/ranges for power and confidence. The description only restates the three required inputs ('current conversion rate, the smallest lift worth detecting and the daily traffic') and adds no format or unit guidance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Plan an A/B test for marketing copy') and immediately details the mechanics: it consumes conversion rate, minimum detectable lift and daily traffic, and returns required sample size per version plus run duration. That scope is clearly distinct from siblings like read_ab_test_result or plan_marketing_copy, though neither sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence gives four concrete trigger intents ('how long to run a test', 'how much traffic they need', 'whether a test is worth running', 'what to test next'), which is stronger than implied usage. It stops short of the 5 bar because it names no alternative tool or exclusion condition (e.g., when to use read_ab_test_result or plan_marketing_copy instead).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.