Skip to main content
Glama
Otha-Labs

Persuasion Taxonomy MCP

Read an A/B test result

read_ab_test_result
Read-onlyIdempotent

Analyze A/B test results to find the winning version, check statistical significance, and flag pitfalls like sample ratio mismatch or low conversions.

Instructions

Read the result of an A/B test. You get each version's conversion rate and lift, whether the difference is statistically significant, the confidence interval, and the chance each version really beats the original. It also warns you about the things that make a test result lie. Those include too few conversions, traffic that didn't split the way it should (a sample ratio mismatch), stopping the moment it looked good, and testing too many versions at once. So use it whenever someone shares test numbers, or asks "did my test win?", "is this significant?", "which version won?" or "should I keep it running?"

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
questionNoThe reader question the versions answered in different ways.
versionsYesEach version with its visitors and conversions. Put the original, the control, first.
confidenceNo
intended_splitNoHow you meant to split the traffic, like [50, 50]. The default is an even split.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description adds substantial behavioral context beyond annotations: it explains the returned statistical measures and warns about reliability pitfalls like too few conversions, sample ratio mismatch, peeking, and testing too many versions at once. This is precisely the kind of added value expected when annotations are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and outputs, then moves to warnings and usage triggers. It is efficient overall, though the list of trigger questions at the end is somewhat expansive. Every sentence contributes useful information, so it is well structured but slightly long.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately explains the returned values and statistical warnings. It is nearly complete for a read-only analysis tool, but it omits any guidance on the confidence parameter and does not clarify how the intended_split default relates to the input. These are minor gaps against an otherwise thorough description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with the question, versions, and intended_split parameters documented in the schema, while confidence has no schema description. The description does not explain any parameter semantics, including the confidence parameter or the intended_split format. When schema coverage is moderately high, a baseline of 3 is appropriate because the schema does the heavy lifting and the description adds no parameter-specific meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: reading the result of an A/B test. It goes further by enumerating the outputs (conversion rate, lift, significance, confidence interval, chance to beat original) and warnings, which clearly distinguishes it from planning siblings like plan_ab_test. An agent can identify the correct tool without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage triggers such as when someone shares test numbers or asks 'did my test win?', 'is this significant?', 'which version won?', and 'should I keep it running?'. It does not explicitly state when not to use it or name the alternative plan_ab_test, so it falls short of full when/when-not/alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.