Skip to main content
Glama

Shipfound

Start an A/B test

experiment_create

Size and start an A/B test on one page against one goal. The site must be in Attribution mode. Reads the page's last 28 days, works out the visitors each variant needs to detect the lift (mde, default 0.2 = 20%) and how many weeks that takes, and refuses a test that would take over 12 weeks or a page under 50 testable visitors a week, with the numbers. On success it returns the code to put on the page (Next.js with @shipfound/next, or plain HTML with sf.variant), which you ship as a PR. The control is the first variant. Free.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
keyYesLowercase key the page's code uses, e.g. pricing-headline
mdeNo
goalYesA goal on the site, by name or id (see goals)
nameYes
pathYesThe page, e.g. /pricing (* as a wildcard)
siteNoThe site's domain, when the workspace tracks more than one
variantsNoDefault ["control", "b"]; the control first
hypothesisYesWhat changes, for whom, why it should move the goal, and the analytics evidence

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does much of it: it reads 28 days of data, computes required sample size, refuses statistically unviable tests with numbers, and returns shippable code. It omits auth/permission requirements and rate limits, but the operational behavior is unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A dense but front-loaded paragraph where nearly every clause adds decision-relevant information (prerequisite, computation, refusal rules, return value, free-tier note). Slightly crowded, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain returns, and it does ('returns the code to put on the page... which you ship as a PR'). Prerequisites, failure modes, and the control convention are all covered, leaving nothing an agent needs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% (baseline 3), and the description adds value the schema lacks: the mde default of 0.2 (20%) is not in the schema, and it clarifies the control is always the first variant. It does not document key, goal, path, or hypothesis beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Size and start an A/B test') with scope ('on one page against one goal'), clearly distinguishing it from experiment_close and experiment_results. An agent can identify the tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a hard prerequisite ('The site must be in Attribution mode') and the failure conditions under which it refuses (over 12 weeks, under 50 testable visitors/week). It does not explicitly route to a sibling alternative the way an ideal definition would, but the usage context is concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources