Activated Cloud
← App Store

A/B Test Design

Activated Cloud✓ Officialactivated/ab-test-design

No ratings yet8 installsv1.0.0Updated Oct 6, 2026● Unknown

Free · MIT

About

Designs and analyses A/B tests on pages, emails, ads and flows: an evidence-based hypothesis, one primary metric with guardrails, sample size and run time from the real baseline, randomisation and QA, stopping rules, a sample ratio check, and a plain-English readout with confidence intervals. Also says when traffic is too low to test. Use when asked to A/B test something, which version won, or how long to run a test. Not for finding what to test (use funnel-analysis or landing-page-review).

Marketing

Documentation

From SKILL.md · v1.0.0 · what the agent reads when it loads this skill3 files: SKILL.md, references/CREDITS.md, references/test-plan.md

A/B Test Design

You run experiments that give answers the owner can trust. That means deciding before launch what will be measured, how many people are needed and when the test ends, then analysing exactly that. The honest result is often "no detectable difference", and you say so. When traffic is too low to detect a realistic effect, you say that before anyone wastes a month.

When to use

  • "Can we A/B test the new headline?"
  • "Which subject line won?"
  • "How long should this test run?"
  • "Is this result significant?"
  • "Set up an experiment programme."

What you need

  • The change and the reason for it (evidence from analytics, research, user feedback).
  • Baseline numbers for the primary metric: rate and weekly volume of eligible visitors, recipients or users. From Google Analytics, PostHog or Mixpanel (connected apps), or the email or ad platform via the owner's browser.
  • The testing tool the owner already has: the website testing or feature-flag tool, the email platform's A/B feature, or the ad platform's experiments feature. If there is none, ask via clarify; do not require a paid tool.
  • A developer or the page owner for implementation, if needed (ask_teammate).

Method

  1. Write the hypothesis. "Because we saw [evidence], we believe [change] for [audience] will [increase/decrease] [primary metric]. We will know when [metric] moves by at least [minimum detectable effect] after [sample size]." No evidence, no test: go and get some first.
  2. Choose one primary metric tied to value (purchase, signup, qualified lead, click on the main CTA), plus 2 to 3 guardrails that must not get worse (revenue per visitor, refund rate, unsubscribes, page load time). Secondary metrics explain results; they do not decide them.
  3. Pick the minimum detectable effect (MDE). The smallest relative change worth acting on, and realistic for the change: copy tweaks rarely move conversion by 20 percent; a new offer might.
  4. Calculate sample size and duration with the function in references/test-plan.md (two-sided, 5 percent significance, 80 percent power by default). Duration = total sample / weekly eligible traffic, rounded up to whole weeks, minimum 1 full week and preferably 2 business cycles. As a guide, a 5 percent baseline needs about 31,000 visitors per variant to detect a 10 percent relative lift, and about 8,200 per variant for a 20 percent lift.
  5. Decide whether to test at all. If the test would take more than about 6 to 8 weeks, do not run it as planned. Options: test a bolder change (larger MDE), test on a higher-traffic page or an earlier funnel step, use a different metric with a higher base rate, or make the change based on qualitative evidence and monitor before-and-after, labelled as weaker evidence.
  6. Design the variants. Control plus one variant is the default. More variants need proportionally more traffic. Change one idea per variant (a whole redesign is one idea; testing it tells you whether the redesign wins, not why).
  7. Randomise properly. Split at the right unit (visitor or user, not pageview), 50/50 unless risk justifies less exposure, sticky assignment so people see the same version. Exclude internal traffic and bots where possible. For emails, split randomly within the same segment and send at the same time.
  8. QA before launch. Both versions render on mobile and desktop, tracking fires for both, assignment is roughly 50/50 in a short smoke test, no flicker of the original before the variant appears.
  9. Pre-register stopping rules. Run to the planned sample and full weeks. Do not stop early because a dashboard shows "95 percent significant" on day 3; repeated peeking inflates false positives. Stop early only for broken implementation or a guardrail clearly harmed. If the tool uses sequential or Bayesian statistics designed for continuous monitoring, follow its rules and state them.
  10. Check sample ratio mismatch first. At the end, test whether the split matches the design (chi-square, code in references/test-plan.md). A p-value below 0.001 means assignment or tracking is broken: do not trust the result; find the cause.
  11. Analyse. Report each variant's rate, the absolute and relative difference, the 95 percent confidence interval of the difference and the p-value; then guardrails. Interpret in plain English: "B increased signups by an estimated 12 percent; the true effect is likely between 3 and 21 percent." A p-value is not the probability that B is better.
  12. Decide and document. Ship, iterate or drop. Record the hypothesis, setup, results and decision in a test log (template in references/test-plan.md), save the learning in memory, and tell the team with brief_team. Segment results (device, new vs returning) are clues for new hypotheses, not conclusions, unless they were planned up front.

Output

  • Test plan (before launch): hypothesis, metrics, MDE, sample size, duration, split, QA list, stopping rules, owner.
  • Readout (after): SRM check, results table with intervals, guardrails, decision, next test.
  • Test log entry. Show the readout on a show_card.

Checks before you finish

  • Hypothesis cites evidence; primary metric and guardrails were fixed before launch.
  • Sample size came from the real baseline and a stated MDE.
  • The test ran whole weeks and reached its planned sample (or the reason it stopped is documented).
  • SRM check passed.
  • Results include confidence intervals, not only "significant / not significant".
  • Claims about segments are labelled exploratory unless pre-planned.
  • Changes to the live site or sends to customers were approved by the owner.

Pitfalls

  • Peeking and stopping early. The most common way to ship false winners. Fix the sample in advance.
  • Testing with too little traffic. Underpowered tests produce noise. Do the maths first and change plan if needed.
  • Too many metrics. With twenty metrics, one will move by chance. One primary metric decides.
  • Ignoring novelty and seasonality. Run full weeks; be wary of tests over holidays or during campaigns that change who visits.
  • Counting clicks when the goal is revenue. A variant can win clicks and lose sales. Use guardrails.
  • Declaring "no effect". A null result means no effect large enough to detect was found. Report the interval.
  • Changing the test midway. Editing a variant or the split resets the experiment.

Versions

v1.0.0currentOct 6, 2026

Listed from the source repository.

Reviews

No reviews yet. Be the first.

Write a review