// AI MARKETING OPERATING MODEL

The Test Before the A/B Test

How much time does your team spend deciding what to A/B test before the A/B test even starts?

That work rarely appears in the media plan, but it can decide whether the budget buys an answer or just another debate. A technically valid test still wastes budget when nobody can explain why the question matters, what result would change the team's mind, or whether the answer will arrive in time to affect a decision.

The cheapest test is the one you reject on paper before it buys a single impression. The harder question is how to reject a weak idea without killing a good one because the brief was sloppy.

This is spoke two in my AI Marketing Operating Model series, which looks at the decisions connecting marketing, revenue, and AI.

What should happen before an A/B test launches?

Write the belief, the one variable that changes, what stays fixed, the audience, the success metric and the stop rule. Then check whether the idea deserves budget at all.

Pre-test validation is the design check before an ad, email, or page variation enters rotation. It is also where a team may discover that its “one test” is really three changes bundled into one campaign brief.

The six things a test needs before launch

A clean test brief should fit on one page. Not because the work is small, but because the decision should be understandable without a meeting to translate it.

  1. 1. The prior beliefWhat do you expect to happen, and why? Write the reason before looking at the result. “The new version is better” is not a belief. “The current page buries the implementation proof, so moving that proof above the form may increase qualified demo requests” is testable. It gives the team something to learn even if the challenger loses.
  2. 2. The changeWhat is the one meaningful difference between control and challenger? Name the audience-facing change, not every asset touched to make it possible. A new offer, message, page structure, and form length all changing at once might produce a result, but the result will not tell you which change mattered.
  3. 3. What stays fixedKeep the offer, audience rules, budget, placement, timing, eligibility, and follow-up process stable where the design allows it. If something must change for operational reasons, record it. The goal is not laboratory purity for its own sake. It is being able to say what the comparison can and cannot tell you.
  4. 4. The eligible audience and unitWho can enter the test, and what gets assigned: a visitor, account, household, campaign, or time period? Define exclusions before the first exposure. For B2B, two people at the same account seeing different versions can contaminate a comparison. For some channels, random assignment may not be available at the person level, so a geo, time-based, or matched-market design may be more appropriate.
  5. 5. The primary metric and guardrailPick the one primary outcome that will decide whether the test answered its question. Then choose the measure that would make a nominal win unacceptable. A landing-page test might use qualified meeting requests per eligible account as its primary metric and form completion quality or sales capacity as guardrails. Click-through rate can help diagnose the path, but it may not be the business outcome you are trying to improve.
  6. 6. The stop ruleDecide how the team will know the test has enough information to read, and what would make it stop early for harm. That is not the same as checking a dashboard every morning and ending the test when a chart looks favorable. Set the sample or duration using the expected traffic, baseline rate, variability, and smallest effect worth detecting. If volume is too low, say so before launch and choose a different design or decision. A calendar deadline does not create statistical power.

For many straightforward A/B tests, changing one main variable makes the result easier to interpret. That is a useful default, not a law. Factorial designs can test multiple factors deliberately, and tests across channels may need to account for interaction and spillover. The design has to match the question. Putting A and B side by side is not enough.

If the business outcome is pipeline, write down how the event will be reconciled in the CRM and, where relevant, Finance. I cover that measurement chain in Proving AI Pipeline to Your Board.

Do not confuse a stop rule with a stopping date

A stopping date is when the team plans to look. A stop rule is the decision logic that says what evidence is enough, what result counts as harm, and what happens if the answer remains unclear.

Repeated peeking can change the chance of a false positive in some fixed-horizon statistical tests. Statsig's current experimentation FAQ recommends selecting a readout date before launch based on power analysis and warns against repeatedly checking for significance before that date. Different statistical methods handle sequential monitoring differently, so the rule should match the method the team is actually using.

There are usually two separate decisions. First, should we pause because a safety or business guardrail is being breached? Second, at the planned readout, did the test provide enough evidence to roll out, reject, or run another test? If those are blurred, an emergency pause can get misread as a result, or a weak result can be dressed up as a win because the deadline arrived.

“Run it for two weeks” is not a complete stop rule. Two weeks may be too short for the test to answer the question, or longer than needed. The amount of traffic, the baseline behavior, the outcome's variance, the minimum effect worth acting on, and the chosen analysis method all matter. Eppo's sample-size guidance spells out the relationship between sample size, a metric's mean and variance, and the effect a test can detect.

Put a budget gate in front of the test

Before launch, ask four blunt questions.

  • If the challenger wins, what decision changes?
  • If it loses, what will the team stop doing or learn?
  • Can the planned audience and measurement setup distinguish between those outcomes?
  • Could the business absorb the result without breaking sales, service, margin, or customer trust?

If the answer to the first two is “nothing,” the test may be interesting but it does not yet deserve paid distribution. If the answer to the third is no, buying more traffic is not automatically the fix. The metric, audience, assignment method, or question may need to change. If the fourth answer is no, the test needs an operational guardrail before it needs more impressions.

This gate does not mean every idea has to clear a spreadsheet full of forecasts. Some tests are cheap enough to run as exploratory work. Label them that way. Keep exploration separate from confirmatory evidence, and do not turn an early directional signal into a company-wide claim.

What the public preflight walkthrough shows

The live Marketing Experiment Preflight walkthrough starts with the question behind this article: how much time does a team spend deciding what to test before the A/B test begins?

It walks through a fictional Company X marketing and revenue decision across paid search, paid social, lifecycle email, a landing page, and sales follow-up. The example keeps the overall budget and capacity boundary in view while changing the proposed journey. It asks the team to define the control, challenger, audience, objective, evidence, primary outcome, guardrails, and what would reverse the recommendation.

The walkthrough is a planning example, and its audience context is aggregate. It does not show individual people or measured customer response. It shows the work I want done before a real test uses live media: put the actual decision on screen, make the assumptions legible, find missing definitions, and decide what must stay fixed.

I built the page because test planning is often scattered across a brief, a spreadsheet, a Slack thread, and somebody's memory. That creates friction before launch and makes it harder to understand the result later. A preflight is useful if it helps the team make the question sharper, not if it adds another approval ceremony.

The page can help a team compare directions, check its design, and surface questions to resolve. Only an actual test with valid assignment and observed outcomes can tell the team how customers responded. The walkthrough makes no claim about campaign lift, conversion improvement, or revenue prediction.

AI makes the gate more valuable

AI can generate headline variants, audience hypotheses, landing-page copy, email sequences, and creative concepts faster than a team can approve them. Drafting options gets cheap. Media distribution, measurement, and buyer attention still cost money.

When producing ten alternatives takes minutes, teams can end up with more things to test than they can measure well. The bottleneck moves. It becomes deciding which question has commercial consequence, what evidence would answer it, and whether the organization can act on the answer.

A model can also make a weak hypothesis sound researched. Fluent copy does not establish that the audience has the problem, that the offer is credible, or that the event in analytics represents value to the business. The brief still needs a human owner who can explain why the test matters and who can stop it if the consequences change.

Use AI to expand the option set, inspect the brief for missing assumptions, propose possible guardrails, and help make the variants consistent. Keep the evidence and decision rule outside the model's prose. A model-generated rationale is not a preregistered hypothesis unless the team adopts it before the result is visible.

I do not want a pre-test tool announcing a winner. It should help the team become more specific before launch. If the test still matters after that review, run it with a suitable design, collect the real outcome, and keep the record so the next decision can use what this one taught.

Pre-test validation is a decision, not another meeting

Pre-test validation is a quick decision: does this test have a question worth answering, an audience you can reach, a metric tied to a business decision, and a stop rule the team will honor?

Statsig's current experiment-creation docs require a hypothesis and at least one primary metric as part of setup. Eppo's experiment-protocol docs describe standardizing metrics, analysis methods, and decision criteria for a class of experiments. The products differ, but the operating point is similar: decide how the experiment will be interpreted before the result is in front of you.

Kohavi, Tang, and Xu's 2020 book remains a useful field reference. Current platform docs turn that discipline into setup choices: primary outcomes, guardrails, and decision criteria.

Marketing teams do not need everyone to become a statistician. They do need to stop spending money on tests where nobody can say what would change if the result went either way.

If neither result changes a decision, do not buy the test.

Frequently asked questions

What should happen before an A/B test launches?

Write the belief, the one variable that changes, what stays fixed, the audience, the success metric and the stop rule. Then check whether the idea deserves budget at all.

How do you decide if a test idea deserves budget?

Fund a test when either result could change a real decision, the audience and measurement design can answer the question, and the business can act on the outcome. If neither a win nor a loss would change anything, the test is not ready for paid distribution.

What is a stop rule in marketing experimentation?

A stop rule defines when the team will read the result, what evidence is sufficient for a decision, and what guardrail breach requires a pause. It should match the test's statistical method and be set before the team sees a favorable or unfavorable result.

How many variables should an A/B test change?

A simple A/B test should usually change one main variable if the goal is to understand that change. A planned factorial or multivariate design can test several factors, but the design must account for how those factors are assigned and interpreted.

Does AI change how marketing teams should prioritize tests?

AI makes it faster to create more variants, so the team has to be more deliberate about which question is worth measuring. It does not remove the need for valid assignment, a meaningful metric, a guardrail, or observed customer outcomes.


Sources

Experiment setup, hypothesis, scorecard, and primary metrics: Statsig, Create an Experiment (documentation checked September 29, 2026). Readout timing and repeated significance checks: Statsig, Experiment results FAQ (checked September 29, 2026). Standardized analysis methods and decision criteria: Eppo, Creating an experiment protocol (checked September 29, 2026). Guardrail thresholds: Eppo, Guardrail cutoffs (checked September 29, 2026). Sample size and detectable effects: Eppo, Sample size calculator usage (checked September 29, 2026). Historical reference: Ron Kohavi, Diane Tang, and Ya Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020.


About the author

Jeff Brokaw is a CMO and Certified Chief AI Officer who works on the commercial systems around AI: the evidence, access, measurement, capacity, and accountability that turn a model into a useful business tool.

Have a test that needs a better question?

Start with the decision, the guardrail, and the thing that would make you change your mind.