Interpreting traditional test results
Start with the question you tested
Once your traditional test has been running, its report brings together the traffic, conversions, and statistical evidence you need to understand how each variation is performing.
Start by returning to the question behind your test.
For our homepage example: Did moving the primary CTA above the fold increase Trial starts for mobile visitors?
The report helps you answer that question by showing how the base and each variation performed against the goal you selected.
Read the results table
The results table gives you the clearest snapshot of how each variation is doing.
[IMAGE: Traditional test results table with Base and Variation A. Call out Sessions, Conversions, Conversion rate, and Statistical significance.]
Here are the main metrics to look for:
- Sessions: How many browsing sessions were attributed to each variation.
- Conversions: How many times visitors completed the action tracked by the selected goal.
- Conversion rate: The number of conversions divided by sessions for that variation.
- Statistical significance: How much evidence there is that a variation’s performance compared with the base reflects a real difference rather than normal fluctuation.
For example:
Base: 1,000 sessions · 82 conversions · 8.2% conversion rate
Variation A: 1,000 sessions · 101 conversions · 10.1% conversion rate
At first glance, Variation A looks stronger.
But a higher conversion rate doesn’t automatically mean you have a winner. You also need enough evidence to trust that the difference isn’t temporary or due to chance.
Read statistical significance
That’s where statistical significance comes in.
In a traditional test, Optimize compares each variation with the base and calculates how reliable that difference appears to be. Higher statistical significance means the results are more stable; lower statistical significance means they’re more likely to change as additional data comes in.
[IMAGE: Early test with fluctuating results → later test with more stable results.]
Three things can affect statistical significance:
- Time: Visitor behavior varies, so a test needs time to account for short-term spikes or unusual periods.
- Sample size: More traffic and conversions gives Optimize more evidence to compare.
- Effect size: A large, sustained difference between a variation and the base may become clear sooner than a very small difference.
When the test reaches high statistical significance, Optimize can identify a winning variation. Until then, treat an early lead as just that — an early lead.
Look beyond the top-line result
The variation table is the starting point, but the rest of the report can help you understand the result in more context.
Switch between goals
If your optimization tracks more than one goal, the report shows the target goal by default. You can switch goals to see how the same variations performed against your supporting goals.
For example, your CTA test might show:
- More CTA clicks
- More pricing-page views
- No meaningful change in Trial starts
That tells a more complete story than the target goal alone.
[IMAGE: Goal selector switching between Trial starts, CTA clicks, and Pricing-page views.]
Look at performance over time
The performance chart shows how conversion rates changed while the test was running. The traffic allocation chart shows how much traffic each variation received over the same period.
[IMAGE: Performance-over-time chart beside traffic allocation.]
Use these views to spot trends or understand whether a particular event — like a campaign or promotion — coincided with a change in performance.
Explore audience insights
Audience insights break conversion rates down across different visitor groups, such as device type, traffic source, or other audience criteria.
[IMAGE: Overall test result → Audience insights showing stronger performance for one segment.]
A variation might not win overall but still perform differently with a particular group. That doesn’t change the result of the test, but it can give you a useful idea for what to investigate or personalize next.
Note: Audience insight cells with fewer than 300 sessions remain empty to reduce noise.
Decide what to do next
Once you understand the report, bring it back to the hypothesis you started with.
You might:
- Apply the result: The test has enough evidence to support the winning variation.
- Keep running: There isn’t enough evidence yet.
- Iterate: The result raises a new question worth testing.
- Stop: The change isn’t producing a useful result.
And if a variation loses or the test stays flat, don’t just discard it. Ask what the result taught you.
- Did visitors have enough opportunity to notice the change?
- Was the change meaningful enough to influence behavior?
- Or did the hypothesis simply not hold?
A traditional test doesn’t just tell you which variation performed best. It gives you evidence you can carry into the next idea on your roadmap.
Ready to keep going?
You now know how to read the core metrics in a traditional test report, use statistical significance to judge whether the result is trustworthy, and dig into supporting goals and audience insights for more context.
Next, you’ll learn how to interpret manual personalization results, where the goal isn’t to find one winning variation at all.