A new version might seem clearer to the team yet still introduce difficulties for the audience. Comparing research rounds helps examine the effect of the change, provided the conditions are visible. If the audience, task, and material change simultaneously, it becomes difficult to explain the result.
Start by recording what was changed and what behavior you expect to improve. “The new version will be better” does not offer an evaluation criterion. “The task entry will be easier to find” allows for preparing a specific comparison.
Differentiate Version Comparison and Experiment
In UXTap, published versions and segments allow for examining results side-by-side. This feature organizes the reading but does not guarantee random distribution between versions nor eliminate recruitment differences.
Document how each group arrived at the study. When data collection occurs at different times, consider changes in the product, familiarity, and participant origin.
Choose the Main Metric Before Looking at the Numbers
Imagine a fictional application where customers change the delivery address. The team moved the action to the order details.
The main criterion can be completing the change. Time, path, and comments help explain the experience but do not need to carry the same weight in the decision.
Also define what constitutes a useful difference. A small change in a measure might not justify a broad product change, even when statistical evidence is available.
How to Compare Versions Without Mixing Conditions
Use the same scenario and equivalent data. Do not simplify the wording of the new version or provide additional instructions to one group.
If the same people experience both proposals, consider learning effects. The first attempt might teach where to look or clarify a business rule. Plan the order and record this condition.
When the goal is to evaluate initial discovery, participants without prior exposure usually offer a condition more aligned with the question. The composition still needs to be planned according to the study.
Read Base, Distribution, and Segments
In UXTap, the comparison can show differences in metrics by block and statistical tests where compatible evidence exists. Not every measure receives the same automatic analysis.
Check the number of responses for each group and which sessions are included in the distributions. The treatment of partials must be consistent across versions.
Investigate relevant segments, such as device or experience. Avoid looking for dozens of cuts until you find a favorable difference. If a hypothesis emerges during the analysis, record it as exploratory and plan its verification.
A Numerical Result Needs Explanation
In a hypothetical example, success goes from 12 out of 20 attempts to 15 out of 20. The observed difference is 15 percentage points, but these numbers alone do not demonstrate that the change caused the improvement.
Consult the unsuccessful attempts. Perhaps the new entry helped locate the action, while a later step remains confusing. The recommendation might be to keep the location and revise the confirmation.
Also examine potential losses. A change that helps new users might hinder a frequent routine for another group.
Write a Short Protocol Before the New Round
Revisit the address change in the fictional application. Record the hypothesis: placing the action in the order details should facilitate its discovery in the context of a delivery. The main measure can be completing the correct change. First choice, returns, and interpretation of the confirmation serve as evidence to explain the result.
Define what will remain the same: scenario, fictional address, order status, intended device, and recruitment criteria. Separately list what changed in the interface. If the new version also has more complete information or a shortened flow, the comparison evaluates that set; it will not be possible to attribute all differences to the button's position.
Decide in advance which sessions will be considered valid. A loading failure might require different treatment than a person who tried and gave up. Do not exclude difficult attempts from only one version. The rule must be applied consistently and remain available for anyone reviewing the result.
Choose Between Different Participants and Repetition with the Same People
Different people prevent the first attempt from directly teaching the second, but groups may have distinct experiences. Compare their composition and record the recruitment. A group of frequent customers does not automatically equate to another composed of people who have never used the product.
Repeating with the same people allows observing individual changes but introduces learning and memory. Someone might find the action faster because they already know the task. If this design is appropriate, plan the order and acknowledge this limitation; do not present the second attempt as an independent initial discovery.
A comparison between rounds is also subject to period differences. Campaigns, audience changes, and external events can influence usage. The dashboard organizes the available results, but the research design remains the team's responsibility. Maintain this distinction in the report.
Separate Difference Size and Conclusion Strength
In the fictional example of 12 successes out of 20 attempts versus 15 out of 20, the observed difference is 15 percentage points. The next question is how much uncertainty exists and what other changes might explain the result. Presenting only an upward arrow hides these issues.
When compatible statistical analysis is available, read it alongside the base sizes and study design. An indication of difference does not correct for incompatibly recruited groups. The absence of an indication also does not prove that the versions are identical; the study might not allow for a sufficiently precise conclusion.
Also discuss practical relevance. A change might aid discovery and worsen the understanding of who will receive the order. Another might reduce time only among those who managed to complete it, while more people give up. Define which effects matter for the decision and avoid choosing only the favorable measure. The combined reading of SUS and tasks can add usability perception when the instrument is appropriate.
Decide and Preserve the Reasoning
Write what the comparison allows to affirm, what doubts remain, and what action will be taken. If the evidence is inconclusive, record the reason and the next necessary design.
To start, choose a small change and an important task in UXTap. Define the main criterion, preserve comparable conditions, and plan the analysis before data collection. Then, combine the numbers with sessions and reports to decide what to keep, revise, or investigate again.
Upon completing the comparison, prepare a UX executive report and consult the results resources to present differences alongside the evidence.
Frequently Asked Questions About Version Comparison
Does Comparing Two Studies Equate to an A/B Test?
Not necessarily. An experiment requires design and control of conditions, including how people arrive at the alternatives. Placing results side-by-side allows comparing what was observed, but it does not retroactively create a random distribution nor eliminate differences in audience and context.
Can I Compare a Task Whose Text Was Changed?
You can analyze both rounds, but record that the instrument changed. A clearer or more specific instruction can alter the difficulty independently of the interface. If the text change is relevant, avoid treating the results as directly equivalent measures of the same problem.
What to Do When No Version Clearly Wins?
Describe what the study allows to affirm and what remains uncertain. Use the sessions to locate concrete advantages and limitations and consider implementation cost, risk, and decision reversibility. Another round might be necessary, but it should resolve a specific doubt, not just repeat the dispute until a favorable number appears.
Create your UXTap account to prepare your first study.
Your next discovery begins with a question.
Test an idea, observe the experience and bring evidence to your team.
Create my first studyFree · 1 study · 50 responses/month · no card
