All articles

Experience Metrics / UXTap Journal

SUS: how to use the usability scale without confusing the score

Understand the SUS structure, result calculation, and necessary precautions to compare versions without losing the experience context.

A paper questionnaire with response scales represents the SUS usability evaluation.

SUS, an acronym for System Usability Scale, offers a standardized way to record usability perception after interacting with a system. It helps summarize an evaluation but does not, by itself, identify which part of the interface needs correction.

The original structure comprises ten items, with responses from 1 to 5 and alternating favorable and unfavorable statements.

Prepare the experience before the questionnaire

Define which system or version will be evaluated and what activities participants will perform. A person who only saw a presentation had a different experience from someone who tried to execute tasks.

Imagine a fictitious reimbursement tool. The study might ask professionals to record an expense and track its status. After these activities, the scale records how they perceived using the system.

Avoid teaching all steps before the test if the intention is to evaluate autonomous discovery. This training changes the experience the person will consider when responding.

Understand the calculation

In the original structure, responses range from 1 to 5. For odd-numbered items, 1 is subtracted from the response. For even-numbered items, the response is subtracted from 5. The sum of these contributions is multiplied by 2.5.

The result is between 0 and 100, but it is not a percentage of completed tasks. In a hypothetical example, contributions totaling 28 produce a SUS score of 70. This does not mean the person correctly executed 70% of the flow.

Avoid inadvertently creating a different scale

Swapping a negative phrase for a positive one requires attention to interpretation. Removing items, inverting labels, or mixing scales modifies the instrument.

In UXTap, SUS items can be configured as a metric battery. The results aggregate the score and present the item readings according to their direction. Check if each question is correctly identified and if the battery is complete.

Conduct a pilot with known responses to review the configuration. Pay special attention to negative items, where agreeing has a different meaning than agreeing with a favorable statement.

Read the average alongside the experience

An aggregated score can hide distinct experiences. Professionals who have used similar systems might find less difficulty than people without that repertoire.

Consult the distribution and relevant segments, keeping the base size visible. If the sample is small, do not treat a modest difference between averages as sufficient proof of improvement.

Also, avoid using general references as mandatory goals for any product. The decision depends on the audience, tasks, system maturity, and the cost of observed problems.

Use tasks to explain perception

In the reimbursement example, an unfavorable evaluation might be linked to the difficulty of understanding which documents to attach. Observe the sessions and consult the reports before concluding that the entire form needs to change.

SUS offers a summary of perception. Clicks, paths, and responses help locate opportunities. The combination allows for formulating a more specific proposal and checking if it improves the experience.

Check a complete response before looking at the average

Use a fictitious response to check the battery configuration. Consider scores 4, 2, 4, 2, 5, 1, 4, 2, 4, and 2, in that order, for the ten original items. The adjusted contributions are 3, 3, 3, 3, 4, 4, 3, 3, 3, and 3. The sum is 32, and the individual result is 80. This is a calculation example, not a research result.

The study average is calculated from valid individual scores. Do not sum all responses without preserving the association between person and battery. Also, do not interpret the response to an isolated phrase as an independent SUS. The instrument combines items to form the global measure.

In UXTap, a session's score requires all ten items and numerical responses. An incomplete battery does not receive an invented score to fill the gap. When presenting the result, check how many people effectively contributed to the SUS, which may have a different base than that of a previous task.

Record the conditions that give meaning to the score

In the fictitious reimbursement tool example, some people might have only used expense submission; others might have managed rules and approved requests. The product name is the same, but the evaluated experience is different. Define which set of tasks precedes the scale and keep this information alongside the result.

Also record familiarity, training offered, and material limitations. A guided presentation can help understand a proposal but produces another evaluation condition. If the first round had guidance and the second was autonomous, do not attribute a score difference solely to the product version.

The distribution helps perceive experiences that the average hides. An intermediate evaluation might gather people with similar perceptions or very distinct groups. Read the reports and tasks of these groups before writing a single conclusion about the interface. Preserve the limits of segments with few responses.

Transform a score into investigation questions

If the usability perception is unfavorable, start with the obstacles recorded during the activity. In the reimbursement tool, look for differences between understanding which documents are accepted, being able to attach them, and tracking the submission. Each difficulty can contribute to the evaluation, but the scale does not automatically identify its cause.

Prepare an action linked to the evidence: review information about documents before attachment, for example. Then, evaluate if the task became more comprehensible and if the global perception changed under comparable conditions. The goal is not to edit the interface until a number surpasses a chosen reference without context.

In the report, present the version, tasks performed, number of valid batteries, and comparison limitations. If there is an external reference, describe what it represents and avoid treating it as a percentage of satisfied people. The executive report guide helps link the metric to product decisions.

Plan the comparison before altering the product

Keep tasks, audience, and conditions as consistent as possible between rounds. Record inevitable changes and consider their effect on the reading.

To start, set up a relevant task in UXTap and add the complete SUS battery. Review configuration and scale in the preview. After data collection, choose an observed obstacle and plan its correction; use the next round to examine both behavior and usability perception.

To track a revision, plan how to compare versions and explore the results resources, preserving the conditions of scale application.

Frequently asked questions about SUS

Does a score of 80 mean 80% usability?

No. SUS uses a transformation of responses to a scale from 0 to 100. This value is neither a percentage of completed tasks nor a proportion of satisfied customers. Present the score as a measure of global perception and use separate behavioral measures to describe execution.

Can I remove questions that seem repetitive?

Removing or rephrasing items alters the instrument and compromises comparability with applications of the original scale. Use the complete battery and a language-appropriate version. If you need a shorter investigation, consciously choose another instrument instead of calling a selection of items SUS.

How to handle a missing response?

Do not silently complete the score with an average or a neutral value chosen by the researcher. Record the valid base and check why the question was left unanswered. In UXTap, a session without the complete numerical battery does not generate a SUS score; this should be visible in the sample reading.

Create your UXTap account to prepare your first study.

Your next discovery begins with a question.

Test an idea, observe the experience and bring evidence to your team.

Create my first study

Free · 1 study · 50 responses/month · no card