System Usability Scale (SUS)

The System Usability Scale (SUS) is a post-test attitudinal survey measuring perceived usability. It provides a numeric score, so it can be used for benchmarking, relative comparisons of different systems, to evaluate the effects of changes to a system, or to track trends and changes over time and to detect unexpected impacts.

SUS usually takes a few minutes to complete, and can be easily scored by hand. Academic review suggests that it can detect differences as accurately as questionnaires two or three times longer. With about forty years of industry usage, it became a de-facto standard survey for usability. For example, Berger’s review of usability research in healthcare found it being used in 92.8% of the cases over 26 years, compared to 6.0% for the next most common instrument. Due to extensive usage, academic research and published benchmarks, SUS is a standard reference point for validating other questionnaires.

This article explains how to run, score and interpret System Usability Scale surveys, and where the method applies and where it does not.

sus score testing template

SUS Questionnaire

SUS consists of ten questions, each with a five-point Likert Scale. The statements are deliberately redundant to detect inconsistent answers and reduce acquiescence bias.

The SUS survey template alternates between redundant positive and negative statements

1. I think that I would like to use this system frequently

1 2 3 4 5
Strongly Disagree Strongly Agree

2. I found the system unnecessarily complex

1 2 3 4 5
Strongly Disagree Strongly Agree

3. I thought the system was easy to use

1 2 3 4 5
Strongly Disagree Strongly Agree

4. I think that I would need the support of a technical person to be able to use this system

1 2 3 4 5
Strongly Disagree Strongly Agree

5. I found the various functions in this system were well integrated

1 2 3 4 5
Strongly Disagree Strongly Agree

6. I thought there was too much inconsistency in this system

1 2 3 4 5
Strongly Disagree Strongly Agree

7. I would imagine that most people would learn to use this system very quickly

1 2 3 4 5
Strongly Disagree Strongly Agree

8. I found the system very awkward to use

1 2 3 4 5
Strongly Disagree Strongly Agree

9. I felt very confident using the system

1 2 3 4 5
Strongly Disagree Strongly Agree

10. I needed to learn a lot of things before I could get going with this system

1 2 3 4 5
Strongly Disagree Strongly Agree

The original phrasing of question 8 used the word “cumbersome” instead of “awkward”, but it turned out to be difficult for many participants to understand. Kraig Finstad evaluated the comprehension, concluding that half of non-native English speakers had trouble understanding it. Bangor, Kortum and Miller reported that about 10% of participants had a question about the meaning of “cumbersome” over two decades of using SUS at AT&T. Recent books such as Quantifying the User Experience show the “awkward” version, and Bangor, Kortum and Miller note that approximately 90% of the data in their own normative study, which is one of the three data sets used for the benchmarks in this article, was collected with that form.

To make the survey more relatable, researchers commonly replace the word “system” with “product”, “website” or “application”, or the name of the product. John Brooke’s research confirms that rephrasing the “system” does not impact results. To ensure consistency, use the same term in all ten questions. Russ-Jara and colleagues suggest that mixing terms in a single questionnaire undermines the comparison of scores.

The canonical SUS questionnaire alternates between positive and negative phrasing to detect if participants pay attention to questions, but this effect may not be statistically noticeable. Jeff Sauro and James R. Lewis ran an experiment with all five negative items phrased positively, and Philip Kortum, Claudia Ziegler Acemyan and Frederick Oswald replicated the comparison across 20 products. Neither experiment found a meaningful difference between the variants. The difference between means was just 0.68 points, with a 95% confidence interval running from −0.72 to 2.08. The reliability for the positive version was slightly higher. The most important change when presenting the positive SUS variant is to adjust the scoring, so all questions are rated on a positive scale.

Some of the ten questions might not make sense in a specific situation. For example, the first question asks about frequent usage, and it may not be meaningful for products such as tax submissions that are intended to be used once per year. Lewis and Sauro tested removing individual questions from 9,156 completed SUS questionnaires from 112 unpublished industrial studies and surveys. They found that the nine-item variants sometimes move the scores marginally (some scores fell from B to B−), but that the internal consistency for all tested variants was above 0.90, and all the changed options correlated above 0.99 with the full instrument, with mean differences within one point. Their conclusion was that, if necessary, any single item could be omitted, but “unless there is a strong reason to do otherwise, use the full 10-item version of the SUS”. Critically for scoring, removing a question requires adjusting the multiplier from 2.5 to 100/36.

The original questionnaire is in English, and most of the benchmark data and academic research is related to it. Translated versions are easy to find online, but may not be as reliable as the English one. Lewis reported a mean reliability of about 0.81 across published translations, lower than the 0.91 typical for the original version. Professionally translated versions of SUS that have been psychometrically validated are available for Chinese, French, Spanish, Arabic and German. Blažica and Lewis present a good model for evaluating a translation in A Slovene Translation of the System Usability Scale: The SUS-SI, showing how to report coefficient alpha, concurrent validity, and sensitivity to frequency of use.

sus scale

Origins of SUS

John Brooke created SUS in the mid-1980s while working at the Digital Equipment Corporation, and published the survey structure in the 1996 compilation book Usability Evaluation In Industry, describing it as a “quick and dirty usability scale”. SUS evolved with usage into a de-facto standard, and according to Brooke “has never been through any formal standardization process”. The survey was based on Kirakowski and Corbett’s 1988 Computer User Satisfaction Inventory, but simplified to evaluate perceived usability and not effectiveness or efficiency.

According to Andrew James Holyer, the great advantage of SUS was its “astonishing ease of use” compared to other usability questionnaires at the time.

Fitting easily on a single side of paper, it is quick and easy to administer, and since it is less intimidating than longer questionnaires a far higher percentage of users are likely to actually complete it - a serious issue when you consider that many more “serious” questionnaires constitute a substantial document in their own right. It is also extremely quick and easy to score… The SUS can be scored by hand in around 30 seconds.

Andrew James Holyer, Methods for Evaluating User Interfaces

Holyer’s paper claims that University College Cork evaluated SUS and compared it to their much more comprehensive 50-question survey Software Usability Measurement Inventory, concluding that the results correlated with a reliability of 0.86, “which considering that it is far simpler and cheaper to administer is very impressive” (that figure was presented with no sample size, comparison method description or further references).

sus score calculator

Calculating the SUS score

Each survey response produces a numerical score, on the scale of 0 to 100. Lewis provides an aggregate formula for the score in The System Usability Scale: Past, Present, and Future:

SUS = 2.5 × (20 + SUM(items 1,3,5,7,9) − SUM(items 2,4,6,8,10))

The score requires all ten items. For partially filled-in surveys, Lewis suggests using the value 3.

To understand the formula, consider that options for even and odd rows alternate in meaning, and that the participants select values between 1 to 5. Each row contributes to the final score differently. For positive questions (odd rows), subtract 1 from the selected option, so they contribute on the scale from 0 to 4. For negative questions (even rows), subtract the answer from 5, so they contribute on the scale from 4 to 0. Add up all the contributions, and this will produce a number between 0 and 40. Multiply by 2.5 to scale it up between 0 and 100.

When aggregating results, it’s critical to score each survey response first, then aggregate. Although the scoring formula itself can easily be applied to question aggregates across all participants, the SUS items are strongly intercorrelated, so a standard deviation rebuilt from question totals will be much smaller than the standard deviation from participant scores, and the confidence intervals from question totals will be incorrectly narrow.

Because SUS scores are on the scale of 0 to 100, people often misread them as relative percentages. In effect there are only 41 possible values, in 2.5-point steps, and they are not directly proportional in magnitude. A score of 40 just says that a product has serious usability issues; it does not make it half as good as a product scoring 80. The conversion to a 0–100 scale has no statistical justification, and Brooke notes it was done mostly as a convenience for easy understanding:

But why is there the rigmarole around converting the scores to be between 0 and 4, then multiplying everything by 2.5? This was a marketing strategy within DEC, rather than anything scientific. Project managers, product managers, and engineers were more likely to understand a scale that went from 0 to 100 than one that went from 10 to 50.

John Brooke, SUS: A Retrospective

Russ-Jara and colleagues warn against intuitively interpreting SUS scores, and suggest reading grades from a comparison curve instead. For example, publicly available data from Sauro’s research put the 50th percentile at about 68, and the published averages across the major data sets range from roughly 68 to 71. An intuitive interpretation of a scale of 0 to 100 would suggest that the average score is 50, which actually sits near the 13th percentile. Higher scores do get closer to intuitive understanding. For example, a score of 80 corresponds to the 88th percentile.

sus scoring

Interpreting SUS Scores

There are two schemes for interpreting SUS scores, and they are somewhat complementary.

The most popular is the curved grading scale from Quantifying the User Experience:

The curved grading scale for SUS surveys
GradeSUS rangePercentile
A+84.1–10096–100
A80.8–84.090–95
A−78.9–80.785–89
B+77.2–78.880–84
B74.1–77.170–79
B−72.6–74.065–69
C+71.1–72.560–64
C65.0–71.041–59
C−62.7–64.935–40
D51.7–62.615–34
F0.0–51.60–14

The second is an adjective scale proposed by Bangor, Kortum and Miller. They added an eleventh question asking the participants to select the general feel about the product from “Worst imaginable” to “Best imaginable”. A pilot with 212 surveys appeared in An Empirical Evaluation of the System Usability Scale, and the mapping below comes from the larger follow-up study of 959 surveys in Determining What Individual SUS Scores Mean: Adding an Adjective Rating Scale:

The adjective scale for SUS surveys, with mean SUS score and standard deviation
AdjectiveMean SUSSD
Worst imaginable12.513.1
Awful20.311.3
Poor35.712.6
OK50.913.8
Good71.411.6
Excellent85.510.4
Best imaginable90.913.4

You can use this scale to put a score into context. Because anchor means have standard deviations of 10 to 14 points, the adjective ranges overlap. The labels are also read more favourably than the scores warrant, particularly the “OK” term: the authors report that project teams took a score of OK to mean the product was satisfactory when scores in that range “were clearly deficient in terms of perceived usability”, and conclude that “It seems clear that the term OK is probably not appropriate for this adjective rating scale”. Lewis suggests, based on communication with Philip Kortum, using “Fair” instead, but such a corrected scale has never been validated.

Benchmarks for SUS

Lewis in The System Usability Scale: Past, Present, and Future suggests that absolute scores might not be that relevant, and that different classes of products would require different realistic targets. For example, “a SUS of 80 is probably unrealistically high” for many types of products, but it “is probably unrealistically low if developing a new search interface”. To put a specific score in context, it might be useful to compare it with products in its own class, and there are many publicly available benchmarks:

Kortum and Bangor created a benchmark scale in 2013 for popular products:

Kortum and Bangor's SUS scores for popular products in 2013
ProductMean SUSn
Excel56.5866
GPS70.8252
DVR74.0276
PowerPoint74.6867
Word76.2968
iPhone78.5292
Amazon81.8801
ATM82.3731
Gmail83.5605
Microwave86.9943
Web browser88.1980
Google search93.4948

The book Quantifying the User Experience has aggregate benchmarks for average scores of different product types:

SUS benchmarks from Quantifying the User Experience (2016)
Class99% confidence interval
All studies66.5–69.5
Interactive voice response75.3–84.5
Internal productivity software71.2–82.2
Mass-market consumer software69.3–78.7
Hardware (phones, modems, cards)65.2–77.4
B2B and enterprise software63.0–72.2
Public-facing websites and intranets64.4–69.6
Cell-phone equipment58.4–71.0
Web/IVR combination43.1–75.3

Note that some of these (such as Web/IVR) are based on very small samples, with huge confidence intervals.

These benchmarks were collected before mobile applications became ubiquitous. Kortum and Sorber collected SUS scores for the frequently used mobile applications:

Kortum and Sorber's SUS scores for top-10 used mobile applications by platform in 2015
Mobile platformMean SUSSD
Android phones82.715.7
Android tablets76.118.7
iPhone79.316.9
iPad72.618.9

For digital health apps, System Usability Scale Benchmarking for Digital Health Apps: Meta-analysis proposes a benchmark of 68.05 with a standard deviation of 14.05. For educational technology, Perceived usability evaluation of educational technology using the System Usability Scale (SUS): A systematic review suggests 70.09 overall, with university websites at 63.82 and multimedia applications at 76.43, and the difference between categories is significant.

Confidence intervals for SUS

Lewis and Sauro suggest reporting the mean with a confidence interval next to the scores.

If using a guide like the Sauro-Lewis curved grading scale for interpreting SUS means, include grades for the lower and upper limits of the confidence interval in addition to the grade for the mean to report the range of plausible SUS grades.

James R. Lewis and Jeff Sauro, Can I Leave This One Out? The Effect of Dropping an Item From the SUS

In Quantifying the User Experience, Sauro suggests using the t-distribution to compute the confidence interval.

To evaluate if your product exceeded a benchmark, both Quantifying the User Experience and A Practical Guide to the System Usability Scale suggest using a one-tailed one-sample t-test against the benchmark.

To compare two different designs or products, with the same group of people rating both, Quantifying the User Experience suggests working on each participant’s difference between the two scores, and putting a two-sided t-interval around the mean of those differences, with n − 1 degrees of freedom. If two different groups rated the variants, the book suggests an interval around the difference between the two group means, using the Welch–Satterthwaite unequal-variance form for the standard error.

With a large enough sample, a difference of 1.1 points can be statistically significant and still not worth acting on, because it is barely 1% of the range the scale can take. There are no established thresholds for the minimum meaningful difference between scores, but there are several rules of thumb. The narrowest bands on the grading scale are 1.4 points wide, so anything much under that cannot reliably move the score. The standard sample-size calculations are built around a ten point difference in A Practical Guide to the System Usability Scale. So a general guideline is roughly that under a point is noise, a few points is real but rarely worth acting on alone, and ten points is what you would normally expect to see with significant changes.

Check the shape of the distribution before reporting any mean at all. Hyzy and colleagues found that the SUS data from digital health applications was bimodal so “this mean score lies between 2 peaks” and may not be suitable for benchmarking.

Be careful how you average

There are two ways to combine results from several studies. You can average the means for all studies, so each study counts the same, or you can weigh the averages based on the sample size, where large studies contribute more. Vlachogianni and Tselios compared the two approaches to 170 surveys from the educational technology literature and got 70.09 unweighted against 63.30 weighted, resulting in a 6.8-point difference that is a grade and a half on the curved scale.

Their review did not suggest which approach to use, so it is a warning to pay attention to the difference more than a guide on how to resolve it.

How many people are needed for a SUS survey?

SUS is a post-test survey, using Likert scales, so many of the common suggestions for such surveys apply to SUS as well. In addition, as it’s testing usability, some of the general UX research guidelines apply.

In 10 Things to Know About the System Usability Scale (SUS), Jeff Sauro suggests that even samples as small as 5 respondents can produce relevant results. The “about five” guidance is usually taken from Robert Virzi’s 1992 experiments, as a guideline for how many users are needed to discover usability problems, not for a precise score estimation. Virzi documented that five users found about 80% of problems given the detection rates he observed, and he suggests 15 to 22 users for catching any problem affecting a tenth of the population. So running SUS with five people who all score the system badly is a good sign that something is wrong, without having to do any further research — but it will not tell you what the score is.

Tullis and Stetson sub-sampled a 123-person study and concluded that “at least 12-14 participants are needed to get” reasonably reliable results.

The book A Practical Guide to the System Usability Scale provides some more guidelines around different types of tests:

Required sample sizes for meaningful results from A Practical Guide to the System Usability Scale
What you wantParticipants needed
A ±10-point margin of error on one score, at 95%20
To detect a 10-point difference, same users on both37 in total
To detect a 10-point difference, separate groups144 (72 per group)
To prove a 5-point edge over a benchmark, at 80% power110
To prove a 10-point edge over a benchmark, at 80% power28

Applicability and limitations

Prior to 2017, a common opinion was that SUS measured two things: how easy a system was to use, and how easy it was to learn to use it. That view came from several academic reports, such as The Factor Structure of the System Usability Scale and On the dimensionality of the System Usability Scale: a test of alternative measurement models, and influenced a whole generation of reporting templates and tools. James R. Lewis and Jeff Sauro, who proposed that structure in the first place, withdrew it in 2017 after published research since 2009 had “consistently failed to replicate that Usability/Learnability” structure. Their instruction is to stop computing the subscales, and that it is best to interpret SUS as a one-dimensional measure of perceived usability.

SUS looks at perceived usability only, and does not reflect other aspects of a product such as task effectiveness (this was the intended purpose behind simplifying more complex survey instruments when designing SUS). Kortum and Peres suggest also collecting traditional ISO metrics for effectiveness, to get a more complete picture, as SUS intentionally only tracks perceived usability. Bangor has a similar suggestion in An Empirical Evaluation of the System Usability Scale: “the SUS score should not be used in isolation to make absolute judgments about the ‘goodness’ of a given product.”

In the article 10 Things to Know About the System Usability Scale (SUS) Jeff Sauro suggests that “SUS scores predicts around 40% of why customers recommend software and websites” and compares SUS scores to Net Promoter Score categories:

Detractors have an average SUS score of 67 (slightly below average usability) and Promoters have an average score of 82 (well above average usability). In independent, large datasets, we’ve seen that you can estimate the Likelihood to Recommend question used in the Net Promoter Score (a 0 to 10 scale) by simply dividing the SUS by 10. For example, an SUS score of 72 would predict a LTR response of 7.2.

Jeff Sauro, 10 Things to Know About the System Usability Scale (SUS)

Peres, Pham and Phillips evaluated eight studies and found that, using SUS, researchers may be able “to say that one system is better than another” but not “speak to how much better”.

Different groups of people might rate the same product significantly differently. Rolf Molich and colleagues reported on a workshop in 2010 where fifteen experienced usability research teams evaluated the same website using different participants at the same time. Seven of them used the standard SUS, but only three ran samples large enough to be individually meaningful (43, 60 and 313 participants); the rest tested between 9 and 20 people, too few to place a score with any confidence. Across those three, mean scores ranged from 62.4 to 78.0, spanning D to B+ on the curved grading scale, and the 95% confidence intervals between the worst and the best score did not overlap at all. The authors of the study conclude that “it is likely that many of the observed differences occurred due to the different participants and evaluation procedures used by the teams.”

McLellan, Muddimer and Peres found that prior experience produced a significant difference in scores, and that users with extensive experience rated the same software 6.62 points above those who had never used it, so testing with the same group multiple times may create differences purely because of more experience.

A translated SUS cannot be compared against English benchmarks. Gao and colleagues compared five validated translations against the English scores for the same five products in Multi-Language Toolkit for the System Usability Scale, reporting that “All mean values were significantly different from the respective benchmark value at p <.001 unless specified”. Some gaps are very large. The biggest is twelve points, for a microwave rated 74.9 in Arabic against an English benchmark of 86.9. Spanish and German scores for some products were above the English figure. A translated SUS can be used to compare products within one language, but for an absolute score or checking against the published English benchmarks.

sus tools

Running SUS surveys

SUS is a post-test survey, so it is usually administered after a larger workflow. Brooke suggested asking people to complete it immediately after evaluating the system, before any debriefing or other discussion, and asking the participants to reflect on their immediate experience and not first consolidate their thoughts. All the questions are mandatory, and if people struggle to answer a specific question, instruct them to pick the middle option.

SUS is not good at detecting how individual steps or tasks contribute to the overall experience, so it’s useful to combine it with a quick post-task survey such as the Single Ease Question.

Get the most out of System Usability Scale with Votito

Get notified about new articles & major updates
Share