Customer Satisfaction Score (CSAT)
The Customer Satisfaction Score (CSAT) is a measurement of satisfaction with a product or a service, typically used to track trends over time. CSAT is popular because it is simple to collect and can be used as a leading indicator of behavioural impacts or business value.
At the same time, the simplicity of CSAT surveys can cause naive interpretation to be misleading. Satisfaction ratings often cluster heavily towards the top end of the scale, and the survey process itself can have an impact on the score, separately from how the survey is structured. It’s not easy to check reliability for single-question surveys. CSAT tends to predict what customers say they will do better than how they actually behave.
This article explains popular formats for CSAT surveys, the most common ways of scoring and interpreting CSAT questions, particularly dealing with skewness and statistical methods for evaluating confidence. It also presents the origins, applicability and limitations of the method.
- CSAT Questions
- How to calculate a CSAT score
- Interpreting CSAT scores
- Statistical analysis
- Origins of CSAT
- What does CSAT measure?
- Using CSAT in practice
Client Satisfaction Survey
CSAT Questions
The acronym CSAT and the name Customer Satisfaction Score are commonly used to describe a whole class of survey types, not a specific question or response wording. Unlike many other popular customer research methods, CSAT emerged organically through use, so it has no single inventor, no standard wording and no single scoring rule.
CSAT surveys usually contain a single question, asking the respondents to select how satisfied they are with a product, service, or an interaction. The response format typically uses a Likert scale between five and eleven points, and the answers are usually scored using top-box share methods.
Eugene W. Anderson, Claes Fornell and Donald R. Lehmann distinguish between transaction-specific satisfaction (measuring a single purchase experience or interaction) and cumulative satisfaction (measuring overall experience with a product, service or company over time).
How many points to use?
Since CSAT is usually presented as a Likert scale, the general recommendations for Likert-scale questions apply to CSAT, so four-point to eleven-point scales are going to be the most practical. In practice, most CSAT surveys today use either 5 or 11 points.
1. Overall, how satisfied or dissatisfied are you with [the product or interaction]?
Neil A. Morgan and Lopo Leotte Rego describe the five-point scale as the one companies typically use to capture customer satisfaction, although they give no source for it. The standard single-question scale is bipolar, with end labels similar to “very dissatisfied” and “very satisfied”. Robert A. Westbrook and Richard L. Oliver gave six satisfaction measures to 125 car owners and found the seven-point bipolar question the weakest at telling apart people who reported different emotions about their cars. Westbrook and Oliver also reported that angry respondents from their survey had an average score of 5.43 out of 7, well above the midpoint. Naively reading the scale as linear would make it seem as if the angry people are actually satisfied. Individual questions about satisfaction and about dissatisfaction separated the groups more clearly. Jochen Wirtz and Meng Chung Lee compared nine survey types (with 257 students) and reached a similar conclusion, that an 11-point scale from “not at all satisfied” to “completely satisfied” was the “best single-item scale”.
A unipolar 11-point Customer Satisfaction Score (CSAT) question, the best single-question format in Wirtz and Lee’s comparison.
1. Overall, how satisfied are you with [the product or interaction]?
The Swedish national satisfaction index composes a score from three questions on 10-point scales with no mid-point. Many other national indices, such as the popular American Customer Satisfaction Index and the Norwegian Customer Satisfaction Barometer, follow the same formula. Note that national indices do not use a single question, but a combination of scores from different questions.
Using non-verbal scale options
Non-verbal rating scales, such as smiley faces, are sometimes used on CSAT surveys instead of labels, particularly on physical terminals with limited space for words, such as in airports or sports venues. In a 2020 article for the Economist magazine, Amelia Tait reports that one terminal manufacturer sold more than 30,000 machines to customers in 135 countries.
Several researchers compared smiley-face surveys to those labelled with words, and found that adding faces changed the answers very little. Tobias Gummer and colleagues “found no convincing evidence that using smiley face scales altered response” behaviour in German online panels, Vera Toepoel and colleagues found that smileys scored like radio buttons in a Dutch panel, and Mathew Stange and colleagues found no difference in the answers in a US eye-tracking study.
Note that adding smiley faces might have unexpected consequences on survey completion. Gummer and colleagues report that their survey participants took longer to answer when faces were added to surveys. Alexandru Cernat and Mingnan Liu found that more people skipped questions on a PC when a row of faces replaced radio buttons. Niels Lassen and colleagues found that dissatisfied people were “more inclined to vote, or vote more often” at a smiley terminal than the rest, so the response averages tend to drift more towards the unsatisfied part of the sample population. The smiley-face terminal results still skew similarly to labelled CSAT surveys. According to Tait, the “Happy Index floats between 70-90%”, and across more than a billion responses, 70% were for the happiest face and 11% for the unhappiest.
When designing surveys with non-verbal options, be aware that people coming from different cultures might interpret them differently. Masaki Yuki and colleagues found that Japanese students judged a face mostly by its eyes, and American students by its mouth, and Hisako W. Yamamoto and colleagues found the same contrast between Japanese and Dutch students. In Cameroon and Tanzania, where the authors expected emoticons to be less familiar, Kohske Takahashi and colleagues found that the participants in Cameroon could not reliably tell a smiling face from a sad one, and those in Tanzania barely distinguished them, concluding that a smiley “does not necessarily look smiling to everyone”.
Including neutral options
It is important to consider whether to allow neutral or not-applicable options when designing a single-question survey. Odd-point bipolar scales often use the mid-point for neutral answers, but the typical result skewness suggests that the participants do not interpret mid-points as neutral.
Some survey formats work around this by labelling the mid-point explicitly, such as “neither dissatisfied nor satisfied”. Some offer an explicit “not applicable” or “don’t know” option.
In a study of US national surveys, Frank M. Andrews found that offering an explicit “Don’t know” answer correlated with better results more than an explicit midpoint (this was not a randomized trial of survey types, so it’s an interesting reported correlation more than a rule). Jon A. Krosnick, reviewing experiments on attitude surveys, reported that many more people say they have no opinion when that option is offered explicitly than when they have to volunteer it, but offering it does not make the data more reliable.
Some researchers place several neutral options outside the scale, to differentiate between ambivalent satisfaction and lack of interest or other reasons why someone could not answer clearly. For example, Andrews and Stephen B. Withey placed three separate answers off the scale: “Neutral (Neither satisfied nor dissatisfied)”, “I never thought about it” and “Does not apply to me”.
George S. Day warns that many purchases happen without a deep evaluation and that “research procedures which assume that an evaluation always takes place may produce one as an artifact of the research”. For satisfaction questions, especially in post-transaction surveys, including a “not applicable” option might be a good practical way to filter out people whose opinion might not have been fully formed.
If you include a “no opinion” option, then exclude such responses from the main sample, but report how many people chose it. A big change in opt-out selections is an interesting finding in itself.
How to calculate a CSAT score
Similarly to how there’s no standard question format, there is no standard scoring system for CSAT. There are several systems commonly used to combine results from multiple CSAT surveys into a single number, so it can be used to monitor satisfaction over time or compare against benchmarks. The most popular methods are:
- Mean: the average of the numeric answers.
- Top-box: the share of people choosing the highest point on the scale.
- Top-2-box: the share of people who chose the two highest points. On a five-point scale this is everyone above neutral, so the top-2-box score on a five-point scale is also called percent satisfied.
- Net: the top-box share minus the bottom-box share (called Net Satisfaction, or NSAT).
The chosen method significantly impacts the outcome when comparing different response sets. For example, Jeff Sauro and James R. Lewis compared different rankings from different score types in a survey comparing 12 airline websites. The mean (6.1) and the top-2-box score (83%) ranked Alaska Airlines best, while top-box ranked American Airlines first, at 41% against Alaska’s 35%. Across the 12 sites the mean correlated .97 with top-2-box but only .76 with top-box.
Some benchmark sources publish several scores. For example, the Centers for Medicare & Medicaid Services publishes top-box, middle-box and bottom-box percentages, while their star ratings use adjusted linear scores on a 0–100 scale.
There’s no universal rule of thumb for which method to use. No scoring rule works best everywhere, and the studies that compared them disagree about the outcome.
Evert de Haan, Peter Verhoef and Thorsten Wiesel compared top-2-box CSAT, mean CSAT, NPS and CES using data about 93 Dutch companies and concluded that top-2-box was the best predictor of retention, ahead of the mean. Morgan and Rego found that “average satisfaction scores have the greatest value in predicting future business performance and that Top 2 Box satisfaction scores also have good predictive value”, comparing data for large US consumer companies.
When using satisfaction ratings as a leading indicator for recommendations or word-of-mouth marketing, it’s important to track both ends of the scale. Eugene W. Anderson found an asymmetric U-shape, concluding that the most dissatisfied customers talk to the most people, but the most satisfied customers also talk to more than those in the middle. Peterson and Wilson argue that it is more useful to “examine the tail of a satisfaction-rating distribution”, the customers who say they are dissatisfied, than to focus on the satisfied ones.
In general, because there are so many different ways of collecting and scoring, if you compare against a benchmark, use the scoring system that the benchmark was created on. If you want to compare against your own data, and track retention, the top-2-box score is probably the most relevant. If that number doesn’t move enough to track improvements, the top-box percentage might be more sensitive. For example, Vikas Mittal and Wagner Kamakura suggest that moving from a rating of 4 to a rating of 5 “has a disproportionately larger impact on repurchase behavior than a corresponding move from a score of 3 to 4”.
If you have a small sample and want to detect changes, showing the mean with a confidence interval is probably best.
To catch problems, report the bottom-box percentage separately.
Net scores (the difference between top and bottom boxes) may be appealing because they capture both the top and bottom ends of the distribution in a single score. They look simple, but they require very large samples to be meaningful, so it’s better to report the different box scores separately.
Interpreting CSAT scores
Many CSAT scoring systems present a percentage or a score on a scale of 0 to 100, but CSAT distributions are not normal, so CSAT scores should not be naively interpreted as percentages.
Virtually all self-reports of customer satisfaction possess a distribution in which a majority of the responses indicate that customers are satisfied and the distribution itself is negatively skewed.
– Robert A. Peterson, William R. Wilson, Measuring Customer Satisfaction: Fact and Artifact
Because there is a huge variance in how CSAT surveys look and how the scores are calculated, generally it’s difficult to compare CSAT scores from different sources. Generic advice should be taken with caution. For example, Salesforce’s guide to CSAT says that “anything above 70% is considered a good customer satisfaction score”, that a score below 50% is poor and that the average across industries is 78%, without giving a source for any of the three figures.
In Hall and Dornan’s meta-analysis of medical patient studies, in the 68 studies that reported satisfaction, the median share satisfied was 84%, ranging from 43% to 99%. So an 80% top-2-box percent-satisfied score is not good; it’s just average. Jones and Sasser argue that managers often misread satisfaction distributions. Most would be “happy to learn that 82% of their customers fell into” the top two categories.
When comparing scores against a benchmark, it’s important to know which method was used for the reference scores, otherwise the results will be meaningless or misleading. Some vendors are very sloppy when it comes to such comparisons. For example, IBM’s guide to CSAT uses unadjusted top-2-box scoring but compares it to the ACSI index, which uses a weighted index from three 10-point questions. An ACSI of 75 and a CSAT of 75% are different measurements that just happen to share a range; they are not directly comparable.
Satisfaction ratings cluster towards the top
Satisfaction ratings are heavily clustered towards the top result. Peterson and Wilson’s review concluded that satisfaction self-reports are invariably negatively skewed and show a positivity bias. Claes Fornell found negatively skewed satisfaction in 28 of 29 Swedish industries and all 40 American ones in national index data.
There are several potential reasons for skew, similar to the Acquiescence Effect.
One common assumption is the response rate bias, that people who are more satisfied are more likely to answer. Peterson and Wilson’s research suggests that “satisfaction percentages are not related to response rate percentages”, and say that there is “no logical reason” satisfied customers should respond more. However, in a study of patients of 82 physicians from the data of a US health insurer, Kathleen M. Mazor and colleagues found that response rate correlated .52 with mean satisfaction. So whether satisfied or dissatisfied customers are more likely to answer seems to depend on the setting.
Peterson and Wilson also evaluated the question form, and suggest that “framing the satisfaction question in positive terms is likely to lead to more favorable associations than framing it in negative terms”, and found statistically significant results. Performing a surveys with car owners and asking one group how satisfied they were, and the other group how dissatisfied they were, they found that 91% of the respondents of the first version said they were somewhat or very satisfied, compared to 82% of the respondents in the second group.
The third potential cause is the timing of the survey, also claimed by Peterson and Wilson to be statistically valid (but without the data source or statistics published in the paper). Peterson and Wilson repeated car owner satisfaction surveys 30 days and 90 days after purchase, and concluded that there “was about a 20 percent decline in satisfaction ratings over the 60-day time period” but even then, the ratings were positively skewed in both sets. Most satisfaction surveys are conducted immediately after an interaction, and they are likely to be influenced by the immediate positive experience.
Adding more points to the scale is a common approach to resolve skewed results for Likert scales, but it doesn’t work fully for CSAT, because the satisfaction questions and timings themselves create the skew. For example, Claes Fornell and colleagues who designed the Swedish national satisfaction index apply three fixes:
In CSB, the problem of skewness was handled by:
(1) extending the typical number (usually 5 or 7) of scale points to 10 (to allow respondents to make finer discriminations),
(2) using a multiple-indicator approach (to achieve greater accuracy), and
(3) estimating via a version of partial least squares (PLS).
– Claes Fornell, A National Customer Satisfaction Barometer: The Swedish Experience
Asking a unipolar question instead of a bipolar one can separate different groups of customers better. In Westbrook and Oliver’s study, the angry and upset owners averaged 5.43 out of 7 on the bipolar question but 5.64 out of 10 on the unipolar satisfaction question (our reading of their table).
Andrews and Withey developed the Delighted-Terrible scale as an alternative to bipolar satisfaction scales. In their 1976 book, Social Indicators of Well-Being, Andrews and Withey report that 54% of the research participants scored to be in the satisfied category on the D-T scales, compared to 65% with the typical Likert surveys. Westbrook suggests that using the Delighted-Terrible scale “reduces the skewness of satisfaction responses”.
Ultimately, the skewness of the results is a consistent attribute of satisfaction surveys, so it just needs to be handled in reporting in some way. One potential way to do that is to report the full distribution, not just the single final score, so researchers can visually compare the changes.
Statistical analysis
When using top-box or top-2-box scoring, Alan Agresti and Brent A. Coull recommend using the adjusted-Wald interval. To compare two box scores from different groups of respondents, Sauro and Lewis recommend the N−1 two-proportion test in Quantifying the User Experience.
When using the mean to score CSAT, the typical approach is to use the t-interval for confidence and the t-test to compare different groups. Gail M. Sullivan and Anthony R. Artino Jr. argue that means are of limited value unless the answers are close to a normal distribution, and most CSAT answers will be close to the top, so the t-interval around the mean can expand past the end of the scale. In theory, the Mann–Whitney U test is safer for satisfaction ratings than the t-test, because it uses only the order of the answers, and it does not assume that the steps between the ratings are equal. In practice, both tests seem roughly equal. In simulations of five-point items, de Winter and Dodou found the two tests had roughly equal power for most comparisons. The Mann–Whitney U test performed better for groups where answers were skewed, and the t-test did better in some comparisons with a polarised group, particularly one split between the two ends of the scale.
For modelling what a score predicts, Mittal and Kamakura found that models treating each scale point as a separate category performed better than models treating the rating as a linear score. Ittner and Larcker treat satisfaction measures as “ordinal rather than cardinal”, and used nonparametric regression.
Origins of CSAT
Richard N. Cardozo’s 1965 paper An Experimental Study of Customer Effort, Expectation, and Satisfaction is one of the earliest recorded research experiments that included measurements of satisfaction. In the paper, he notes that the marketing and economics literature at the time did not provide exact definitions nor rigorous discussion for the terms he evaluated. Cardozo’s experiment was an attempt to evaluate the connection between satisfaction and consumer expectations, and not an investigation of satisfaction in a real market. Cardozo asked students to rate ballpoint pens compared to other pens in a catalog, on a scale of 0 to 100 with labels “Very inferior”, “Rather inferior”, “Somewhat superior” and “Vastly superior”. Cardozo concluded that satisfaction depends not just on the product under evaluation, but also the “experience surrounding acquisition of the product”, noting that “the definition and measurement of total satisfaction pose a complex problem”.
H. Keith Hunt described how the US Federal Trade Commission tried to rank consumer problems by dissatisfaction in the mid-1970s, and found that “no dissatisfaction measure existed”, so its staff fell back on complaint counts. In 1973, Rolph Anderson attempted to model dissatisfaction, explaining that at the time “No satisfactory literal definition has yet been developed for consumer satisfaction or dissatisfaction in the literature of marketing”.
The earliest systematic consumer satisfaction survey we found started in 1971, when the researchers at the University of Michigan conducted a survey of 574 households asking people to rate products attributes on a seven-point scale from “very satisfied” to “not at all satisfied”, using letters instead of numbers “in order not to suggest a particular order or specific quantitative relation between points on the scales.” This was a precursor to the most commonly used CSAT survey format. This survey already discovered potential problems with bipolar ratings, and Anita B. Pfaff explains that the team rejected the “very satisfied to very dissatisfied” form as a “mixed scale, derived from two separate scales”, one measuring satisfaction and one dissatisfaction. James C. Lingoes and Martin Pfaff dealt with the problem of assigning numerical scores to satisfaction scales in 1972, trying to create a methodology for a satisfaction index, and provide early discussions on aggregating subjective ratings.
In 1975, a US Department of Agriculture report by Charles R. Handy and Martin Pfaff, on a survey of 1,831 households, reported that 66.2% of respondents were “always or almost always satisfied” with food products, documenting the clustering and the negative skew typical of satisfaction surveys. The authors called that level “somewhat surprising in the face of substantial evidence that consumers are disgruntled”. Handy and Pfaff used a five-point numeric scale, with 1 corresponding to “always satisfied”, 2 to “almost always satisfied”, 3 to “sometimes satisfied”, 4 to “rarely satisfied” and 5 to “never satisfied”, and used the top-2-box score for percent satisfied and the mean score to summarize satisfaction with different types of food. This is an effective precursor to the most popular CSAT scale and two of the most popular scoring types.
In The Quality of American Life, published in 1976, Angus Campbell, Philip E. Converse and Willard L. Rodgers report on research that included questions such as “Which number comes closest to how satisfied or dissatisfied you feel?” on a seven-point scale labelled “Completely dissatisfied” and “Completely satisfied” at the ends, matching common CSAT question scales. The same year, Andrews and Withey describe that scale as producing “markedly skewed distributions”, with half to two-thirds of respondents in the two most satisfied categories, and propose a Delighted–Terrible scale to reduce the skew.
By 1977, enough different competing models had emerged that Alan R. Andreasen published A Taxonomy of Consumer Satisfaction/Dissatisfaction Measures, including what he called “Simple Satisfaction Scales” that are a precursor to modern CSAT surveys. Andreasen warned that such scales “underreport actual problems”.
In 1979, Ernest R. Cadotte described a push-button terminal at hotel check-out, allowing customers to choose between dissatisfied, satisfied or very satisfied, and scoring answers 0, 50 and 100, which can then be numerically analyzed.
Westbrook opens his 1980 paper by noting that “consumer researchers have used rather simple measures, most often single-item rating scales of four to seven points between the extremes of ‘very satisfied’ and ‘very dissatisfied’”, so modern CSAT surveys emerged somewhere around that time through practice.
John A. Quelch and Stephen B. Ash reported the percentages of satisfied and dissatisfied consumers for each product in 1980, from a four-point scale with no midpoint, effectively suggesting tracking top-2-box and bottom-2-box scores separately, without calling them that. The earliest use of the “top box” and “two top boxes” terms we found is a 1998 study by Christopher D. Ittner and David F. Larcker, who list both among the satisfaction measures companies commonly use. They describe a bank that had scored each branch since 1995 on the share of customers answering 6 or 7 on a seven-point scale, which they call the “two top boxes”.
By 1992, when Bob E. Hayes published his handbook for the American Society for Quality Control, the five-point CSAT from Very Dissatisfied to Very Satisfied had become standard practice. Hayes lists it as one of the three standard response types.
What does CSAT measure?
Research suggests that self-reported satisfaction metrics, such as CSAT, predict what customers say they will do better than their actual behaviors. In addition, satisfaction metrics tend to capture opinions that are broader than satisfaction alone.
Transactional surveys completed shortly after a purchase capture someone’s opinion both about the product and about the whole purchase experience. Caroline Ardelet and Christophe Benavent studied 314,194 customer contacts with 96 brands, and found that “the less effort customers exert when interacting with a brand, the more satisfied they will be”, concluding that the effort in an interaction lowers the satisfaction ratings (with exceptions, notably “hedonic industries” where “asking the customer to invest intense effort in interacting with the brand can have a positive effect on satisfaction”). Chezy Ofir and Itamar Simonson found that “expecting to evaluate leads to less favorable quality and satisfaction” ratings. For example, the satisfaction ratings for computer services differed significantly depending on whether customers were told upfront about a post-service survey or not. Those who were told in advance scored the service 3.7 out of 5 on average, compared to a 4.2 average rating for customers who did not get the survey announcement. Notably, asking customers to provide a positive rating actually produces lower ratings.
William Boulding and colleagues note that cumulative satisfaction surveys are analogous to measures of perceived quality. Maurice LeVois, Tuan D. Nguyen and C. Clifford Attkisson report that the survey channel itself also has an impact on ratings, and that “oral administration of the CSQ produced 10% higher satisfaction ratings than written administration”. Peterson and Wilson, from undisclosed data, report that “telephone interviews consistently result in satisfaction percentages for automobiles approximately 12 percent higher than those obtained through mail interviews”.
Satisfaction is a somewhat leading indicator of customer recommendations and retention, but not very strong. Vikas Mittal and colleagues found that satisfaction correlated .65 with customers’ stated intention to stay (intention for retention), but only .21 with retention measured as behaviour. In a study of 6,649 Dutch consumers who evaluated 93 companies, de Haan, Verhoef and Wiesel concluded that the top-2-box CSAT score on a seven-point scale correlated .184 with retention, better than NPS or any other single-item metric they observed, so CSAT can be a weak indicator of retention. Timothy L. Keiningham and colleagues, using an online panel of more than 8,000 US customers of banks, large retailers and internet providers, found that overall satisfaction was more closely related to recommendations (median correlations .26 to .36 across three industries) than to retention (.10 to .20). Keiningham and colleagues conclude that “no one metric best predicts all behaviors associated with customer loyalty”, and that “each dimension is likely to be affected by differing aspects of the customer experience”. Kathleen Seiders and colleagues studied 945 customers of a US retailer, and found that satisfaction correlated .53 with stated repurchase intentions, but only .07 with visits and spending over the next 52 weeks. Seiders and colleagues warn that relying on satisfaction and intentions “may create false security”. Satisfaction and retention studies mostly do not account for switching costs, particularly among virtual monopolies. Banks, internet and telecom providers are sometimes not easy to switch. Jones and Sasser describe the local telephone service as an example of a market where customers stayed regardless of satisfaction ratings.
Importantly, sometimes the act of surveying for satisfaction correlates with different customer behaviour. Utpal M. Dholakia and Vicki G. Morwitz compared customers of a financial services company, 945 of whom completed a satisfaction survey against 1,064 that were not surveyed. Over the next year, 51% of the surveyed customers opened a new account, compared to only 13.3% in the group that was not included in the survey. In the group that completed the survey, 6.6% churned, compared with 16.4% for the group that was not surveyed. Sharad Borle and colleagues found that customers of a US car servicing chain who answered a post-visit satisfaction call later spent about 5.6% more per visit than those who did not. Both studies were not randomized and only report on correlation, not causation.
When satisfaction does affect retention, the relationship between the two is not linear. Ittner and Larcker analyzed a survey of 2,491 small-business customers of one US telecom, comparing satisfaction to one-year retention. The bottom decile retention was 60%, the sixth decile retention was 81%, and “scores above 70 produced no increase in” retention. Mittal and Kamakura used a five-point scale and concluded that the change from 4 to 5 (“very satisfied”) had disproportionate weight, and that a linear model would understate the “impact of a change in score from 4 to 5 by 64%”. Eugene W. Anderson and Mary W. Sullivan found that satisfaction was more sensitive to falling short of expectations than to exceeding them, studying more than 22 thousand customers of 57 companies in Sweden. Mittal, William T. Ross Jr. and Patrick M. Baldasare found that “negative performance on an attribute has a greater” impact on overall satisfaction than positive performance.
There is also some evidence that satisfied customers purchase more from the same provider. Keiningham, Tiffany Perkins-Munn and Heather Evans analyzed the relationship between satisfaction ratings and share-of-wallet at a financial services company, and concluded that clients giving top-2-box satisfaction ratings on a 10-point scale gave the financial provider 15.4% of their business, compared to 10.8% for the rest.
There were plenty of studies trying to find a link between customer satisfaction and overall financial performance of companies, with conflicting results. Ittner and Larcker found “modest support for claims that” satisfaction leads financial performance.
Fornell and colleagues reported in 2006 that a theoretical ACSI-weighted trading portfolio from 1997 to 2003 returned 40% against 13% for the S&P 500, but in their words that portfolio was “back tested with perfect hindsight”. Don O’Sullivan, Mark C. Hutchinson and Vincent O’Connell evaluated the same strategy and period adjusted for risk, and found that “risk adjusted returns from portfolios based upon high ACSI scores, low ACSI scores or changes in ACSI scores are not significantly positive”. Lerzan Aksoy and colleagues tested a theoretical trading strategy based on the ACSI customer satisfaction index from 1996 to 2006. According to their data, $100 invested in the high-satisfaction companies would have grown to $312 in 10 years, compared to $205 for the same amount invested in stocks with high customer satisfaction and short selling those with low customer satisfaction, has a significant, positive abnormal return ranging from 11% to 14% per year”. Robert Jacobson and Natalie Mizik evaluated Aksoy’s trading strategy and found that the extra returns came only from ten computer and Internet companies over a period that included the dot-com boom. High satisfaction ratings did not create a significant difference in value for other companies in the same period. In a 2016 paper, Claes Fornell, Forrest V. Morgeson and G. Tomas M. Hult claim that stock market investments in a fund driven by high-satisfaction companies recorded “cumulative returns of 518% over the 15 years studied (2000–2014), compared with a 31% increase for the S&P 500”, without going into the details of fund, the fees or the trading strategy. Kapil Tuli and Sundar Bharadwaj compared stock prices with ACSI data for 129 companies from 1994 to 2006 and found that rises in satisfaction were correlated with lower stock-return risk, although they do not report confidence intervals in their research.
Ashley S. Otto, David M. Szymanski and Rajan Varadarajan’s meta-analysis of 251 correlations from 96 studies suggests that there is a small association (.101) between satisfaction and company performance. Their analysis is also notable for the conclusion that single-item measures such as CSAT are not statistically significantly different from ACSI-type indices in predicting overall company performance. Ittner and Larcker also found that a single 1–10 overall satisfaction question predicted retention as well as a more complex three-item index, although nothing in their data explains more than 2.4% of the variance in retention.
Using CSAT in practice
Generally, for evaluating overall trends in cumulative satisfaction, the choice of a specific metric or scoring system doesn’t matter as much as consistency. Vikas Mittal, Sonam Singh and Ashwin Malshe surveyed about a thousand US consumers online, rating brands, and found that the single ten-point question “Overall, how satisfied are you with the brand?” correlates highly with more complex multi-item satisfaction scales, and concluded that researchers can “use the different CSAT scales interchangeably”.
CSAT scores are only comparable against scores compiled using the same process. If you want fine-grained distinctions between customer satisfaction categories, use the unipolar 11-point scales. Wirtz and Lee’s results suggest that those are the best at discriminating individual categories for single-question surveys. For general tracking, the five-point scales with top-2-box scoring (percent satisfied) are the most common, and your customers will already be used to such questions, so that seems a reasonable default even though it’s not the most precise. Avoid smiley faces or other non-verbal labels. For a mixed-language audience, if the space allows, use translated verbal labels. If you must use faces, test the selection for understanding with your target audience before running the survey. Report the whole distribution next to the single score, and provide a confidence interval for the score.
For transactional satisfaction surveys, make the interaction as low effort as possible, do not announce the survey upfront, and track the top-box and bottom-box results separately. Individual ratings are noisy, so treat individual responses as prompts for follow-up questions rather than as measurements. Perhaps automatically add a free-text explanation question when someone provides a low score, or follow up with the person directly to resolve the issues.
A single-question satisfaction score tends to combine several types of consumer opinions, such as matching expectations, efficiency and effectiveness of a product or service and the effort of the survey process itself, so it is not very discriminating. If you want to track individual aspects of customer satisfaction or use it as a leading indicator of consumer behaviors, a more complex survey will probably give you more reliable results.