What is a Likert scale? It is a measurement method in which respondents rate how far they agree with a series of statements on ordered response options, typically from 'Strongly disagree' to 'Strongly agree', and the item scores are combined into a single score. Rensis Likert introduced the technique in 1932 to measure attitudes. This guide shows how to design a Likert scale that stands up to peer review.
What is a Likert scale? Scale versus Likert-type item
A Likert-type item is a single statement with ordered response options. A Likert scale is a set of items measuring the same construct (say, test anxiety) whose scores are summed or averaged, hence its other name, the summated rating scale. A lone 'How satisfied are you with the service?' question is a Likert-type item, not a Likert scale, even with a five-point format. The distinction sets your terminology and your level of analysis (ordinal items versus an approximately interval scale score), as our guide on how to analyse Likert data explains.
Anatomy of a Likert item and its anchor labels
Every Likert item has three parts: the stem (e.g. 'I find it hard to concentrate before an exam.'), the response options, and the verbal anchors that give them meaning; numeric codes (1–5) mostly serve data entry. A fully labelled format gives every option a verbal label, whereas an end-point labelled format names only the extremes and numbers the points in between. Full labelling is generally recommended for five- and seven-point scales because respondents then interpret each category more consistently; end-point labelling suits longer, near-numeric rating formats.
Anchors must fit the stem ('How often…' cannot take agreement anchors) and should not change within a subscale. Midpoint wording matters too: 'Undecided' implies indecision, 'Neither agree nor disagree' a neutral attitude. When adapting a scale, preserve the meaning of the original label.
| Response type | Five-point anchors (low to high) | Suitable stems |
|---|---|---|
| Agreement | Strongly disagree, Disagree, Neither agree nor disagree, Agree, Strongly agree | Attitudes and opinions |
| Frequency | Never, Rarely, Sometimes, Often, Always | Behaviours and experiences |
| Satisfaction | Very dissatisfied, Dissatisfied, Neither satisfied nor dissatisfied, Satisfied, Very satisfied | Service ratings |
| Importance | Not at all important, Slightly important, Moderately important, Very important, Extremely important | Priorities and expectations |
| Self-description | Not at all true of me, Slightly true, Moderately true, Mostly true, Completely true of me | Personality and self-report |
How many points? 4-, 5- or 7-point Likert scales
The number of points trades discrimination against respondent burden. More categories capture finer differences, but the gains fade beyond about seven and labelling gets harder. The choice also shapes the analysis: items with five or more categories can usually be treated as continuous in CFA, while four or fewer call for WLSMV-type estimators.
| Format | Advantage | Drawback | Typical use |
|---|---|---|---|
| 4-point (no midpoint) | Forces a direction | Pushes neutral respondents to a pole | Topics expecting a clear stance |
| 5-point | Familiar, easy to label fully | Midpoint may absorb 'no opinion' answers | Default in education and social sciences |
| 6-point (no midpoint) | Forced choice with finer gradation | Six clear labels are hard to write | Self-report and personality measures |
| 7-point | Finer discrimination, more variance | Labels blur in meaning; can be tiring | Psychology, marketing, organisational behaviour |
On the midpoint debate, keep the midpoint if the topic genuinely allows a neutral stance. Place 'Don't know / No opinion' outside the scale rather than at its centre, and treat it as missing. Changing the number of points on an adapted instrument stops the original validity evidence from transferring directly; if unavoidable, report and justify it.
Writing items: reverse-worded, double-barrelled and leading items
Acquiescence bias is the tendency to agree whatever an item says. Reverse-worded items help detect it, but careless readers misanswer them, and in CFA they can form a separate 'method factor'. Use them sparingly, and prefer opposite-meaning statements ('I stay calm during exams') to easily missed negations with 'not'. Recode them before scoring: new value = 6 − old value on a five-point scale, 8 − old value on a seven-point scale. A forgotten recode usually shows up as a low or even negative reliability coefficient; see our guide to Cronbach's alpha and McDonald's omega.
- Double-barrelled: 'My supervisor is accessible and supportive.' Two judgements, one answer; split it.
- Leading: 'As everyone knows, online learning is ineffective.' Write it neutrally: 'Online learning works well for me.'
- Double negative: 'I do not think the classes are uninteresting.' Disagreement becomes uninterpretable.
- Absolutes: 'always' or 'never' combined with agreement anchors push nearly everyone into one category.
- Factual: 'I study more than ten hours a week.' is a fact, not an attitude; ask it separately.
Number of items, pilot testing and scale adaptation
Aim for at least three, preferably four or more, items per dimension. In a one-factor CFA, three items make the model just-identified (zero degrees of freedom), so its fit cannot be tested; four or more also leave room to drop a weak item after piloting. For a new instrument, write a generous item pool; our scale development guide covers the path to EFA and CFA.
- Cognitive interviews: have a few members of the target population think aloud while answering, to catch misread wording.
- Pilot: in a small sample, check response distributions, piling-up in a single category and corrected item-total correlations; review items below 0.30.
- Translation–back-translation: obtain the original authors' permission; two independent translators produce the target version, another back-translates it, and an expert panel resolves discrepancies.
- Equivalence and structure: where feasible, give both versions to a bilingual group, then confirm the structure with CFA and reliability analysis in a new sample.
Scoring, measurement level and reporting in the method section
After recoding, a subscale score is either a sum score, if the original scoring instructions require it, or a mean score, which stays on the response metric (1–5) and makes subscales of different lengths comparable. In SPSS, Transform → Recode into Different Variables handles reverse coding, and MEAN.n returns a mean only for respondents with at least n valid answers, e.g. MEAN.5(m1 TO m6). Item scores are ordinal and multi-item scale scores approximately interval, so choose tests accordingly. A method-section description might read:
Test anxiety was measured with the 12-item [Scale Name], developed by [Author] ([year]) and adapted for the present population by [Author] ([year]). Items were rated on a 5-point Likert-type format (1 = strongly disagree, 5 = strongly agree), and three items were reverse-scored. Scores were computed as item means, with higher scores indicating greater anxiety. McDonald's omega in the present sample was .87.
For help with scoring, factor or reliability analysis, see our quantitative data analysis service.
A good Likert scale is built when you write the stem, not when you run the analysis.
Frequently Asked Questions
How many points should a Likert scale have?
Five or seven points suit most thesis research: five is familiar and easy to label fully, while seven captures finer differences. With an existing or adapted scale, keep the original number of points so its validity evidence still applies.
Is a Likert-type item the same as a Likert scale?
No. A Likert-type item is one statement; a Likert scale is a set of items measuring one construct, summed or averaged. Single items are analysed as ordinal data, multi-item scale scores usually as approximately interval.
How do you reverse score Likert scale items?
Subtract each response from the sum of the scale's lowest and highest values: 6 − x on a five-point scale, 8 − x on a seven-point scale. Recode before computing scores or running any reliability analysis.
What Likert scale design and analysis support does Celsus offer?
We review item wording and response formats, analyse pilot data, help document the translation–back-translation process, and run EFA, CFA and reliability analyses in SPSS, AMOS or Mplus, with thesis- and journal-ready reporting.