CSConjoint Survey

How to Interpret VADER Sentiment Scores

Read VADER compound scores responsibly, distinguish sentiment from agreement or stance, and interpret individual responses, sample means, and experimental differences.

The Short Version

Read VADER compound scores responsibly, distinguish sentiment from agreement or stance, and interpret individual responses, sample means, and experimental differences. A practical way to begin is to start with the compound score. VADER combines the sentiment intensity in a response into one normalized value from -1 to +1. Values nearer +1 indicate more positive language, values nearer -1 indicate more negative language, and values near 0 indicate neutral or sentiment-balanced language, then use common classification labels only when they help the research question. A frequently used convention labels scores of 0.05 or higher positive, scores of -0.05 or lower negative, and scores between those values neutral. Treat these as classification rules, not universal boundaries for a meaningful effect. Finish by making sure you validate consequential analyses. Compare a sample of VADER scores with blinded human coding or a construct-specific measure, preregister the scoring and comparison when possible, and retain the raw text so readers can audit the interpretation. The full walkthrough below explains each part in order.

Is This Guide for You?

Researchers using Conjoint Survey's open-ended text analysis who need to explain what a VADER score does and does not show.

Follow These Steps

Work through these 9 steps at your own pace. The wording is intentionally practical, and you can return to the checklist at the end when you are ready to review your work.

  1. Start with the compound score. VADER combines the sentiment intensity in a response into one normalized value from -1 to +1. Values nearer +1 indicate more positive language, values nearer -1 indicate more negative language, and values near 0 indicate neutral or sentiment-balanced language.
  2. Use common classification labels only when they help the research question. A frequently used convention labels scores of 0.05 or higher positive, scores of -0.05 or lower negative, and scores between those values neutral. Treat these as classification rules, not universal boundaries for a meaningful effect.
  3. Interpret an individual score as a feature of that response's wording. For example, a score near 0.75 contains strongly positive language, a score near 0 may be neutral or contain offsetting positive and negative language, and a score near -0.60 contains strongly negative language. Always read examples from the actual responses before attaching a substantive label.
  4. Interpret a sample mean as the average tone of the analyzed responses, not the percentage of respondents who agree. A positive mean can arise from many mildly positive responses, a smaller number of very positive responses, or a mixture of positive and negative responses.
  5. In a randomized survey experiment, interpret a treatment effect as treatment mean minus reference-arm mean. A positive difference means the treatment caused more positive wording on average; a negative difference means it caused more negative wording on average. Read the confidence interval, p-value, sample sizes, and raw examples alongside the difference.
  6. Remember what VADER is designed to notice. Its rules account for features such as negation, intensifiers, capitalization, punctuation, and some informal expressions, but no rule-based lexicon understands every context.
  7. Do not translate sentiment into a different construct. Positive language is not automatically support, agreement, persuasion, factual accuracy, trust, ideology, or policy preference. Negative language can express concern about a problem while supporting the proposed solution.
  8. Inspect failure cases. Sarcasm, technical vocabulary, long mixed passages, quoted material, ambiguous wording, domain-specific meanings, and languages other than the English lexicon being used can produce misleading scores.
  9. Validate consequential analyses. Compare a sample of VADER scores with blinded human coding or a construct-specific measure, preregister the scoring and comparison when possible, and retain the raw text so readers can audit the interpretation.

What This Helps You Accomplish

VADER, introduced by Hutto and Gilbert (2014), is a transparent rule-based measure of expressed sentiment. Its compound score is useful as a reproducible text feature, especially for comparing randomized groups, but the number cannot identify why a respondent used that language or replace a validated measure of the construct a study actually cares about. Responsible interpretation keeps the statistical claim aligned with what the scoring method measured.

A Quick Confidence Check

  • State that the compound score ranges from -1 to +1 and that higher values indicate more positive language.
  • If using positive, neutral, and negative labels, report the classification thresholds and justify their use.
  • Report sample sizes, group means, the mean difference, and uncertainty for experimental comparisons.
  • Do not describe sentiment as agreement, support, stance, persuasion, ideology, or truthfulness without separate validation.
  • Read raw examples from low, middle, and high scores and check surprising cases.
  • Check whether the language and domain fit the English VADER lexicon used by the app.
  • Use human coding or a construct-specific measure for high-stakes or publication-defining claims.
  • Keep raw text, word counts, VADER scores, exclusions, and scoring notes in the replication bundle.

Common Questions

Can I use this guide if I am new to this?

Yes. Researchers using Conjoint Survey's open-ended text analysis who need to explain what a VADER score does and does not show. Follow the steps in order, start with a small test, and use the final checklist before you field or report anything important.

What is the simplest way to get started?

Begin by start with the compound score. VADER combines the sentiment intensity in a response into one normalized value from -1 to +1. Values nearer +1 indicate more positive language, values nearer -1 indicate more negative language, and values near 0 indicate neutral or sentiment-balanced language. Next, use common classification labels only when they help the research question. A frequently used convention labels scores of 0.05 or higher positive, scores of -0.05 or lower negative, and scores between those values neutral. Treat these as classification rules, not universal boundaries for a meaningful effect. You do not need to perfect every setting before running a small preview or pilot.

How do I know when I am ready?

Use the confidence check above. In particular: State that the compound score ranges from -1 to +1 and that higher values indicate more positive language. If using positive, neutral, and negative labels, report the classification thresholds and justify their use. Report sample sizes, group means, the mean difference, and uncertainty for experimental comparisons. Do not describe sentiment as agreement, support, stance, persuasion, ideology, or truthfulness without separate validation. When the decision is consequential, keep your study documentation and ask a qualified colleague or methods reviewer to examine the design as well.

Related Guides

Research artifact: sample data, codebook, and scripts

Research Pathways

Product names and fielding platforms mentioned in this guide belong to their respective owners. Conjoint Survey is an independent academic research tool and is not affiliated with, sponsored by, or endorsed by those services.