Why Two Personality Quizzes Can Give You Two Different Results

By Big Time Trivia Editors. Published .

Take two personality quizzes on the same afternoon and they can disagree. One calls you an extravert and the other leans introvert. One love language test says words of affirmation, another says quality time. The easy conclusion is that one of them is broken. Often neither is. Two quizzes can use the same name for different things, ask in different formats, and catch different versions of you. Five of those differences are below, with the research on each, followed by what they mean for the quizzes on this site.

This piece is about two different quizzes. Getting a different result when you retake the same quiz is a related but separate problem, and the piece on what a close split tells you covers it.

The same name for different traits

A quiz is a set of questions with a label on top, and two quiz writers can put the same label on quite different questions. Victoria Pace and Michael Brannick tested how far that goes among commonly used personality scales. In a 2010 meta-analysis in Personality and Individual Differences, they pooled correlations between different scales with similar labels, such as different measures of extraversion. The scales were about equally reliable, but "scales of the 'same' construct were only moderately correlated in many cases." Limiting the comparison to scales built on the Five-Factor Model improved the agreement somewhat, and the authors close by saying that questions remain about how similarly the traits are defined and measured.

Those were published research scales, and the gap can be easy to see. The Big Five Inventory, a widely used research questionnaire, counts agreement with "Is outgoing, sociable" toward Extraversion (the item is quoted in a 2008 study by Soto, John, Gosling and Potter). The Myers-Briggs framework, popular but of contested scientific standing (the evidence is in the science behind MBTI), describes its Extraversion and Introversion pair as "Opposite ways to direct and receive energy", and asks whether you prefer to focus on the outer world or your own inner world. Someone who is sociable but recharges alone could fairly answer as an extravert on the first kind of question and lean introvert on the second. MBTI and the Big Five goes through where the two systems line up and where they part.

A habit of agreeing

Many quizzes hand you a statement and a scale from "strongly disagree" to "strongly agree." Some people lean toward agreeing whatever the statement says, and some lean the other way. Psychologists call it acquiescent responding, which Soto and colleagues, citing a 1958 paper by Jackson and Messick, describe as "the tendency of an individual to consistently agree (yea-saying) or consistently disagree (nay-saying) with questionnaire items, regardless of their content."

Whether that habit moves your result depends on how the questionnaire was built. A statement such as "Is talkative" counts toward Extraversion when you agree with it. Its opposite, "Tends to be quiet," counts toward Extraversion when you disagree. When a scale has more statements of the first kind than the second, the same paper spells out the consequence: a respondent who tends to agree "will score too high on a scale with more true-keyed items than false-keyed items and too low on a scale with more false-keyed items." In the original Big Five Inventory, the paper notes, every scale had more of the first kind. Among the 230,047 online volunteers aged 10 to 20 in that study, people varied considerably in this habit, and the variance in it at age 10 was twice the variance at age 20.

The habit can be designed around. The authors of the revised Big Five Inventory-2, published in 2017, describe it as controlling for individual differences in acquiescent responding. Plenty of quizzes don't. Take one quiz written mostly in statements you agree with and another balanced between the two kinds, and a person who agrees easily can come out differently on each without anything about them having changed.

Which situation you pictured

"I am talkative" doesn't say where. At work, at home, with strangers, with old friends? Each person fills in a setting, and not always the same one from question to question. Researchers call the setting a frame of reference. A quiz that says "at work" supplies one, and a quiz that says nothing leaves you to supply your own.

That choice changes how people answer. In two studies published in 2008, with 337 and 105 participants, Lievens, De Corte and Schollaert found that giving a frame of reference "reduces within-person inconsistency" in how people respond to generic items. Put plainly, tying the statements to a setting made each person's answers less inconsistent than leaving them generic. An earlier 2004 replication describes the same effect: a common frame of reference "standardizes item interpretation" and has been shown to reduce measurement error. So a quiz framed around work and an unframed quiz can be reading two different versions of you.

Who you compared yourself with

Rating yourself on a scale means rating yourself against someone, usually without noticing. Steven Heine and his coauthors Lehman, Peng and Greenholtz built a paper around that, published in 2002 in the Journal of Personality and Social Psychology (read here in the authors' preprint). Their example is height. The same height, 5 feet 9 inches, "will be seen as tall in some contexts (e.g., elementary school children, or among Japanese women) and short in others (e.g., professional basketball players, or among Dutch men)." They add that these comparisons tend to happen spontaneously, without deliberate awareness.

Their concern was comparisons between countries. Cultural experts agreed that East Asians are more collectivistic than North Americans, yet cross-cultural comparisons of self-rated traits, attitudes and values failed to show it, and the paper's explanation is that each group was rating itself against the people around it. The same mechanism can work on one person. Rate how organized you are with your most chaotic friend in mind and you'll answer one way; with your most meticulous colleague in mind, another. Two quizzes taken in different moods or settings can call up different comparison groups.

Rating scales and forced choices

Some quizzes let you rate every statement on its own. Others make you choose between options, so picking one means not picking the others. The two formats behave differently.

Choosing has a real advantage. According to a 2013 paper by Brown and Maydeu-Olivares, comparative formats "can reduce the impact of numerous response biases" that affect rating scales. A 2019 meta-analysis by Cao and Drasgow compared forced-choice personality scores between low-stakes settings and high-stakes ones such as job selection, where people have a reason to make themselves look good. The overall score inflation for forced-choice measures came to an effect size of 0.06, and in selection scenarios it was much lower than for single-statement measures across most personality facets.

Choosing has a cost too. When forced-choice answers are simply counted up, Brown and Maydeu-Olivares point out, every person ends up with the same total, so "it is impossible to achieve all high or all low scale scores." A count like that tells you how your options rank against each other, not how strongly you feel about any one of them. A rating quiz can say you care a lot about two things at once. A counting quiz can only say which of them you picked more often.

What this means for the quizzes here

The quizzes on this site ask you to pick one answer per question. None of them uses a scale from disagree to agree, so the yea-saying habit has much less to work on than it would on a rating quiz. The other differences above apply too, in different ways.

  • Love languages. Every answer is one point for one of the five, so your points always add up to the number of questions you answered. The result says which language you picked most often, relative to the other four, and shows the count for each. If a rating-style love language test puts touch first and the love language quiz here puts quality time first, the two may not disagree at all, since one asks how much each matters and the other asks which wins when you can only have one. The guide to the five love languages looks at this forced-choice limit in Gary Chapman's own quiz.
  • Game of Thrones characters. The character quiz counts the same way: one point to one character per answer, the character with the most points named (every one of them, if two or more tie), and the count for every character shown beside it. So a close second shows up on the page instead of disappearing behind one name.
  • Personality type. Each question on the personality type quiz is about one of the four pairs, and each pair is counted separately, so a choice on one pair can't take points from another. Many of the questions set a scene, such as three hours into a conference where you don't know anyone, which supplies a frame of reference instead of leaving you to pick one. A different quiz with different scenes can catch a different side of you. One question on the full version comes close to agree or disagree: it asks whether a friend's description of you as always in your head is fair, and the two answers that accept it most fully count toward Intuition, so a habit of agreeing could nudge that one answer. A pair split evenly or narrowly enough is shown as balanced, with both counts, instead of being forced into a letter, and the quiz page says which splits count as balanced on each version.

How to compare two results

When two quizzes disagree, a few questions usually explain more than retaking either one.

  1. Read the questions, not only the label. If one quiz's introvert questions are about energy and the other's are about sociability, they measured different things.
  2. Check the format. A scale from disagree to agree and a pick-one question can rank the same person differently, for the reasons above.
  3. Notice the setting you pictured. If you answered one quiz with work in mind and the other with weekends in mind, each result may be right about its own setting.
  4. Look at the margins. A result won by one answer is weak evidence either way, and the counts say more than the name at the top.

Two results that disagree aren't proof that personality quizzes are useless, or that one of them got you wrong. More often they were asking slightly different questions. The quizzes here put a note under every question that says why its answers point where they do, so you can check what each one is asking.

Sources

  1. Pace, V. L., and Brannick, M. T. (2010). How similar are personality scales of the "same" construct? A meta-analytic investigation. Personality and Individual Differences, 49(7), 669 to 676. Accessed October 10, 2026.
  2. Soto, C. J., John, O. P., Gosling, S. D., and Potter, J. (2008). The developmental psychometrics of Big Five self-reports: Acquiescence, factor structure, coherence, and differentiation from ages 10 to 20. Journal of Personality and Social Psychology, 94(4), 718 to 737. Accessed October 10, 2026.
  3. Soto, C. J., and John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117 to 143. Accessed October 10, 2026.
  4. Myers-Briggs Overview. The Myers & Briggs Foundation. Accessed October 10, 2026.
  5. Lievens, F., De Corte, W., and Schollaert, E. (2008). A closer look at the frame-of-reference effect in personality scale scores and validity. Journal of Applied Psychology, 93(2), 268 to 279. Accessed October 10, 2026.
  6. Bing, M. N., Whanger, J. C., Davison, H. K., and VanHook, J. B. (2004). Incremental validity of the frame-of-reference effect in personality scale scores: A replication and extension. Journal of Applied Psychology, 89(1), 150 to 157. Accessed October 10, 2026.
  7. Heine, S. J., Lehman, D. R., Peng, K., and Greenholtz, J. (2002). What's wrong with cross-cultural comparisons of subjective Likert scales? The reference-group effect. Journal of Personality and Social Psychology, 82(6), 903 to 918. Authors' preprint. Accessed October 10, 2026.
  8. Brown, A., and Maydeu-Olivares, A. (2013). How IRT can solve problems of ipsative data in forced-choice questionnaires. Psychological Methods, 18(1), 36 to 52. Accessed October 10, 2026.
  9. Cao, M., and Drasgow, F. (2019). Does forcing reduce faking? A meta-analytic review of forced-choice personality measures in high-stakes situations. Journal of Applied Psychology, 104(11), 1347 to 1368. Accessed October 10, 2026.