What a Close Split on a Personality Quiz Tells You

By Big Time Trivia Editors. Published , updated .

You answer every question, ask for your result, and it comes back with a letter you didn't expect, an x where a letter should be, or two names sharing the top spot. The count underneath can explain it: one pair split 3 to 2, or one outcome finished a single answer ahead of the next. A close split is easy to misread in both directions, as a firm answer or as a sign the quiz broke. It's neither. This article covers what measurement error is, what test-retest reliability checks, why a 3 to 2 split on five questions settles very little, how the quizzes on this site report close results, and how to read any quiz result with all of that in mind.

Every score carries some error

The Standards for Educational and Psychological Testing, prepared by a joint committee of the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, make a plain observation early in their chapter on reliability: "different samples of performance from the same person are rarely identical." A person's responses to test questions vary, in the Standards' words, "from one sample of tasks to another and from one occasion to another, even under strictly controlled conditions."

The Standards sort the causes of that random error into two groups: those inside the person and those outside. Inside are changes in "motivation, interest, or attention." Outside are differences in testing conditions, such as the "time of day" or the "level of distractions." None of it can be subtracted after the fact: "Because random measurement errors are unpredictable, they cannot be removed from observed scores." It can be kept down to some extent, the Standards add, for example "by averaging over multiple scores."

Classical test theory gives the idea a name. Your true score on a test is "the hypothetical average score over an infinite set of replications of the testing procedure," as the Standards put it, and the score you actually get "fluctuates around the true score." Nobody takes a test an infinite number of times, so any score you see is one reading taken somewhere around that average. The size of the error can be summarized in several ways. One is the standard error of measurement, which "provides an indication of the expected level of random error over score points and replications for a specific population." The quizzes on this site haven't been through that kind of testing, so they have no such figure. What they can show is the count behind each result, which says how close the call was.

Personality questions add a layer of their own, because behavior itself moves around. In three experience-sampling studies that tracked behavior relevant to the Big Five traits over two to three weeks of everyday life, the psychologist William Fleeson found that within-person variability was high: "the typical individual regularly and routinely manifested nearly all levels of all traits in his or her everyday behavior." The same 2001 paper found that differences between people in their typical level of behavior were "almost perfectly stable." Both halves matter for a quiz. If behavior ranges that widely, a question about one situation catches one point in a wide spread, and any single answer can land some way from your usual. Your average is the steady part, and it takes many answers to see it.

What test-retest reliability checks

A basic way to see measurement error is to give the same test twice. The Standards describe the design: the test is given, then given again "after a brief period during which the examinee's standing on the variable being measured would not be expected to change," on the assumption that "the first administration has no influence on the second administration." When both conditions hold, "more variation across the two administrations indicates more error in the test scores and therefore lower reliability/precision." Reliability/precision is the Standards' term for how consistent scores are across repeated rounds of testing.

Both conditions matter. If the thing being measured has changed in between, a different score isn't error. The Standards say as much: some changes from one occasion to another "are not regarded as error (random or systematic), because they result, in part, from changes in the construct being measured." They also draw a line between a trait and a state. For "a state variable (e.g., mood or hunger), where fairly rapid changes are common, scores generated on two successive days would not be considered replications." And a retake five minutes after the first sitting makes it hard to claim the first had no influence on the second.

One retest study of a personality questionnaire shows the gap between a score built from many questions and a single question. In a 2022 study in PLOS ONE, Sam Henry, Isabel Thielmann, Tom Booth and René Mõttus had 416 people, recruited and paid through the online research platform Prolific, take the HEXACO-100, a 100-item questionnaire covering six broad personality domains, twice, about 13 days apart. They measured agreement between the two sittings as a correlation, where 1 would mean the scores lined up perfectly. The median was .88 for the six domain scores, .81 for the narrower facet scores, which have four items each, and .65 for single items. Every domain held steadier than every single item: the lowest domain figure, .86, was above the highest item figure, .84.

A quiz that sorts people into categories raises a second question, which is whether the categories agree as well as the scores. The Standards call this decision consistency, "the extent to which the observed classifications of examinees would be the same across replications of the testing procedure." Their Standard 2.16 says that when a test is used to classify people, estimates should be provided of "the percentage of test takers who would be classified in the same way on two replications of the procedure" (Standards, chapter 2). Consistent scores and consistent labels are different things, and the site's article on the science behind MBTI shows how far apart they have been in research on the Myers-Briggs Type Indicator.

Why the line is where the trouble is

Error doesn't threaten every label equally. In the Standards' explanation, test takers who are "far above or far below the cut score" can have "considerable error in their observed scores without any effect on their classification decisions." The people near the line are another matter: "Errors of measurement for examinees whose true scores are close to the cut score are more likely to lead to classification errors."

A quiz that hands out letters has a cut on every pair, at the point where one side's count passes the other's. Someone whose answers lean hard one way can lose an answer or two and keep the same letter. Someone near the middle can't, so their letter rides on exactly the kind of small, random variation described above.

A cut has a second cost, which is what it throws away. A letter keeps the direction of a lean and drops its size, so a narrow lead and a landslide come out looking the same, and two people one answer apart can end up on opposite sides. The site's article on MBTI and the Big Five goes through what researchers have said about turning a score into two groups.

Why fewer questions means more wobble

How many questions sit behind a result matters as much as where the line falls. In general, the Standards say, "if the assessment is shortened (e.g., by decreasing the number of items or tasks), the reliability is likely to decrease," and lengthening a test with comparable items "is an effective and commonly used method for improving reliability/precision." Each added question is one more reading to average over. The retest study above shows the same pattern from the other end. By the median, single items held up less well than four-item facets, and facets less well than the broad domains.

Very short measures cause trouble even in professional research. In 2012 Marcus Credé and colleagues noted that researchers often use "very abbreviated (e.g., 1-item, 2-item) measures of personality traits," and tested what that costs, using data from 437 employees and 355 college students. They found that the practice, "particularly the use of single-item measures," can lead researchers "to substantially underestimate the role that personality traits play in influencing important behaviors." They argued that "even slightly longer measures can substantially increase the validity of research findings."

A pair measured by a handful of quiz questions sits much nearer that short end than a research questionnaire does. Each answer is a large share of the total, so a single answer that goes the other way moves the result a long way.

Why 3 to 2 on five questions is weak evidence either way

Take a hypothetical pair measured by five questions, and a result of 3 to 2. It's the narrowest split five questions allow, since an odd number can't tie. It also fits several quite different people.

  • Someone with no lean at all. If each answer is as likely to go one way as the other, five answers end up 3 to 2, in one direction or the other, more often than any other way, because there are more ways to arrange three and two than four and one, or five and none. For someone sitting right in the middle, 3 to 2 is the result to expect.
  • Someone with a slight lean toward the side that got three. The result points the right way, but it looks exactly like the no-lean result above.
  • Someone with a clear lean toward the side that got three, on a day when one answer drifted. If everyday behavior ranges as widely as Fleeson's findings suggest, an answer that strays from your usual is no surprise.
  • Someone with a slight lean toward the side that got two. One answer drifting the other way is all it takes to hand the three to the other side.

Five answers can't tell these people apart, and that's what "weak evidence either way" means here. The 3 to 2 is weak evidence that you lean toward the three. It's also weak evidence that you sit in the middle. Compare a 5 to 0 split on the same pair: two answers could change and the same side would still lead. A 3 to 2 flips with one.

What would firm it up is more evidence, either more questions on the pair, for the reason the Standards give, or a second sitting on another day. A pair that keeps coming out close across sittings tells you more than any single split can.

What a close split doesn't mean

A close split isn't a malfunction, and it isn't a verdict that you're indecisive. The Myers-Briggs Type Indicator is a popular framework whose scientific standing is contested (the site's article on the science behind MBTI goes through the evidence), but on this point its publisher's guidance fits any quiz, because the publisher doesn't present every letter as equally certain. On the self-scorable version of its Step I assessment (Form M), each preference comes with a preference clarity category of "Slight, Moderate, Clear, or Very Clear." In a 2016 article on the publisher's site, Patrick Kerwin, an MBTI Master Practitioner, explains that clarity describes "how clearly or consistently a person chose his or her preference in each pair of opposites," which in turn tells "the likelihood that what is reported is the person's true preference."

A Slight result, Kerwin writes, "simply tells us that the person is less sure about whether that preference describes him or her." It "does not mean that he or she is good at using both preferences" or "uses both preferences well, or that he or she is confused!" A close split on any quiz deserves the same reading. It says the answers didn't settle the question. It isn't a finding that the person who gave them is undecided in general.

How the quizzes here report close results

Every result on this site shows the count behind it, so a narrow lead looks narrow, and close calls aren't settled quietly. How we build quizzes covers the method in general.

The personality type quiz scores the four Myers-Briggs preference pairs one at a time and shows the count on each. Two settings in the site's scoring code decide what happens near the line, and each is a share of a pair's answers. The first is the balanced line. If the side in the lead holds that share of the pair's answers or less, the quiz doesn't pick a letter for you. It calls the pair balanced, puts a lowercase x where the letter would go, and says which side the split leaned toward, with both counts. An even split is always balanced. The second is the close line. A letter won with that share or less gets a note calling it a close split that one answer could flip. When the two settings are the same, every split narrow enough to be close is already balanced, so no letter carries that note.

With one or two pairs balanced, the result lists the types closest to your answers instead of a single type. With more, so many types fit equally well that the quiz shows no type description and lets the split on each pair stand as the result.

Whether a split counts as balanced depends on how many questions the pair has, so the short and full versions can treat the same narrow lead differently. The "How this quiz works" section on the quiz page applies the rule to each version and names the splits it covers. If no split on a version can count as balanced, the page says so and says why. The same section works out how often random answers would leave a pair balanced (or, where no split can be balanced, close) and how often they would name a full four-letter type, and its last part shows how often changing a single answer changes the result. Those figures are recalculated from the quiz's own questions every time the site is built, so they stay correct if a question changes. Real answers aren't random, so the figures show how much room the scoring leaves for close calls, not what anyone's result is likely to be.

The Game of Thrones character quiz gives one point per answer to the character that answer points to and shows the count for every character, so a one-answer lead reads as one answer. If two or more characters share the top count, the result names all of them instead of breaking the tie. The quiz page works out how often random answers would end in a tie, how often changing one answer changes the result, and how few changed answers it takes.

The love language quiz works the same way. Each answer gives one point to the language it points to, the result shows the count for every language, and every language tied at the top is named. Its page reports the same figures for ties and changed answers. The site's guide to the five love languages explains why a close count is no surprise on that quiz, and why the whole spread is worth reading, not only the winner.

How to read any quiz result

These habits work on any quiz, here or anywhere else.

  1. Find the margin, then the number of questions behind it. A label is the last step of the scoring. The margin shows how much it rests on, and the number of questions shows how much each answer weighs. The fewer there are, the more one answer can swing, and shorter tests tend to be less reliable.
  2. Ask whether one answer would flip it. If it would, treat both sides as live. The question Kerwin suggests for a Slight result works on any quiz: "What might have pulled you in two directions when you were responding?"
  3. Retake it on another day, not right away. The retest design in the Standards assumes the first sitting has no influence on the second, which is hard to claim while your earlier answers are still fresh. If the result changes, the first one was probably close, or something changed in between. A passing state like mood can be one of those changes, and it isn't the same as a change in who you are.
  4. Read a balanced or tied result as unsettled, not as an answer. It says these answers didn't decide it. Reading both descriptions is a better use of it than picking one.
  5. Match the weight to the use. The Standards note that "the need for precision increases as the consequences of decisions and interpretations grow in importance." A quiz is fine for a conversation with friends. It isn't a basis for decisions about work, health, money or relationships, or for judging other people. For anything like that, a qualified professional is the right person to ask.
  6. Keep the description and the count apart. A paragraph can feel accurate whether or not the count behind it was close. The site's article on why personality tests feel so accurate explains why.

A close split isn't the quiz failing. It marks the place where a short set of questions ran out of evidence, and that's worth knowing too.

Sources

  1. American Educational Research Association, American Psychological Association and National Council on Measurement in Education, Standards for Educational and Psychological Testing (2014), chapter 2, Reliability/Precision and Errors of Measurement. Accessed October 9, 2026.
  2. William Fleeson, Toward a structure- and process-integrated view of personality: Traits as density distributions of states, Journal of Personality and Social Psychology 80(6), 1011-1027 (2001), PubMed record. Accessed October 9, 2026.
  3. Sam Henry, Isabel Thielmann, Tom Booth and René Mõttus, Test-retest reliability of the HEXACO-100: And the value of multiple measurements for assessing reliability, PLOS ONE 17(1), e0262465 (2022). Accessed October 9, 2026.
  4. Marcus Credé, Peter Harms, Sarah Niehorster and Andrea Gaye-Valentine, An evaluation of the consequences of using short measures of the Big Five personality traits, Journal of Personality and Social Psychology 102(4), 874-888 (2012), PubMed record. Accessed October 9, 2026.
  5. Patrick Kerwin, Clarifying Clarity, The Myers-Briggs Company (March 3, 2016). Accessed October 9, 2026.