What makes a great quiz question

Here's a bad question. "Which of these is the largest animal? A) Blue whale B) African elephant C) Giraffe D) A banana."
It's bad in four separate ways at once, and every one of them is a mistake we've made ourselves and had to go back and fix. The banana is doing nothing except insulting the player. The elephant and the giraffe aren't close enough to be tempting. Nobody has said whether "largest" means longest, heaviest or tallest. And even if you get it right, you leave knowing exactly what you knew when you arrived.
Anyone can ask a question. Writing one that's actually worth answering is a craft, and after a few thousand of them we've become fairly opinionated about how it works.
A question is mostly its wrong answers
The correct answer is the easy part. You look it up, you check it against a second source, you move on. It takes a couple of minutes.
The three wrong ones take the rest of the afternoon, because they're what decide whether the question measures anything. They have a proper name in the assessment world — distractors — and the name is honest about the job. Each one is there to catch a specific way of being wrong.
The test is simple and unforgiving. Can somebody who knows nothing about the subject eliminate options just by looking at them? If yes, the question isn't measuring knowledge. It's measuring exam technique, which is a real skill and not the one anybody came here for. Assessment people call this test-wiseness, and a question vulnerable to it is broken no matter how good the correct answer is.
Good distractors are near-misses. If the answer's the pancreas, the others are organs. If the answer's 1066, the others are plausible medieval dates and not 1912. If the answer's Lisbon, the others are European capitals of similar size and prominence — not Lisbon, Paris, Tokyo and a village.
Where a good wrong answer comes from
The best distractors aren't invented. They're collected. Each one should be an answer a reasonable, half-informed person would genuinely give, and there are only a handful of places those come from.
- The famous near-thing. The biggest city that isn't the capital. Sydney, Istanbul, Rio, Casablanca — these do more work than any invented option ever will.
- The person who gets the credit. Edison for the light bulb. Bell for the telephone. Whoever popularised the thing rather than whoever built it first.
- The date everyone remembers slightly wrong. Usually a year or two off, or the year of the announcement rather than the discovery.
- The thing in the same category that's better known. The other planet, the other war, the other Brontë.
- The literal reading. The answer you'd get if you took the word at face value rather than knowing the term of art.
If you can't say out loud what confusion a particular distractor is catching, it isn't a distractor. It's filler, and it's making the question easier without making it more interesting.
The tells that give it away
Question writers leak information without meaning to, and the leaks are depressingly consistent. These are the ones we check for every time.
Length. Writers hedge the correct answer because they want it to be defensible, so it ends up longer and more qualified than the others. "A large flightless bird native to Australia" against "an eagle", "a bat", "a swan". The longest option is the answer far more often than chance allows, and players notice.
Grammar. If the stem ends with "an" and only one option starts with a vowel, you've written a giveaway. Same with singular and plural agreement. Rewrite the stem so it ends cleanly, or make every option fit.
Absolutes. Options containing "always", "never", "all" or "none" are usually wrong, because reality rarely cooperates that neatly, and experienced players know it. If you're going to use an absolute, use it in the correct answer sometimes.
The joke. A four-option question with a comedy option in it is a three-option question wearing a hat.
The odd one out. Three options in one register and one in another — three place names and a person, three round numbers and one precise one. Whichever one is shaped differently draws the eye, and if that's the answer you've told everyone.
Why "all of the above" is a weak option
We don't use it, and the reasons are worth setting out because it looks so harmless.
First, it lets partial knowledge score full marks. If a player is sure that two of the three listed options are correct, "all of the above" is proved right without them knowing anything about the third. That's the opposite of what the question is meant to do — it rewards knowing two-thirds of the material with the same mark as knowing all of it, and it can't tell the difference.
Second, and equally fatal, it lets partial knowledge score zero unfairly. A player who knows one option is definitely false can eliminate "all of the above" instantly, which narrows a four-way question to a two-way guess for someone who knows almost nothing.
Third, it's usually the answer. Writers reach for it at the end of a long session when they can't think of a fourth option, and players work that out within about ten questions.
"None of the above" has a related problem. When it's correct, the player has demonstrated that four things are wrong without ever demonstrating they know what's right — which means the question has confirmed an absence rather than a presence. That's occasionally what you want. It usually isn't.
While we're here: the research on how many options a question needs keeps landing somewhere uncomfortable for people who like tidy fours. Three well-built options generally do the same measuring job as four, because the fourth is nearly always the weak one nobody picks. We still write four, mostly because a four-option grid looks right on a screen — but the fourth has to earn its place like the others, not fill a slot.
Ambiguity kills a question outright
If two answers can both be defended, the question is broken. Not clever, not tricky, not "testing careful reading". Broken. And the player who got it wrong is right to be annoyed.
The usual culprit is a hidden qualifier — a word doing silent work that the player had no reason to notice. "What's the largest animal?" silently means largest living animal, or largest ever, and those have different answers. "What's the longest river?" has a genuine, unsettled disagreement behind it. "Who invented the telephone?" has a patent dispute attached. "How many countries are there?" depends entirely on who's counting.
The fix costs three words and no difficulty. Say which you mean. The person who knows the subject finds the question exactly as hard as before; everyone else stops being punished for a reading of the sentence that was perfectly reasonable.
The other form is false precision. There's a real difference between knowing the Moon's surface gravity is about a sixth of Earth's and knowing it to three decimal places. One of those is knowledge and the other is a lookup, and asking for the second doesn't make a question harder, only more annoying. Ask for the figure at the resolution a person would actually carry it.
And avoid double negatives entirely. "Which of these is not a country that has never had a monarchy" is a reading comprehension test with a geography costume on.
Difficulty and obscurity aren't the same thing
Making a question hard is trivial. Ask about something almost nobody has encountered and your success rate drops immediately. That's obscurity, and it produces a very specific kind of deflation: the player learns only that there exists a fact they had no route to.
Real difficulty is different. It comes from questions where the player very nearly knows the answer — where they have a belief, the belief is reachable, and it's wrong.
"Which planet is hottest?" is the classic. Almost everybody reaches confidently for Mercury, because it's nearest the Sun and that's how heat works. The answer is Venus, because a thick carbon dioxide atmosphere traps heat and Mercury has essentially no atmosphere at all. Get it wrong and you've gained a mechanism you can use elsewhere. Get it right and you feel the specific pleasure of a fact that earned its keep.
Misconceptions make the best material for exactly this reason. The player arrives holding a belief and leaves holding a corrected one, which sticks far better than filling an empty slot. There's a documented effect here — errors made with high confidence get corrected more reliably than errors made with low confidence, provided the correction actually arrives. Being confidently wrong and then told why is one of the more efficient ways a person learns anything.
Which brings us to the part most quizzes skip.
The explanation is half the question
A question you get wrong should feel like a small gift rather than a trap, and that's entirely down to what appears after you answer.
This isn't sentiment. Retrieval — being asked and having to produce an answer from memory — is one of the strongest known ways to make something stick, considerably stronger than reading the same material again. Adding a correction on top of the retrieval is what turns a wrong answer from a dead end into the useful part.
So the explanation changes how a question gets written, and it changes it early. If we can't write an interesting explanation, we usually bin the question.
Take "what's the chemical symbol for gold?" On its own that's arbitrary notation. The explanation is what makes it worth asking: Au, from aurum, the Latin for gold — which is also why silver is Ag, iron is Fe and lead is Pb, all of them Latin rather than English. One question, and suddenly the whole periodic table's odd corners have a reason. Next time someone meets Sn or W they've got a way in.
We have a blunt internal test for this. If the explanation boils down to "because that's the answer", the question isn't finished.
Faults that keep coming back
- Two defensible answers. Almost always loose wording. "Largest" without saying by what.
- The answer sitting in the stem. A word repeated in only one option, or grammar that fits only one.
- Facts with a shelf life. Current record holders, populations, office-holders, tallest buildings. These rot quietly and are wrong for months before anyone reports it. Prefer facts that were settled a century ago.
- Testing reading rather than knowing. Long stems, nested clauses, negation. If the hard part is parsing the sentence, the question is about the sentence.
- Trivia with nothing behind it. No mechanism, no story, no correction — just an isolated fact you either happened to have or didn't.
That last one is the one we guard hardest against, because it's the easiest kind of question to produce in volume and the least worth reading. A hundred of them can be written in an afternoon. None of them will be remembered by Thursday.
How we actually check one
The process is dull, which is rather the point.
Read the stem with the options covered and try to answer it cold. If it can't be answered without seeing the options, it's a recognition test rather than a knowledge one, and it'll behave differently from how you expect. Then look at the four options and ask, honestly, whether a well-informed person could argue for a second one — not a pedant, a well-informed person. Then check the lengths and the grammar. Then read the explanation and ask whether it says anything.
Then, once the thing is live, look at what players do with it. Two numbers tell you nearly everything. The proportion who get it right shows whether the difficulty is where you thought. More usefully, whether the people who score well overall get this question right more often than the people who score badly — if a question is answered correctly just as often by weak players as by strong ones, something in it is wrong regardless of how good the fact is. That's usually a distractor problem, and it's usually visible the moment you look at which wrong option people are picking.
A distractor nobody ever selects is dead weight. A distractor selected more often than the correct answer is either a brilliant question or a mistake, and you have to go and find out which.
When to break all of it
A quiz made entirely of hard, surprising, misconception-correcting questions would be exhausting to play and slightly smug to read.
Easy questions do real work. They set a rhythm, build a bit of confidence and give a player somewhere to stand before the ground tilts. A quiz that never lets you feel clever isn't a good quiz however rigorous it is, and opening with your hardest question is a reliable way to lose people at question one.
So the rules describe the average rather than the individual case. Across ten questions we want a shape — two that nearly everyone gets, two that nearly nobody does, and a middle stretch where knowing the subject genuinely pays. The order matters as much as the mix.
Get that right and the thing stops feeling like a test. It starts feeling like a conversation with someone who knows a lot and would rather show you something than catch you out.


