Blog
Small Samples Lie, and So Does Your Gut
Twenty draws tell you almost nothing, and your intuition will insist otherwise.
Twenty observations tell you almost nothing, and your intuition will insist otherwise with considerable confidence.
This is a short account of how many you would actually need, and why the wrong answer feels so convincing.
What twenty observations can distinguish
Suppose you suspect a picker favours one of five options. You watch twenty draws and that option comes up seven times against an expected four.
The typical deviation on twenty draws with a one-in-five chance is about 1.8, so seven against four is about 1.7 standard deviations out. That happens by chance roughly one time in eleven, which is not remotely enough to conclude anything.
To detect a genuine shift from 20% to 35% with reasonable confidence, you would need somewhere around a hundred and fifty observations. Twenty is not a small sample for this question — it is the wrong instrument entirely.
Why the wrong answer feels right
Human pattern detection is tuned for a world where most processes have structure, and it is very good at finding it. Applied to genuinely random data it finds structure anyway, because random data contains runs, clusters and apparent trends by construction.
The failure is asymmetric in a way that matters: we notice the pattern and do not notice the many opportunities it had to appear. Seven-out-of-twenty is striking; the fact that you were watching five options, any of which could have produced a striking count, is not.
That is the same error behind coincidences generally. The specific event is unlikely and the class of events like it is large, so something from the class happens almost every time.
The square-root rule
The single most useful thing to carry is that the typical deviation in a count of n observations is about the square root of n.
Twenty draws: typical deviation about 4.5 on a count. A hundred: about 10. Ten thousand: about 100. Note what that means proportionally — the absolute error grows and the relative error shrinks, which is the whole of the law of large numbers in one sentence.
Applied to a fifty-fifty question: two hundred coin flips has a typical deviation of about seven heads, so anything from 93 to 107 is unremarkable. A thousand flips has a typical deviation of about sixteen, so 484 to 516 is unremarkable.
How to actually check something
Decide the sample size before you start looking, and stick to it. Stopping when you have seen enough to confirm your suspicion, or continuing until you do, both invalidate any conclusion — and both feel entirely reasonable at the time.
Count frequencies rather than noticing runs. A frequency table over a few thousand observations is far more sensitive to bias than any amount of watching for streaks.
And compute what you would expect before you look at what you got. A prediction made in advance is a test; a threshold chosen after seeing the data is not.
The uncomfortable implication
Most of the confident claims people make about random processes are made from samples far too small to support them. That includes claims that a tool is rigged, and it includes claims that a tool is fine.
Which means the honest response to "is this picker biased" is usually that a few dozen observations cannot tell you, in either direction — and that if you actually want to know, the frequency test is a few thousand draws and an afternoon.
This is not a satisfying answer and it is the correct one, which is roughly the position most questions about randomness end up in.
The asymmetry nobody applies consistently
There is a bias in how the small-sample argument gets used, and it is worth being honest about because everyone does it including people who know better.
The demand for more data appears when a result is unwelcome and disappears when it is not. Twenty observations suggesting a tool is fine are accepted; twenty suggesting it is broken produce a request for a proper test. The evidential standard moves with the conclusion, which is exactly the failure mode statistical thinking is supposed to prevent.
The discipline that helps is deciding what would change your mind before looking. Writing down the number of observations and the threshold in advance turns a moving standard into a fixed one, and it is uncomfortable in exactly the way that suggests it is working.
Frequently asked questions
How many observations do I need?
Far more than feels necessary. Distinguishing a shift from 20% to 35% with reasonable confidence takes around a hundred and fifty.
What can twenty observations tell me?
Very little. Seven of twenty against an expected four is under two standard deviations out and happens by chance about one time in eleven.
What is the square-root rule?
The typical deviation in a count of n observations is about the square root of n. Two hundred coin flips has a typical deviation of about seven heads.
Why does my intuition disagree?
Because pattern detection is tuned for structured processes and finds structure in random data anyway — and we notice the pattern without noticing its many chances to appear.
Should I stop once I have seen enough?
No. Fixing the sample size in advance is what makes the conclusion valid; stopping early or continuing until you find something both invalidate it.
Is watching for streaks a good test?
No. A frequency count over a few thousand observations is far more sensitive to bias than any amount of watching for runs.