LuckPicker

Picking Teams for Pickup Sports

A random split is fair to players and unfair to the match. Pick which one you want.

Pickup sport is the clearest case where the fair draw and the good outcome pull against each other: an unbiased 5v5 split gives every player equal odds of any team and puts three or four of the strongest four on one side about 52% of the time.

The mechanics of a snake draft and what a random split costs are on the splitter's own page. This is about the twenty minutes before anyone kicks anything, which is where pickup football actually goes wrong.

Nobody objects to the algorithm. They object to being rated.

Somebody has to be the one who rates people

This is the entire social cost of a balanced split and it is rarely acknowledged. A rating is a judgement about somebody made by somebody, and in a group of friends that is an awkward thing to be seen doing.

The workable arrangement is that one person does it, privately, and the numbers are never shown. That sounds evasive and it is the arrangement most established groups arrive at independently, because the alternative — a public rating discussion on a cold pitch — is genuinely worse than an unbalanced game.

Rotating who rates is worth doing where the group is stable enough. It stops the ratings calcifying around one person's opinion from two seasons ago, which is the failure mode nobody notices because the teams keep coming out plausible.

What to do about self-ratings

Asking people to rate themselves feels fairer and produces worse teams. Self-ratings compress toward the middle from both directions: strong players understate to avoid seeming arrogant, and weaker players overstate to avoid being the one everyone knows is weak.

Since it is exactly the extremes that determine whether a match is even, compressing them is the worst available distortion. A group of ten where everyone rates themselves a six or seven produces teams that are balanced on paper and lopsided on grass.

If self-rating is politically unavoidable, ask for a rank order rather than a score. People are much better at saying who is better than whom than at placing themselves on a scale.

Why the extremes matter most

  • Ten players, true ratings spread 2 to 9.
  • Self-rated, the same ten cluster between 5 and 7.
  • Balancing on the compressed numbers is close to balancing at random.

The positions problem, on a real pitch

A single rating cannot express that somebody is a defender. Balance the numbers and one side can end up with every player who will actually track back, which produces an even total and a one-sided match.

The practical fix takes about a minute: split each position group separately and combine. Defenders into two, midfielders into two, forwards into two — three runs, and the teams are even on both axes.

What does not work is folding position into the rating, which people try first. A number that means both good and defensive cannot be balanced on, and the resulting teams are even on a quantity that means nothing.

When to stop balancing

There is a group size below which none of this is worth doing. With six or eight players who all know each other, a random split and a willingness to swap one person after ten minutes is faster and produces a better evening than any rating exercise.

There is also a skill spread below which it is unnecessary. If everyone is roughly comparable, a random split will be fine most of the time and the occasional lopsided game is not worth the social overhead of a rating system.

Balancing earns its place specifically where the spread is wide and known — one or two players clearly stronger than the rest — which is precisely the case where a random split produces the match everybody remembers for the wrong reason.

Frequently asked questions

Who should do the ratings?

One person, privately, with the numbers never shown. A public rating discussion on a cold pitch is worse than an unbalanced game.

Why not let people rate themselves?

Self-ratings compress toward the middle from both directions, and it is exactly the extremes that determine whether a match is even.

Is there an alternative to self-scoring?

Ask for a rank order. People are much better at saying who is better than whom than at placing themselves on a scale.

How do I handle positions?

Split each position group separately and combine — three runs for defenders, midfielders and forwards. It takes about a minute.

Can I fold position into the rating?

No. A number meaning both good and defensive cannot be balanced on, and the resulting teams are even on a quantity that means nothing.

Should ratings be revisited?

Yes, and rotating who does them helps. Ratings calcify around one person's opinion from two seasons ago without anyone noticing.

When is balancing not worth it?

Small groups who know each other, or a narrow skill spread. A random split plus a willingness to swap someone after ten minutes is faster.

When does it clearly earn its place?

A wide, known spread — one or two players clearly stronger than the rest. That is exactly when a random split produces a memorable one-sided game.

Tools for this job

Background reading

← All use cases