← notesLearning Through Imitation · core lessonAgranov, Lopez-Moctezuma, Strack, Tamuz · why redundant guesses make a group smarter15 slides · 9.5 min at 1× · built 2026-09-17 06:44
Paused so you can answer in your head. Press space for the answer.

1 of 15 · Q: what does the paper find?

Seeing each other's guesses makes a group more accurate, though the guesses add no information

group A sees: every draw, by everyonegroup B: same draws+ everyone's past guessesthe guesses are computed from the draws.To a Bayesian they carry zero new informationabout the urn.Standard worries say they should hurt: herding, double counting, overload.share of correct guesses, round 20A · draws only≈ 0.84B · draws + guesses≈ 0.88group B does better, and agrees more

Two groups are trying to learn the same hidden fact. Both groups see every piece of raw evidence that anyone collects. The second group also sees what everyone guessed in earlier rounds.

Those guesses were computed from the same evidence. To a perfect reasoner they carry no new information at all. And the usual worries in this literature say they should hurt: herding, double counting, information overload.

In the experiment they help. The group that sees guesses ends up right more often, and it agrees with itself more.

2 of 15 · Q: what is the task?

An urn, eight players, twenty rounds: guess, then draw

the urn: 6 of one color, 4 of the otherwhich color leads is a fair coin flip123456788 players, 20 rounds, same urn all gameeach round, each player1 · guessesthe urn's majority color2 · draws one ballright color with probability 0.6pay: one random round of one random game,$20 if correct, $5 if wrongso the best play each round is simply your best guess.A table on screen keeps the full history: memory is not tested.606 subjects, 31 sessions, pre-registered

Here is the task. An urn holds ten balls. Six are one color and four are the other, and a fair coin decides which color has the majority.

Eight players face the same urn for twenty rounds.

In every round each player first guesses the majority color, and then draws one ball with replacement. A single draw points the right way only sixty percent of the time, so one ball tells you little and a pile of them tells you a lot.

One random round is paid: twenty dollars if the guess was right, five if it was wrong. So there is nothing to hedge, and the best play each round is simply your best guess. A table on screen keeps the full history.

3 of 15 · Q: what is varied?

Four information conditions, and one comparison that matters most

others' drawsothers' past guesseseveryone always sees their own drawsno infohiddenhiddenactionshiddenseessignalsseeshiddenallseesseesthe headline comparisonsame information about the urnprivate-information sidecan people decode guesses alone?plus actions and all again with groups of 4, to test group size

The experiment varies what you can see about the other seven players. In the signals condition you see all of their draws. In the all condition you see their draws and their past guesses.

That pair is the headline comparison, because both conditions give exactly the same information about the urn.

Two more conditions cover private information. In no info you see only your own draws. In actions you see your own draws plus everyone's past guesses, and never their draws.

Finally, the actions and all conditions are repeated with groups of four, to test whether group size matters.

4 of 15 · Q: what would a perfect reasoner do?

A Bayesian just counts, reaches 0.96 by round ten, and ignores everyone's guesses

0.50.60.70.80.91.015101520roundprobability the guess is correctthe Bayesian rulecount reds minus greens overeverything you can see.Guess the color that is ahead.own draws only 0.81all 8 players' draws: 0.96 by round 10for a Bayesiansignals = allothers' guesses are functions ofdraws she already sees, so theycannot move her posteriorany gap between the two is behavioral

What would a perfect reasoner do? She would count. Reds minus greens, over every draw she can see, and she guesses whichever color is ahead.

With only her own draws that is slow. With all eight players' draws pooled it is fast: ninety six percent correct by round ten, and above ninety nine by round eighteen.

Notice what this implies. In the all condition, other people's guesses are functions of draws she already sees. They cannot move her beliefs. For a Bayesian, signals and all are the same experiment, so any gap between them measures something about real human reasoning.

5 of 15 · Q: what happened?

Both groups fall far short of Bayes, and the group that sees guesses does better

the paper's Figure 1. Purple: all. Blue: signals. Green dashes: Bayesian benchmark.accuracyboth groups fall far short of 0.99.all ends near 0.88, signals near 0.84consensussize of the majority: all 0.89,signals 0.79 by the endeffect of seeing guesses, with 95% intervals0accuracy+0.051 (SE 0.028)consensus+0.095 (SE 0.028)treatment varies by session: 11 sessions here

Here is the paper's main figure. Purple is the group that sees guesses. Blue is the group that sees only draws. The dashed green line is the Bayesian benchmark.

On the left is accuracy. Both groups fall far short of the benchmark. They end near eighty eight and eighty three percent, where a Bayesian would be above ninety nine. And purple sits above blue.

On the right is consensus, the size of the majority. Seeing guesses pulls the group together.

How big are these effects? About five points on accuracy and about ten on consensus. Be careful with the first one. Its standard error is almost three points, from eleven sessions, so its ninety five percent interval runs from just below zero to about ten and a half points. It passes at the ten percent level and misses at the five percent level. The consensus effect is the statistically firm one.

6 of 15 · Q: where do people go wrong?

People are a smooth S where a Bayesian is a step, and the damage is on close tallies

a Bayesian is a stepone ball ahead: always guessthat colorpeople are a smooth Sa lead of two: about 2 in 3 follow it.Even 30 to 40 balls ahead, only80 to 90 percent follow the data.both groups are worst in the middleweak tallies: 0.73 optimal withsignals, 0.79 with all.all wins in every bin by 3 to 6 points.

People go wrong in one specific place. This figure plots how often subjects guess red against how far red is ahead in the tally.

A Bayesian is the dashed step. One ball ahead is enough to commit.

Real subjects trace a smooth S. When the tally is close they are barely better than a coin flip: with a lead of two, only about two in three follow it. Even with one color thirty or forty balls ahead, only about eighty to ninety percent follow the data.

The right panel shows the chance of the optimal guess by how lopsided the evidence is. Seeing guesses helps in every bin, and the weak bin is where both groups are worst: seventy three percent optimal with draws alone, seventy nine with guesses as well.

7 of 15 · Q: when do people imitate?

Imitation switches on when the evidence is close, and off when it is lopsided

x: share of others who guessed red last roundy: probability you guess red nowlopsided tally: flat lineswhat others did barely mattersclose tally: the steep yellow line0.13 to 0.86 as the share goes 0 to 1a caution from our surveyinside the close bin, the tally and theothers' share still move together, so thisslope overstates imitation. The model'sstructural effect is 0.31.

When do people imitate? This is the mechanism figure. The horizontal axis is the share of other players who guessed red last round. The vertical axis is the chance that you guess red now.

When the tally is lopsided the lines are flat. You follow the data, whatever the others did.

When the tally is close, the yellow line is steep. Your chance of guessing red runs from thirteen percent to eighty six percent as the others go from all green to all red.

One caution from our own survey of the paper. Inside the close bin, the tally and the others' guesses still move together, so this slope mixes the two. The structural model puts the pure effect of a peer, at an even tally, at about thirty one points. That is still large.

8 of 15 · Q: how can redundant information help?

A peer's guess is a second noisy read of the same tally

the tally Spublic, exactyou: a noisy decoderright with prob. your guessa peer: another noisydecoder of the same Stheir last guess A no new information about the urn enters.Two noisy reads of one number beat one.when S is near zeroyour own read is a coin flip,and the peer term decides when S is largethe tally term dominates andthe peer is ignored“actions are redundant”assumes a noiseless decoder.Real people are noisy ones.our phrasing, not the authors'

How can information that is redundant help? Think of each subject as a noisy decoder. The tally is public and exact. The subject reads it with noise, and the closer the tally, the noisier the read.

A peer is another noisy decoder of almost the same number, one round earlier. Their last guess is a second read.

The paper's model says you respond to a weighted sum of the two, your read of the tally and one sampled peer's last guess. Nothing new about the urn enters. Two noisy reads of one number simply beat one.

This also explains the switch. When the tally is near zero your own read is a coin flip, and the peer decides. When the tally is large it dominates, and the peer is ignored. The argument that guesses are redundant quietly assumes a reader with no noise.

9 of 15 · Q: what is the model?

Three parameters: weight on the tally, weight on a peer, and how the tally is scaled

S = 0tally strongly redstrongly green10.5probability of guessing redsampled peer said redpeer said green 0.31at a dead-even tally, flipping thepeer moves you by 0.31 the Bayesian count the sample proportion a lead persuades less as the pile growsestimates, all treatment

Here is the model in full. The probability of guessing red is a logistic function of beta times the tally signal, plus gamma times the sampled peer's last guess. The two curves are the two possible peers.

The vertical distance between them at an even tally is the pure imitation effect: thirty one points.

The third parameter, psi, scales the tally. The signal is reds minus greens, divided by the number of draws to the power psi. Psi equal to zero is the Bayesian count. Psi equal to one is the sample proportion. The estimates land between one half and about zero point six, so a fixed lead persuades less as the pile of evidence grows.

In the all condition the estimates are beta one point two seven, gamma zero point six five, and psi zero point six three. Simulated at these values, the model reproduces the ordering of the two groups and lands within about two points of the data in the last round. In the middle rounds it runs about five points low.

10 of 15 · Q: how much imitation is best?

In the model, accuracy peaks at a peer weight around two; subjects sit at 0.65

0.600.700.800.90012345 round-20 accuracy in the paper's model (our simulation)no imitation: 0.82subjects: 0.65best: about 0.90, a flat top around 2herdingreading the curvesome imitation helps.More helps more, up to a point.Past it the group copies its ownearly mistakes and stopslistening to the data.This sweep is ours. The paperreports the estimates and leavesthe optimal weight as future work.

How much imitation is best? We ran the paper's model ourselves, holding the other parameters at their estimates and sweeping the weight on a peer. With no imitation, accuracy in round twenty is about eighty two percent.

Subjects sit here, at zero point six five.

Accuracy keeps rising beyond that point. It peaks at about 90 percent, on a flat top around a weight of two, three to four times what subjects use. That is for the last round. Averaged over all rounds, which is what subjects are paid on, the best weight is lower, about one and a half, and the gain over what subjects do is about two points.

Push the weight higher and accuracy collapses. The group starts copying its own early mistakes and stops listening to the data. That is herding. So in this model people imitate less than would be best for the group. This sweep is ours. The paper leaves the optimal weight as future work.

11 of 15 · Q: who benefits?

Stronger subjects read the data better, and imitation is not a crutch for the weak

weight on the data, low IQ score0.99high IQ score1.64stronger subjects read the tally much betterweight on a peer, low IQ score0.61high IQ score0.68the model cannot tell these apart;a subject-level measure finds higherscorers imitate somewhat more (p = 0.004)test-score split: gains of 4.7 and 3.7 points, difference 1.0 (SE 3.1).performance split (Appendix L): the gain is largest near the 25th percentile.

Who benefits? The paper splits subjects by a short reasoning test. The weight on the data differs a lot: about one for the lower scoring group, and one point six for the higher scoring group.

The weight on a peer differs much less: zero point six one against zero point six eight, and the model cannot tell them apart. A second measure in the paper, built subject by subject, does find a difference: higher scorers lean on others somewhat more when the tally is close. So imitation is not something only weak subjects do. If anything the stronger subjects do more of it.

Both groups gain in the point estimates, about five points for higher scorers and about four for lower scorers, though once the sample is split neither gain is statistically firm on its own. Split by test score, weaker subjects do not gain more. Split by how well people actually play, the paper's appendix finds the opposite: the gain is largest in the lower part of the distribution, around the twenty fifth percentile.

12 of 15 · Q: what about private information and group size?

Guesses alone teach a lot, and a bigger group helps only when raw draws are shared

round-20 accuracy, from the paper's Figure 5no info · own draws onlyabout 0.66actions · own draws + others' guessesabout 0.79signals · everyone's drawsabout 0.83guesses alone carry real informationdecoding private draws from repeatedguesses is a hard inference problem,and people still get a large gain from itgroup of 8 versus group of 4all: bigger group learns fastermore raw draws land in the shared tallyeach round (+6.5 points, SE 2.3)actions: group size makes no difference+0.8 points (SE 2.2). The authors link this,loosely in their own words, to a theoremon rational groupthink.

Now the private information side. With only your own draws you reach about sixty six percent by round twenty. Add other people's guesses, and never their draws, and you reach about seventy nine. By the last round that is about three quarters of the way to seeing the draws themselves. Averaged over the whole game it is half way.

That is a large gain from a hard inference problem: working out what others must have drawn from how their guesses change over time.

Group size tells the two channels apart. When draws are shared, eight players learn faster than four, by about six and a half points, because more raw evidence lands in the tally each round. When only guesses are shared, the estimated difference is under one point. The authors connect this to a theorem on rational groupthink, and they call that connection loose themselves. Our own check: the interval runs from minus three and a half to plus five points. That just reaches the four to five point gain that Bayesian groups would show, so the data cannot rule it out.

13 of 15 · Q: how solid is it, and where next?

A real mechanism with a modest headline effect and two open design questions

the accuracy effect is modest+0.051 with SE 0.028: significantat the 10% level, 11 sessions.The consensus effect is the firm one.partly handledlab versus onlinesignals ran only online; all ran inboth. The paper never tests the gapwithin online sessions alone.openrandom-acting subjectsabout a quarter of subjects in twotreatments act as if random; theheadline is not re-run without themopenthe screen shows no running tallysubjects see a table of draws and mustcount. Part of the counting noise thatimitation repairs comes from the interface.our inference

How solid is all this? The mechanism has good support: imitation rises as the tally gets closer, in every cut of the data. One caveat is that the paper's own decomposition leaves about six points of the treatment gap unattributed. The headline accuracy effect is modest: five points, significant at the ten percent level, from eleven sessions.

Two design questions stay open. The signals condition ran only online, while the all condition ran both online and in a lab, and the paper never compares them within online sessions alone. And about a quarter of subjects in the no info and signals conditions act as if random, while the headline is never re-run without them.

One more observation, which is our inference. In the screenshots the paper prints, which come from the all condition, the screen shows a table of draws and no running tally, so subjects have to count. Some of the counting noise that imitation repairs comes from the interface.

14 of 15 · Q: can you reconstruct the two main ideas?

Check yourself: two questions before going deeper

question 1Why would a perfect Bayesian do exactlyas well in the signals condition as inthe all condition?answerin both she sees every draw. Others'guesses are functions of those same drawsplus noise unrelated to the urn, so theycannot change her posterior.question 2In the mechanism figure, why are thetop and bottom lines flat while themiddle line is steep?answerthe guess follows beta times the tally plusgamma times one peer. With a lopsided tallythe first term dominates. With a close tallyit is near zero and the peer term decides.

Before the wrap up, two questions. Pause after each one and answer it in your head. First. Why would a perfect Bayesian do exactly as well in the signals condition as in the all condition?

The answer. In both conditions she sees every draw. Other people's guesses are functions of those same draws, plus noise that has nothing to do with the urn. Conditioning on them cannot change her beliefs.

Second question. In the mechanism figure, why are the top and bottom lines flat, while the middle line is steep?

The answer. The guess responds to beta times the tally signal, plus gamma times one peer. When the tally is lopsided, the first term dominates and the peer hardly matters. When the tally is close, the first term is near zero, and the peer term decides.

15 of 15 · Q: what should I take from it, and where next?

A redundant channel of conclusions is a cheap denoiser, within limits nobody has tested

where the result should carry over (our reading)a redundant channel of peers' conclusionsis a cheap denoiser when• every reader decodes the evidence with noise• their errors are independent• the weight on peers stays moderatenever testeda peer who is confidently wrongevery peer in the experiment was an honest,noisy reader of the same data. The authors listagents whose guesses ignore the data as future work.read next1 · the benchmark and the modelTable 2 re-derived, the three parameters, the unitsinconsistency, what Figure 3 overstates, the optimal weight2 · the auditeffect sizes with intervals, lab versus online, random-actingsubjects, the interface, a verdict table

Our reading of the result: a redundant channel of other people's conclusions is a cheap denoiser when every reader is noisy and their errors are independent. It was never tested against a peer who is confidently wrong. The authors list something close as future work: agents whose guesses ignore the data.

Two deep dives follow. One builds the benchmark and the model in detail. The other audits the evidence.

1.00×
keys
space play / pause
slide · , . beat
[ ] speed · c captions · d deeper
m mute · f fullscreen · r replay slide