← notesLearning Through Imitation · deep dive: does the headline survive an audit?a referee's pass through Appendices A to N · what is handled, what is partly handled, what stays open12 slides · 6.5 min at 1× · built 2026-09-17 06:44
Paused so you can answer in your head. Press space for the answer.

1 of 12 · Q: what is being audited?

The claim: seeing redundant guesses raises accuracy by about five points

Figure 1 of the paper, p.12. Purple: all. Blue: signals.Dashed green: the Bayesian benchmark, the same for both.same balls on both screenssignalsevery ball drawnallballs + all past betsheadline: all beats signals by about 5 pointsthe audit1how big, once uncertainty is attached2same place, same kind of subjects3is the gain produced by imitation4what the procedure itself could create

Groups of eight guess which of two urns was chosen, over twenty rounds. In signals, everyone sees every ball drawn. In the all treatment, they also see every bet the others have made. A Bayesian gains nothing from those bets, because the balls already say everything.

The headline is this gap. Accuracy in all runs about five points above signals, and the paper reads it as imitation repairing noisy counting.

This lesson audits that claim on four points. The size of the effect with its uncertainty. Where, and on whom, each treatment ran. Whether the gain can be attributed to imitation. And what the procedure itself could create.

2 of 12 · Q: how big is the effect, with its uncertainty?

Five points with a standard error near three, from eleven sessions

-0.050+0.05+0.10+0.15+0.20all minus signals, as a share of guessesno effect8 sessions of all3 sessions of signals treatment varies only across the 11 sessionsaccuracy5.1 points, SE 2.8Table A.1, clustered by session normal 95%t, 10 d.o.f. intervals: our arithmeticconsensus9.5 points, SE 2.8 clear of zero

The regression has forty four thousand rows, but treatment varies by session: eight of the all treatment, three of signals. By our count, the sample for this coefficient is eleven.

Table A one gives five point one points of accuracy, clustered standard error two point eight.

By our arithmetic, a ninety five percent interval runs from just below zero to ten and a half points, and widens with ten degrees of freedom. The paper prints one star, which in its other tables means the ten percent level, and the main text omits it.

The consensus effect, nine and a half points with the same standard error, is clear of zero. Agreement rises. Accuracy probably does.

3 of 12 · Q: were the two treatments run in the same place?

Every signals session ran online at Ohio State; half of all ran in a San Diego lab

no infoactionsallsignalstreatmentUC San Diegophysical lab4 sessions80 subjects8 sessions136 subjects4 sessions64 subjectsOhio Stateonline, after thelab closed4 sessions88 subjects3 sessions82 subjectsnever runthe headline:152 subjects from two sites against 82 from one. Treatment is tangled with mode and with subject pool.the clean contrast:the online row alone, 88 against 82, which is 4 sessions against 3.Footnote 5, p.8. It does not say where the two groups-of-four treatments ran.

Next, where the sessions ran. Data collection began in a physical lab at U C San Diego. No info, actions, and four sessions of the all treatment ran there.

Then the pandemic closed the lab. Its other four sessions, and every session of signals, ran online at Ohio State. One cell of the grid is empty. Signals never ran in the lab.

So the headline compares one hundred fifty two subjects from two sites with eighty two from one. As we read it, treatment is tangled with lab versus online, and with which university the students attend.

Only the online row holds the site fixed. That is four sessions against three.

4 of 12 · Q: does Appendix F remove the confound?

Appendix F tests whether all looks alike across sites, which is a different question

Figure F.1, p.54: the all treatment only, split by site.Purple: Ohio State online. Blue: UC San Diego lab.what Appendix F testsall in the lab against all online"p > 0.10 in all comparisons": no table, no coefficientwhat it never testsall online against signals onlinethe one comparison that holds the site fixedall minus signals (always online), rounds 2 to 20, pointsall, both sites pooled5.0all, online sessions5.4all, lab sessions4.6read off Figures 1 and F.1 by the survey agent7 sessions in the clean contrast, so no usable standard error

Appendix F is the paper's answer. Its figure splits the all treatment by site.

The curves sit close, and the text reports no significant difference without printing a coefficient. Four sessions against four gives little power.

The larger problem is the missing test. The headline needs all minus signals with site held fixed. Only the online sessions allow that, and the appendix never reports it.

Read off the printed figures by the survey agent, the online gap is about five point four points. That is an estimate from pixels, with no standard error. It suggests the lab sessions do not create the gap, and nothing stronger.

5 of 12 · Q: were the subjects alike, and does matching fix it?

The samples differ on STEM and IQ, and the matched-sample figure cannot be what its caption says (our reading)

Figure M.1 redrawn, p.75standardized difference, all vs signals0.10.250STEM majorIQ scorefemaleriskoverconfidence31% vs 21%one-to-one propensity matching82 signals82 all kept70 droppedpaper: the finding "remains unchanged".No number is printed.Figure M.3, p.77. Matched gap about 3.3 points; the band covers zero in most rounds.gap digitized by the survey agent1231 round 1 at 0.63: that guess precedes any ball, so it is a coin toss2 round 2 near 0.80: the Bayesian expectation is 0.71 (unlikely, not ruled out)3 signals ends near 0.96: the same 82 subjects end at about 0.835 in Figure 1

Appendix M compares the two samples. In the all treatment, thirty one percent of subjects study science or engineering, against twenty one percent in signals.

The fix is one to one propensity matching. The text says the finding remains unchanged, and prints no number.

Digitized from the matched figure, the gap is about three point three points, with zero inside the band in most rounds.

The survey also found values that should not occur. Round one, a coin toss, is drawn at sixty three percent. Round two sits nine points above the Bayesian expectation. Signals ends near ninety six percent, against eighty three in Figure one. The check is unusable.

6 of 12 · Q: what about subjects who ignore the balls?

About a quarter of signals and no info subjects act as if random, and the headline is never re-run without them

Appendix N, p.78: one sensitivity per subject "random" subject: signals19 of 82= 23.2%no info20 of 80= 25.0%all152 subjects, never classifieddrop the random subjectsre-run Table A.1NEVER RUNIf the random share differs betweenall and signals, part of the 5 pointsis who showed up.our reasoningFigure N.4's caption says "all"; its 82 subjects are signals.

Appendix N looks for subjects who ignore the balls. It fits one sensitivity per subject, beta i, on the Bayesian log odds. Random means the credible interval for beta i includes zero.

In signals that is nineteen of eighty two subjects. In no info, twenty of eighty. About a quarter of each.

Subjects in the all treatment are never classified, so the random share cannot be compared across the two headline treatments.

Nobody drops these subjects and re-estimates the five points. Our reasoning on why it matters: a random subject scores fifty percent, so a modest difference in their share between arms moves the gap by a point or two.

7 of 12 · Q: what could the procedure itself create?

The screen shows every draw and no running tally, so counting is costly by design (our inference)

Figure E.2, p.51: the all screen in round 2, the only screenprinted. Cell: a rectangle (bet), a circle (ball). No tally.the same table at round 20 (simulated draws) last round's 7 bets: one row, read at a glancep.23: a "low-cognitive-cost alternative to counting signals"Part of the counting noise that imitation repairs is madeby this screen. A red-minus-green counter would test that.OUR INFERENCE · the comparison stays fair only if signalssubjects saw the same grid; the paper never shows theirs

This is the screen from the instructions. Each cell holds a rectangle for the bet and a circle for the ball. There is no running count of red against green.

By round twenty, the Bayesian statistic means counting one hundred fifty two circles.

Last round's bets are one row of seven rectangles. The paper calls imitation a low cognitive cost alternative to counting.

Our inference: part of the counting noise that imitation repairs comes from this interface. An on screen counter would test that. It limits how far the result travels. The comparison stays fair if signals subjects saw the same grid, which the paper never shows.

8 of 12 · Q: what do subjects say they did?

Many say the bets add nothing, and the group that says it most uses them most

high IQlow IQ (at most 3 of 6 items)free text from all and all4 only; none from signals"what strategy did you use?"count the balls0.330.19balls plus others' bets0.370.42other rules, or none0.110.16topic shares of answer text, p.40"were the bets useful?"yes, useful0.360.29no, no extra information0.440.17the correct Bayesian answer, p.44what they do, weak signals, p.19response to others' bets0.410.33The group that most often saysbets add nothing leans on betsmost when the tally is close.OUR READING · the appendix creditsthe answers to a Bayesian subset, p.38Appendix C: beliefs about others' accuracyreported: low-IQ subjects predict round-20 accuracy worse, squared error +1.10 (SE 0.54), 476 subjectsnot reported: mean belief by treatment. Subjects in the all treatment were never asked about signals.

Subjects in the two all treatments described their strategy in free text. About forty percent of that text describes using balls with others' bets.

Asked whether bets were useful, high I Q subjects put forty four percent of their text on the Bayesian answer, no extra information. Low I Q subjects, seventeen percent.

Yet in the choices, high I Q subjects respond more to bets when the tally is close. The appendix credits those answers to a Bayesian subset. We read a mismatch, and text is never linked to choices.

On beliefs, one regression shows low I Q subjects predict accuracy worse. Subjects in the all treatment were never asked about signals.

9 of 12 · Q: how much of the gap to Bayes does imitation close?

Appendix K credits actions with one third of the value of signals, and leaves six points unexplained

258111417200.050.100.150.20game roundRMSE against the Bayesian curve (Table K.3, p.71)baselineone pooled logit for "guess is correct" then switch terms on: none, signals, signals + actionssignalssignals + actions0.075baseline0.045signals0.035+ actions the paper's "one third"rounds 3 to 5: actions add errorbaseline 0.004 at round 8:the intercepts already carrymost of the fitleft over, p.67a 6 point treatment effect that "cannot bedirectly attributed" to signals or to actionsraw accuracy shortfall from the Bayes curvesignals 0.148all 0.098SURVEY DIGITIZATION · one third closed

One logit predicts a correct guess from the Bayesian log odds and the share of others who were right. Its error against Bayes is computed with those terms off, then on.

Signals cut the average error from point oh seven five to point oh four five, and actions to point oh three five, the paper's one third.

In rounds three to five, actions worsen the fit. The baseline nearly reaches zero at round eight, so the report distrusts its scale.

The model leaves a six point treatment effect that the paper says cannot be attributed to either source. By the survey's digitization, imitation closes a third of the raw shortfall.

10 of 12 · Q: which checks does the result pass?

The direction of the effect holds across games, cutoffs and ability quantiles

G · first five games vs last fiveaccuracy gap, points5.34.8games 1 to 5games 6 to 10headline 5.1digitized from Figure G.1, p.57H · narrower "weak" bandsP(bet red) vs share of others on red00.5100.5125-7540-60endpoints read off Figure H.1, p.59;straight lines are schematicL · who gains, by ability quantileFigure L.1, p.7325th percentile: about 0.84 vs 0.70median: about 3 pointstop quarter: both near 1The direction holds under every cut. None of the three speaks to the site confound or to the size of the standard error.

Some checks come back clean. Appendix G splits the ten games in half. Measured from its figures, the gap is five point three points early, four point eight late.

Appendix H narrows the definition of a weak tally. The response to others' bets stays steep, and flattens only in the narrowest band.

Appendix L asks who gains. The gap is positive from the tenth percentile up, and largest near the twenty fifth, about eighty four percent against seventy. At the top both reach one.

So the direction is stable under each cut. All three reuse the same eleven sessions, and none speaks to site or to the standard error.

11 of 12 · Q: what is the verdict, threat by threat?

Three threats handled, two partly handled, six open: the direction is solid, the size is soft

handledlearning across gamesgap about 5 points in both halves (G)weak / strong cutoffsimitation stays under every band (H)who gainspositive from the 10th percentile up (L)partly handledlab versus onlineF: all alike by site; clean contrast unreportedcovariate imbalancematched gap about 3.3; Figure M.3 anomalousopenstatistical strength5.1 (SE 2.8), 11 clusters, 10% level onlymoves with controls5.5 (2.6) at IQ score 0, 5% level; 4.1 (2.9)plus covariates. Both carry an IQ interaction.random-acting subjectsa quarter found; headline never re-runattribution to imitation6 points unattributed in Appendix Konly 3 signals sessionsno appendix can repair thisno tally on the screenlimits external validity; oursbottom lineDirection: survives every cut in the appendices.Size: about 5 points, standard error near 3. Re-estimates run from3.3 (matched sample, digitized) to 5.5 (at IQ score zero, Table J.1).Firmer than the headline: the consensus effect and the mechanism.Ratings are the survey's, from the verdict table of the appendix report. The paper does not rate itself.

Here is the scorecard, in the survey's ratings. Handled: learning across games, the cutoffs, and who gains.

Partly handled: the site confound, where the clean contrast is omitted. And covariate imbalance, where the matched figure is anomalous.

Open: the ten percent significance level. A coefficient between four and five and a half points across the appendix regressions, where an I Q interaction shifts its meaning. The random subjects. Six unattributed points in Appendix K. Three signals sessions. And the missing tally, our own point.

The verdict is mixed. The direction survives every cut. The size is about five points, standard error near three. Consensus and the imitation mechanism are firmer than the accuracy headline.

12 of 12 · Q: what would a clean follow-up look like, and where next?

Cross treatment with site, run about eleven sessions per arm, and add a tally arm

everything on this slide is ours, not the paper'ssignalsalllabonlinem/2 sessionsone poolm/2 sessionsone poolm/2 sessionsone poolm/2 sessionsone poolcross treatment with site;randomize sessions to armsinside one subject pool481216201234sessions per armstandard error of the gap, pointstoday: 2.8needed: 1.8about 11 gap of 3.3: about 25 per arma third arm: signals + tallya red-minus-green counter on screentests the interface storyclassify random subjects in all armspre-registered rule; report the headlinewith and without themreport the intervalin the abstract, next to the 5 pointsread nextTable A.1, p.34 · Appendix F, p.53 to 56 · Figure E.2, p.51 · Table K.3, p.71 · Appendix N, p.77 to 80

Everything here is our proposal. First, cross treatment with site. Both treatments run in each mode, from one subject pool, sessions randomized to arms.

Second, sample size. The reported standard error implies session means varying by about four points. Eighty percent power against five points needs a standard error near one point eight, so about eleven sessions per arm, and about twenty five if the true gap is nearer three.

Third, add an arm with signals plus a running tally. If the gap closes, the interface was doing the work. Classify random subjects in every arm by a preset rule.

Read next: Table A one, Appendix F, the Appendix E screenshot, then Table K three.

1.00×
keys
space play / pause
slide · , . beat
[ ] speed · c captions · d deeper
m mute · f fullscreen · r replay slide