Exposing Selection Bias and Analytic Choices in Madison

Sai walks through a synthetic dataset with a known $1,400 true effect and shows how twelve legitimate analytic choices can make it look like an $8,000 gain or a $5,600 loss.

2:07 video3 min readWatch on YouTube

Sai, a Humanitarians AI fellow, walks through a post built entirely around a synthetic dataset where the true answer is already known. The point isn't to catch someone faking numbers. It's to show the far more common and much harder to catch version of data manipulation: choosing, after the fact, which correct analysis to publish.

Two press releases, one dataset, opposite conclusions

Sai generated the data himself, so the true effect of the program he's modeling is exactly $1,400 a year, the same for every participant. Every other number that comes out of the same 12,000 rows is also real, computed correctly using standard analytical methods. Applying twelve common analytical choices to that one dataset produces results ranging from about $1,100 to nearly $8,000, and two of those results point in opposite directions. A program-office framing claims graduates outearn everyone else by $7,988. An opposition framing claims participants fall $5,647 behind. Both numbers are computed correctly. Both are publishable. Neither is anywhere near the true $1,400.

The choice hiding underneath the effect

The reason both flawed results are possible comes down to one thing: people don't enroll in a program at random. In Sai's model, the program recruits hardest where need is greatest, meaning participants start out poorer than non-participants before the program does anything at all. Ignore that fact and the estimated effect can land anywhere from roughly negative $6,000 to positive $8,000. Account for it, using standard adjustment methods, and nearly every version of the analysis lands close to the true $1,500 figure. Sai is explicit that none of this requires a bad actor: it's simply what happens when someone with a preferred conclusion gets to make analytic choices after already seeing the data. Which rows to include, which time window, which denominator, which average, where the axis starts: each individual choice has a justification someone could say out loud in a meeting, even when the combination of choices is misleading.

The fix is about twenty lines of code

Sai's proposed antidote is straightforward: run every defensible version of the analysis and plot all of them together rather than picking one. When you do that with this dataset, the versions that ignore selection scatter across both sides of zero, while the versions that account for it cluster tightly within a few hundred dollars of the true value. Once every analysis is visible at once, quoting only the most flattering one requires an explanation that wasn't needed before.

Three questions for reading any chart

Sai closes with a practical checklist for anyone reading a data-backed claim, not just a nonprofit's own reporting: who is in each group and how did they get there, what happened to everyone who dropped out, and how many analyses were run before the one being shown. His underlying argument is that a number, on its own, is a claim about a procedure, what was measured, who was included, what was tried first, and that the procedure is usually the part that isn't shown alongside the result.

Key takeaways

  • A dataset with a known true effect can produce a wide range of "correct" results depending on which analytic choices are made.
  • Selection bias, when participants aren't random, can flip an estimate's direction even when every calculation is technically valid.
  • Running and plotting every defensible analysis, rather than one, exposes which results are outliers driven by choice rather than signal.
  • Ask who is in each group, what happened to dropouts, and how many analyses preceded the one shown.
  • A published number is a claim about a procedure, not just an outcome; the procedure is what deserves scrutiny.

Who this is for

This video is for anyone who reads or produces data-driven claims, including nonprofit researchers, program evaluators, and Humanitarians AI fellows working with real-world datasets where selection effects are common.

Chapters

  1. 0:00The Common Manipulation: Post-Data Analytic Choices
  2. 0:28The Synthetic Test: One Dataset, Two Opposite Conclusions
  3. 0:55Underneath the Effect: Why Selection Bias Distorts the Numbers
  4. 1:20The 20-Line Antidote: Plotting Every Analysis for Full Transparency
  5. 1:45Three Questions to Ask Every Chart You Read This Week
Full transcript(auto-generated, with timestamps)

The Common Manipulation: Post-Data Analytic Choices

[0:00]Most people picture data manipulation as making numbers up. That's the rare version and it's the one that gets caught. This is sigh. I wrote a post that does the common version instead. One synthetic data set where I know the true answer. 12 standard analytical tricks and two press releases that reach opposite conclusions. Every number in both of them is real. I wrote the data generating process myself. So the true effect is exactly $1,400 a year. The same for everyone. Every other number on this axis is also real. computed from

The Synthetic Test: One Dataset, Two Opposite Conclusions

[0:28]The same 12,000 rows. They run from about $1,100 to nearly $8,000 and two of them point in opposite directions. Same 12,000 rows. The program office says graduates outturn everyone else by $7,988. The opposition says participants fall 5,647 behind. Both are computed correctly. Both are publishable. The truth is $1,400 and neither side is anywhere near it. One choice sits underneath all of

Underneath the Effect: Why Selection Bias Distorts the Numbers

[0:55]It. People don't enroll at random. The program recruits hardest where need is greatest. So, participants start out poorer before it does anything at all. Ignore that and your estimate lands anywhere from minus 6,000 to plus 8,000. Account for it and every version lands near 1500. None of this needs a bad person. It's what happens when someone with a preferred conclusion gets to make analytic choices after seeing the data. Which rows, which window, which

The 20-Line Antidote: Plotting Every Analysis for Full Transparency

[1:20]Denominator, which average, where the axis starts? Each one has a justification you could say out loud in a meeting. The antidote is about 20 lines of code. Run every defensible analysis and plot all of them. The ones that ignore selection scatter across both sides of zero. The ones that account for it cluster within a few hundred dollars of the truth. Then quoting only the flattering one needs an explanation. Here's yours. Take one chart you actually read this week and

Three Questions to Ask Every Chart You Read This Week

[1:46]Ask it three things. Who is in each group and how did they get there? What happened to everyone who dropped out? How many analyses were run before this one? What you can't answer from the page is the finding. A number is a claim about a procedure. What was measured? Who was included? What was tried first? Numbers don't lie. The procedure does. And it's usually the part that isn't shown. This is sigh.

More from Madison

Humanitarians AI Lyrical Literacy Project