Simpson's Paradox: Why Aggregate Metrics Can Mislead You | Ushasvi

Using a hypothetical marketing example, Ushasvi Rachel shows how Simpson's Paradox can make one campaign look better overall while the other wins in every customer subgroup.

4:29 video3 min readWatch on YouTube

An aggregate KPI can say campaign A is the clear winner while campaign B actually performs better within every customer segment that matters. Both numbers can be calculated correctly at the same time. Using a hypothetical marketing example, Ushasvi Rachel walks through how this happens and why it is not a calculation error.

The reversal, in numbers

Across all customers in the example, campaign A converts at 68 percent and campaign B converts at 42 percent. Looking only at that aggregate, campaign A looks like the obvious winner. But that comparison hides how differently the two campaigns were distributed. Campaign A was shown mostly to high-intent customers, and campaign B was shown mostly to low-intent customers, two groups that start with very different propensities to convert before either campaign even runs.

What the subgroups actually show

Split the results by customer intent and the picture flips. Among high-intent customers, campaign A converts at 80 percent and campaign B converts at 90 percent, so B wins that group. Among low-intent customers, campaign A converts at 20 percent and campaign B converts at 30 percent, so B wins that group too. Campaign B is stronger in both meaningful segments, even though campaign A leads once everyone is aggregated together. This is Simpson's Paradox: a pattern observed in aggregated data changes, weakens, or fully reverses once the data is split into the groups that actually matter. Nothing here is a math error. The aggregate rate and the subgroup rates are each calculated correctly; they just answer different questions.

Why the population mix causes it

The reversal happens because of audience composition. Campaign A received a much larger share of high-intent customers, who convert more easily under either campaign. Campaign B received mostly low-intent customers, who are harder to convert regardless of which campaign they see. That mix pulled campaign A's overall average up and campaign B's overall average down, independent of how good either campaign actually was. Customer intent matters here because it is tied to both the composition of each campaign's audience and to the outcome being measured. That can create a composition effect and hint at confounding, but Simpson's Paradox by itself does not prove a specific confounder or establish causation.

Two correct answers to two different questions

The aggregate comparison weights each segment by the campaign's actual audience mix, which is the real-world exposure. The subgroup comparison asks how the campaigns differ among customers with similar intent, holding that variable constant. Both are legitimate statistical views. Which one should drive a decision depends on which question the decision actually needs answered. This is not limited to marketing; the same composition problem can distort healthcare comparisons, hiring metrics, product experiments, and regional sales figures whenever groups differ in baseline outcomes and are unevenly distributed across whatever is being compared.

Five questions before trusting an aggregate KPI

The video closes with a checklist to run before acting on any single aggregate number: which populations were combined, do the groups have different baseline outcomes, were the options exposed to different group proportions, does the pattern persist within meaningful subgroups, and which comparison actually answers the decision at hand. Simpson's Paradox does not mean aggregate metrics are always wrong, and subgroup analysis does not automatically hand you a causal answer either. Selection, measurement, sample size, and decision context still need scrutiny either way.

Key takeaways

  • Simpson's Paradox happens when a trend in aggregated data reverses, weakens, or disappears once the data is split into meaningful subgroups, without any calculation error.
  • In the example, campaign A wins the aggregate 68 percent to 42 percent, but campaign B wins both the high-intent subgroup (90 to 80) and the low-intent subgroup (30 to 20).
  • The cause is uneven audience composition: campaigns that reach different mixes of easy-to-convert and hard-to-convert customers will show different aggregate averages regardless of true performance.
  • Aggregate and subgroup comparisons answer different questions, and picking the right one depends on the decision being made, not on which number looks better.
  • Before trusting a KPI, check what populations were combined, whether their baseline outcomes differ, and whether the pattern holds up once the data is broken into meaningful groups.

Who this is for

This is for anyone who reads marketing, product, healthcare, or hiring metrics and needs a way to check whether an aggregate number is hiding a reversed pattern underneath it.

Chapters

  1. 0:00How campaign A beats campaign B overall but loses in every subgroup
  2. 0:25Separate subgroup performance: High-intent vs low-intent customers
  3. 0:50Defining Simpson's Paradox and trend reversals
  4. 1:15How audience composition pulls overall averages up or down
  5. 1:40Weighting actual audience mix vs comparing similar baseline groups
  6. 2:05Five critical questions to ask before acting on an aggregate KPI
Full transcript(auto-generated, with timestamps)

How campaign A beats campaign B overall but loses in every subgroup

[0:00]What if the overall data says campaign A performs better, but within every important customer group, campaign B actually performs better? That sounds impossible, but both conclusions can come from correctly calculated data. The difference is hidden in how the populations are combined. Consider a hypothetical marketing example. Across all customers, campaign A converts 68% while campaign B converts 42%.

Separate subgroup performance: High-intent vs low-intent customers

[0:28]Looking only at this aggregate KPI, campaign appears to be the obvious winner. But that comparison leaves out an important part of the story. The campaigns reached very different customer populations. Campaign was shown mostly to high intent customers while campaign B was shown mostly to low intent customers. Those groups begin with different propensities to convert

Defining Simpson's Paradox and trend reversals

[0:51]Before we even compare campaign performance. Now separate the results among high intent customers. Campaign A converts 80% and campaign B converts 90%. Within this group, campaign B performs better. These are illustrative values for this hypothetical example, not results from a research study. The low intent group tells the same story.

How audience composition pulls overall averages up or down

[1:15]Campaign A converts 20% while campaign B converts 30%. Campaign B also performs better here. So B leads within both meaningful customer segments even though a leads when every customer is aggregated. The apparent conclusion has reversed overall a look stronger within high intent customers. B is stronger within low intent customers. B is

Weighting actual audience mix vs comparing similar baseline groups

[1:40]Stronger. Again the arithmetic is not broken. The aggregate rates and subgroup rates are each calculated correctly. This is Simpsons paradox. A pattern where a trend observed in aggregated data changes, disappears or reverses after the data is divided into meaningful groups. A complete reversal is the most striking version. But a trend can also weaken or vanish. Why does it happen here? Campaign A received

Five critical questions to ask before acting on an aggregate KPI

[2:06]A much larger proportion of highintent customers who were easier to convert under either campaign. Campaign B received mostly low intent customers who were harder to convert. The population mix pulled A's overall average upward and B's downward. Customer intent matters because it is related to the composition of each campaign's audience and to the outcome being measured. Conversion. This can create a composition effect and may signal confounding. But Simpsons paradox alone does not prove a specific confounder or establish causality. Both statistical views can be mathematically correct because they answer different questions. The aggregate comparison weights the segments according to each campaign's actual audience mix. The subgroup comparison asks how the campaigns differ among customers with similar intent. Interpretation depends on which question supports the decision. This matters far beyond marketing. Population composition can complicate healthcare comparisons, hiring metrics, product experiments, regional sales performance, and customer behavior analysis. Whenever groups differ in

Their baseline outcomes and are unevenly distributed, an aggregate KPI can hide important structure. Simpsons paradox is not necessarily a calculation mistake or deliberate data manipulation. It does not mean aggregate metrics are always bad and subgroup analysis does not automatically reveal causal truth. Selection, measurement, sample size, and the decision context still require careful examination. Before acting on an aggregate KPI, ask five questions. Which populations were combined? Do the groups have different baseline outcomes? Were the options exposed to different group proportions? Does the pattern persist within meaningful groups? And which comparison actually answers the decision you need to make? Aggregate statistics can be correct and still mislead interpretation when population composition is hidden. Subgroup statistics can reveal the hidden pattern, but they still require thoughtful analysis rather than automatic causal conclusions. Always ask what happens when the data is broken into the groups that actually matter. I'm Ashazv Rachel with humanitarians doi.

More from Humanitarians AI Fellows

Humanitarians AI Lyrical Literacy Project