No, you aren't the only one. It has become even worse now that Twitter will not ...

anonymousiam · on March 24, 2021

Okay, I googled it. A non-hostile site hosts a definition here: https://en.wikipedia.org/wiki/Berkson%27s_paradox

junippor · on March 24, 2021

Thank you. Does anyone understand the difference between this and Simpson's paradox?

kgwgk · on March 24, 2021

The latter appears when analyzing subgroups gives a different result than analyzing the pooled data.

The former is about correlations that appear in samples which are not representative of the general population, due to the way that those samples are selected.

junippor · on March 24, 2021

> The latter appears when analyzing subgroups gives a different result than analyzing the pooled data.

> The former is about correlations that appear in samples which are not representative of the general population, due to the way that those samples are selected.

You just said the same thing twice. Think about it.

For one you used terms like "subgroups" and "pooled data" and for the other "samples" and "general population". Those are the same things.

Then you used "[the effect] appears in" and in the other "correlations". Well, Simpsons paradox can also manifest itself in correlations. So you just said the same thing twice.

FeepingCreature · on March 25, 2021

Simpson's paradox: analyzing trends per subgroup can give a different result than pooled data.

Berkson's paradox: analyzing a single subgroup selected with a function aggregating two traits (additively?) will indicate an anticorrelation between the traits.

Simpson's paradox says you can't judge group trends from subgroup trends. Berkson's paradox says given a group selected in a specific way, it will have a certain property in itself. They're just different statements.

Pyramus · on March 25, 2021

Yes and no.

Berkson's paradox is a special case of Simpson's for the two subgroups selected and non-selected.

The difference is that Berkson's paradox involves selecting the subgroup a posteriori and in a particular way, Simpson's paradox assumes a selection a priori.

kgwgk · on March 25, 2021

Another difference is that Simpson's "paradox" involves all the subgroups that the full population is partitioned into, unlike Berkson's "paradox".

junippor · on March 25, 2021

I like how you put paradox in quotes. I also annoys me when people call these things paradoxes. They're more properly called counter-intuitive phenomena. I wonder if there's a single-word name for that.

kgwgk · on March 25, 2021

paradox :-)

junippor · on March 25, 2021

So I actually checked and... turns out you're right.

According to Wikipedia, "paradox" can either mean "logically self-contradictory statement" or a "statement that runs contrary to one's expectation". I always thought that it meant the former only.

These two concepts should really really have separate words.

getlawgdon · on March 25, 2021

You "checked Wikipedia," is that it? You're done now?

junippor · on March 25, 2021

You sound like you're trying to make a point. Make a point.

doubleunplussed · on March 24, 2021

Eh. There is intentional splitting into subgroups, and there is accidental selection bias. I think that's the difference.

caddemon · on March 25, 2021

I think Berkson's paradox is more specific than just correlations arising from non-random sample selection. Correlations that are not representative of the general population could still be useful, if it's a meaningful correlation within some subgroup of interest. The problem is when the features you are correlating relate too closely to the features that were used for sample selection - then you can end up with a trivial result.

I've always learned of Simpson's paradox as relating more to different sample sizes when partitioning data, which can happen entirely arbitrarily - for example a baseball player getting injured part way through the season.

The fact that one player's at bats get partitioned differently than another's is not caused by the on field performance, so there's no "double dipping" going on like I would imagine with Berkson's. Conversely I'm having trouble fitting a Berkson's example into the framework of Simpson's paradox, since there's no reason the poorly-selected subpopulation can't theoretically be exactly half of the general population. And if all of the samples are of equal size Simpson's paradox doesn't exist anymore (because with equal bin sizes the mean of means is equivalent to the overall mean).

kgwgk · on March 25, 2021

> I've always learned of Simpson's paradox as relating more to different sample sizes when partitioning data

When you look at proportions based on binary outcomes it may be related to imbalanced groups but it's more general than that.

In the context discussed here of correlations between continous variables the groups can be of similar size.

See for example the chart here: https://towardsdatascience.com/simpsons-paradox-d2f4d8f08d42

caddemon · on March 25, 2021

Interesting, I only ever heard of Simpson's paradox in the context of comparing overall averages versus subgroup averages.

I guess this paradox could then be thought of as a special case of Simpson's paradox? Since the out group will exclude people with both traits there should also be a negative correlation there, which disappears in the overall population. But in Berkson's case it seems they're implying the subgroup correlation is spurious whereas with Simpson's it could go either way.

kgwgk · on March 25, 2021

> Since the out group will exclude people with both traits there should also be a negative correlation there

Not necessarily. Imagine the traits are distributed uniformly and independently in [-1 1]. There is no correlation:

    ******
    ******
    ******
    ******
    ******
    ******

If you select people with at least one positive trait you will find negative correlation in the group + but the correlation will still be zero in the group -.

    ++++++
    ++++++
    ++++++
    ---+++
    ---+++
    ---+++

caddemon · on March 25, 2021

Makes sense, I was picturing more of a diagonal boundary but you're right the paradox doesn't specify the shape of the boundary. Thanks!

kgwgk · on March 25, 2021

I don't think so.

In the first one you have a partition in subgroups A and B (or more than two) which show similar correlations, different from the correlation seen in A+B.

In the second one you have only a subgroup A (the implicit complement notA is not observed) where the correlation is not the same as in the (unobserved) full population A+notA. Nothing is said about the correlation in notA. It could be at either side of the correlation in the full population, while in Simpson’s paradox both subgroups are in the same side.

Edit: and I also mention "due to the way that those samples are selected" for Berkson's paradox where the selection is based on the variables of interest while in Simpson's paradox the subgroups are "external" (but influence the correlation between those variables).

1vuio0pswjnm7 · on March 29, 2021

You do not need to enable Javascript, you only need to change your User-Agent header to one that is acceptable.