Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

'The fundamental problem is that p values don't mean what we "need" them to mean, that is p(null | significant effect).'

From Bayes' theorem, this more useful probability is given by p * x, where x = p(null) / p(significant effect). Maybe we could just lower the accepted threshold for statistical significance by several orders of magnitude so that, for statistically significant p, p * x is still small even for careful (i.e. big) estimates of x (e.g. maybe a Fermi approximation of the total number of experiments ever performed in the field in question). This doesn't necessarily imply impractically big sample sizes, although obviously this depends on the specifics (I believe the p value for a given value of the t-statistic decays exponentially with sample size).



I don't follow your argument. You've got two premises:

1) You are saying that people are committing the transposing the conditional fallacy: p(H0|data) != p(data|H0):

- OK

2) You say to use Bayes theorem to get the value we want:

- OK, but actually a better formulation is

  p(H_0|data) = p(H_0)*p(data|H_0)/[p(H_0)*p(data|H_0) + p(H_1)*p(data|H_1) + ... + p(H_n)*p(data|H_n)]
You probably don't need to add up all the way to hypothesis n since the terms eventually become negligible and can be dropped from the denominator. The point is that you have to compare how likely the result would be under other hypotheses, not just H_0.

3) You propose lowering the threshold for "significance"

- How does this follow from the premises? Lets say you get a very low value for p(H_0)p(data|H_0), this can still be much higher than p(H_1)p(data|H_1), etc so it is still the best choice. Ie, you can get a low p-value given H_0 but if there is no better model out there you should still keep H_0.


I’m assuming the question we are trying to answer is not “which H_n is most probable”, but rather “how safe is it to conclude that H_0 is not true”. For example, say we are concerned with whether the difference between two groups is less than a certain amount or greater than a certain amount.


>“how safe is it to conclude that H_0 is not true”

I would just take "very safe" as a principle, there is even the truism "all models are wrong".

>"the difference between two groups is less than a certain amount or greater than a certain amount"

You are ignoring a lot of the model being tested here (eg, normality, independence of the measurements, etc) and only considering one parameter.


“how safe is it to conclude that H_0 is not true”

What I meant here by H_0 was the hypothesis that the difference between groups is less than some particular threshold. I think if you made the threshold large enough then it would not be safe to conclude that H_0 is not true.

"You are ignoring a lot of the model being tested here (eg, normality, independence of the measurements, etc) and only considering one parameter. "

I said enough so that if you were arguing in good faith you could fill in the gaps yourself.


I am arguing in good faith. There may be some cases in physics where there is some theoretical distribution derived related specifically to the problem, and they believe the model to be actually 100% true. Otherwise, the model should be assumed to only be an approximation at best.


Yes, the t-test does assume normality and you can never be sure of perfect normality if that's what you are getting at (although I believe that simulation tests of the robustness of the t-test against deviations from normality generally show that this isn't too much of a practical concern). I wasn't trying to address every potential weakness with the t-test (or p-values in general); I was addressing the one stated in the article.


I'm saying that unless you have bothered to derive a statistical model from your theory, and you believe that theory may actually be correct, then you know that you will reject your model if enough time/money is spent on testing it.


Ok, I will assume that by "model" here you mean a probability distribution on the parameters relevant to your experiment. In that case I agree with what you just said: knowing exactly the correct model is impossible in a similar way that knowing someone's height to an infinite degree of precision is impossible. But I never said anything to the contrary. The H_0 I gave corresponds to an infinite set of models and not a single one (note that I said the difference is less than a certain threshold, not that the difference is 0 (although it still would be an infinite set in that case, but the probability would be 0)).


And in anticipation of the rebuttal that the probability is still 0 because it's never exactly normal, what I really meant originally but didn't write out explicitly for brevity and because I assumed it would be implicit: when I talk about giving an upper bound on p(null | data) from p(data | null), what I really mean is giving an upper bound on p(null | data, normal) from p(data | null, normal) where normal is the assumption that the distribution of whatever parameter we are looking at is normally distributed and null is the event that the difference in means between the two groups we are looking at it is less than some predetermined positive threshold. Or, for a 1-sample test, that the mean of a single group is within that threshold of some default value.


If you write out the actual calculation you will see normality (which was just one example of an assumption) is actually part of the null model being tested. It is not something different or outside of it.


That is just a trivial semantics issue, and yes I am familiar with the calculation.

Where do you think the flaw is specifically? Say we are doing a 1-sample test.

0. (setup) Suppose we have a real number mu and a positive epsilon. Define the interval I as [mu – epsilon, mu + epsilon]. For each “candidate mean” within this interval, we have a corresponding t-statistic. Let the statistic t_0 be the inf of all these t-statistics. Let T be the event that t_0 is at least as big as the observed value.

1. You can use Student’s t-distribution to compute an upper bound for the probability of T under the assumptions that the observations are iid normal and the mean lies in I. I will call this probability p(T | null, normal, iid), where “null” is the event that the mean exists and is in I. It makes no difference that it is more typical to lump these assumptions together as “null” because in math you can define things however you want as long as you are consistent.

2. We have that p(null | T, normal, iid) = p(T | null, normal, iid) * p(null | normal, iid) / p(T | normal, iid).

3. Therefore, if we have an upper bound for x = p(null | normal, iid) / p(T | normal, iid) then we can get an upper bound for p(null | T, normal, iid). That is my main claim.

Which of the above statements do you object to?


>"You can use Student’s t-distribution to compute an upper bound for the probability of T under the assumptions"

I'm not sure what you are arguing anymore. I am saying you will never test a parameter value in isolation, it is always part of a model with other assumptions. There is simply no such thing as testing a parameter value alone. To define a likelihood you need more than simply a parameter...

You seemed to be disagreeing with that, but are now acknowledging the presence of the other assumptions.


“I'm not sure what you are arguing anymore.”

It’s the claim I make in 3, and then the secondary claim that making our upper bound on p(null | T, normal, iid) small for significant p-values (i.e. p(T | null, normal, id)) could be used as a criterion for whether our threshold for statistical significance is small enough.

“You seemed to be disagreeing with that”

I’m not sure what I said that gave that impression. I didn’t mention anything about the normal / iid assumptions initially not because I thought we weren’t making these assumptions but because I didn’t think these details were essential to my point.


"Let the statistic t_0 be the inf of all these t-statistics. Let T be the event that t_0 is at least as big as the observed value." Oops, I meant the inf of their absolute values, and T is the event that t_0 is at least as big in absolute value as the observed value.


Also, every probability mentioned should also include in the list of conditions that the number of samples observed matches our experiment.


>"The H_0 I gave corresponds to an infinite set of models and not a single one"

How do you calculate a p-value based on this infinite set of models? Normally it is done using just one.


Ok for simplicity let’s assume this is a 1-sample test. So there is a certain “default” mean, say mu, and we are concerned with whether the mean of some random variable on whatever population we are sampling from is within, say, epsilon of this default value. For every number in [mu – epsilon, mu + epsilon] we can get a p-value giving the probability that we would have observed the data if this was the true mean. In order to get a probability that we would have observed the data given that the mean was somewhere in this interval, we need some prior distribution on the means in this interval, which we don’t have. However, we can just take the sup of all the p-values for each mean in this range to get an upper bound. (I think this is similar to how confidence intervals work also but take that with a grain of salt)


Looking closer at this you are describing one model but testing different values of one of the model parameters.


“Looking closer at this you are describing one model but testing different values of one of the model parameters.”

I am inferring from context that this is probably supposed to be a criticism but it doesn’t make much sense to me. Of course we have to consider different values of the mean, the whole point is to get a p-value corresponding to a range of potential different means.

But anyways, I do think that an explanation based on the t_0 statistic I defined in my other post is better.

1. We can define a statistic t_0 that is the infimum of the absolute value of all t-statistics for every candidate mean in the interval.

2. Suppose the mean is some value mu’ in the interval. Whenever the t_0 statistic is at least as big as its observed value, the t-statistic corresponding to mu’ is also at least as big in absolute value as the observed value of t_0, by definition.

3. So we can give an upper bound for the probability of the former event by the probability of the latter event. But the probility of the latter event is the same no matter what mu’ is, and can be computed using Student’s T distribution.

4. Therefore, we have an upper bound for the probability of t_0 attaining a value at least as big as the observed value, assuming that the mean is somewhere in the specified interval (plus the other standard assumptions). This is the p-value.

Do you agree with those assertions? If not, specifically where is the problem?


“Looking closer at this you are describing one model but testing different values of one of the model parameters.”

Furthermore, this is even done with the usual 1-sample t-test because the variance can be anything.


By the way, it's not normally done using one either. With a 1-sample t test for example, the null hypothesis is that the underlying distribution has a certain prescribed mean, but the variance can be anything.


Ok, you seem to accept that there is an assumption that the data is generated by a distribution with a mean, so start with that. This is not necessarily true: https://en.wikipedia.org/wiki/Cauchy_distribution

I use that only as an example. If you look closer you will find many other assumptions being made as well that are used to derive the actual calculation (for whatever statistical test you choose to look at).


I've already addressed that. See "in anticipation of the rebuttal" post.


Minor point, but you are missing p(H_0)*p(data|H_0) in the denominator in bullet point 2.


Thanks, fixed.


This was basically the suggestion here:

https://www.nature.com/articles/s41562-017-0189-z.epdf

Also previous HN discussion:

https://news.ycombinator.com/item?id=15192610


Responding to this but getting rid of the intense nesting:

  “I'm not sure what you are arguing anymore.”
  
  It’s the claim I make in 3, and then the secondary claim that making our upper bound on p(null | T, normal, iid) small for significant p-values (i.e. p(T | null, normal, id)) could be used as a   criterion for whether our threshold for statistical significance is small enough.

  “You seemed to be disagreeing with that”

  I’m not sure what I said that gave that impression. I didn’t mention anything about the normal /   iid assumptions initially not because I thought we weren’t making these assumptions but because I   didn’t think these details were essential to my point."

Please give some example code or calculation steps for what you are talking about.


You mean for computing the p-value associated with a range of means rather than a single null mean value? That’s the only thing I can think of for which providing code or calculation steps would be applicable (well, there’s also getting a Fermi approximation for the total number of experiments ever performed in a discipline, which I gave as a preliminary suggestion for a cautious estimate of x, but I don’t have time to do that and the main difficulty of that isn’t the math anyways). Anyways I would be happy to provide that if you can confirm that this is what you meant and that you are asking out of genuine curiosity.


>"our upper bound on p(null | T, normal, iid) small for significant p-values (i.e. p(T | null, normal, id))"

This sounds like nonsense to me so I would like to see an example of what you mean.


Ooops, I forgot to include something in the list of conditions. For every probability I mentioned in the post you just quoted and its grandparent post, we should add to the list of conditions that the number of samples used is the same as the number of samples used in our experiment.

Anyways, if it sounds like nonsense then I will try rephrasing. Right now the threshold for statistical significance is p = 0.05, sometimes lower depending on the field. Let’s say we want to lower this threshold, and we need some way of determining how low it should be. I am suggesting that this could be done by deciding on an acceptably safe threshold, say q, for p(null | T, num samples, normal, iid), and an acceptably safe upper bound, say x_0, for p(null | num samples, normal, iid) / p(T | num samples, normal, iid). Then we use q / x_0 as the threshold for statistical significance.

We would also be using p-values for a null hypothesis corresponding to a range of means rather than a single value (or a range of differences in means, or something else suitably adapted to the type of test we are doing), because with a single value, having an upper bound on the probability of the null hypothesis is vacuous, as you very eagerly point out.

You could argue that this is useless because what we really want to bound is p(null | num samples, T), not p(null | num samples, T, normal, iid). I would respond that, yes this would be the more desirable quantity to know, but we need to make some simplifying assumptions in order to make this problem tractable, and while the normal and iid assumptions probably aren’t true, there are at least situations where we can be confident that they approximate reality reasonably well.


Also, as an aside, I’m not sure what you mean specifically when you say “model”. The usual definition is a set of probability distributions on a common sample space. But when you say things like “all models are wrong”, I’m assuming you are just using “model” as a synonym for “distribution”, because this claim is vacuously false using the usual definition of model (just take the set of distributions to be every distribution on that sample space; this must contain the true distribution by definition). But then when you say “Looking closer at this you are describing one model but testing different values of one of the model parameters”, this seems to make more sense if you are using the usual definition of model. So I’m not sure what definition you are using.


Also, in case this clears up some confusion, I was using the nonstandard definition in some posts because I inferred that this was the definition you were using.


Here is an example of a model.

The other day I got a robocall from one of those spoofers that uses the same area code + three digits as the number being called. This call happened to be a number I did know. Did they just happen to hit upon this number, or is something more sinister going on (eg using a hacked address books)?

Say I've gotten nCalls calls like this so far, whats the probability at least one of them would be from a number I know?

The probability of any given number being used will be 1/9999, if I know 5 numbers with the same first digits as my own it would be 5/9999, etc. This leaves 1 - nKnown/9999 other possible numbers to be used. The probability they keep using unknown numbers will then just be eg, (9998/9999)^nCalls and we take 1 minus this value to get the probability that at least one of the calls will be from a known number. Here is the model:

  model_1 = 1 - (1 - nKnown/9999)^nCalls
Lets say I've gotten about one call like this every day for the last two years. So nCalls = 730, and I know 3 numbers that share the same digits including my own (I am assuming they also spoof the number they are calling). Then the probability of getting at least one call from the known numbers would be 20%. I made simulation in R and see the same results:

  # Returns percent of calls coming from a known number for nSim experiments
  sim <- function(nSim, nCalls, nNum, nKnown, replace = FALSE){
    res = replicate(nSim, sample(1:nNum, nCalls, replace = replace) %in% 1:nKnown)
    return(colMeans(res))
  }

  model_1 = mean(sim(1e4, 730, 9999, 3, T) > 0)
I would guess that robocallers are pretty cheap so are probably going with the simplest approach possible. But perhaps they are slightly more advanced and they avoid reusing the same number (ie, if it went unanswered and I never answer unknown numbers) and avoid using the callees number. This can be done by changing a few arguments in the sim:

  model_2 = mean(sim(1e4, 730, 9998, 3, F) > 0)
The two models predict nearly the same thing in the range of 730 calls, so p(data|model_1) = p(data|model_2) = 20%. I think the simpler model is still more probable though, so lets say p(model_1) = 0.75 and p(model_2) = 0.24 and p(model_x) = .01. Model x is that something more shady like using the hacked contact info is going on, this is pretty vague and can explain anything so give p(data|model_x) = 1.

Then we use Bayes' rule:

  p(model_1|data) = .75*.2/(.75*.2 + .24*.2 + .01*1) = 72%
  p(model_2|data) = .24*.2/(.75*.2 + .24*.2 + .01*1) = 23%
  p(model_x|data) = .01*1 /(.75*.2 + .24*.2 + .01*1) = 5%

EDIT:

Actually, I made an error for model 2. Since I am not including my own number, nKnown should be 2 instead of 3.

    model_2 = mean(sim(1e4, 730, 9998, 2, F) > 0)
This gives p(data|model_2) = .14, and:

  p(model_1|data) = .75*.2 /(.75*.2 + .24*.14 + .01*1) = 77%
  p(model_2|data) = .24*.14/(.75*.2 + .24*.14 + .01*1) = 17%
  p(model_x|data) = .01*1  /(.75*.2 + .24*.14 + .01*1) = 5%




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: