Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Also, as an aside, I’m not sure what you mean specifically when you say “model”. The usual definition is a set of probability distributions on a common sample space. But when you say things like “all models are wrong”, I’m assuming you are just using “model” as a synonym for “distribution”, because this claim is vacuously false using the usual definition of model (just take the set of distributions to be every distribution on that sample space; this must contain the true distribution by definition). But then when you say “Looking closer at this you are describing one model but testing different values of one of the model parameters”, this seems to make more sense if you are using the usual definition of model. So I’m not sure what definition you are using.


Also, in case this clears up some confusion, I was using the nonstandard definition in some posts because I inferred that this was the definition you were using.


Here is an example of a model.

The other day I got a robocall from one of those spoofers that uses the same area code + three digits as the number being called. This call happened to be a number I did know. Did they just happen to hit upon this number, or is something more sinister going on (eg using a hacked address books)?

Say I've gotten nCalls calls like this so far, whats the probability at least one of them would be from a number I know?

The probability of any given number being used will be 1/9999, if I know 5 numbers with the same first digits as my own it would be 5/9999, etc. This leaves 1 - nKnown/9999 other possible numbers to be used. The probability they keep using unknown numbers will then just be eg, (9998/9999)^nCalls and we take 1 minus this value to get the probability that at least one of the calls will be from a known number. Here is the model:

  model_1 = 1 - (1 - nKnown/9999)^nCalls
Lets say I've gotten about one call like this every day for the last two years. So nCalls = 730, and I know 3 numbers that share the same digits including my own (I am assuming they also spoof the number they are calling). Then the probability of getting at least one call from the known numbers would be 20%. I made simulation in R and see the same results:

  # Returns percent of calls coming from a known number for nSim experiments
  sim <- function(nSim, nCalls, nNum, nKnown, replace = FALSE){
    res = replicate(nSim, sample(1:nNum, nCalls, replace = replace) %in% 1:nKnown)
    return(colMeans(res))
  }

  model_1 = mean(sim(1e4, 730, 9999, 3, T) > 0)
I would guess that robocallers are pretty cheap so are probably going with the simplest approach possible. But perhaps they are slightly more advanced and they avoid reusing the same number (ie, if it went unanswered and I never answer unknown numbers) and avoid using the callees number. This can be done by changing a few arguments in the sim:

  model_2 = mean(sim(1e4, 730, 9998, 3, F) > 0)
The two models predict nearly the same thing in the range of 730 calls, so p(data|model_1) = p(data|model_2) = 20%. I think the simpler model is still more probable though, so lets say p(model_1) = 0.75 and p(model_2) = 0.24 and p(model_x) = .01. Model x is that something more shady like using the hacked contact info is going on, this is pretty vague and can explain anything so give p(data|model_x) = 1.

Then we use Bayes' rule:

  p(model_1|data) = .75*.2/(.75*.2 + .24*.2 + .01*1) = 72%
  p(model_2|data) = .24*.2/(.75*.2 + .24*.2 + .01*1) = 23%
  p(model_x|data) = .01*1 /(.75*.2 + .24*.2 + .01*1) = 5%

EDIT:

Actually, I made an error for model 2. Since I am not including my own number, nKnown should be 2 instead of 3.

    model_2 = mean(sim(1e4, 730, 9998, 2, F) > 0)
This gives p(data|model_2) = .14, and:

  p(model_1|data) = .75*.2 /(.75*.2 + .24*.14 + .01*1) = 77%
  p(model_2|data) = .24*.14/(.75*.2 + .24*.14 + .01*1) = 17%
  p(model_x|data) = .01*1  /(.75*.2 + .24*.14 + .01*1) = 5%




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: