It's not flamebait at all. The data is all about people's responses to generic questions about what they're looking for in a date. They don't say "I don't want to date Muslims" and have that reflected in match scores. The match score is based on things like "I hate people who squeeze the toothpaste from the middle."
To put data based on 'race & religion' out in the open to me is already a thing that is not excusable (besides the lousy definition of such things as race and religion, both of which are floating, not discrete).
There is a massive pre-selection problem here (people that frequent dating sites, rather than people in general) and the link to 'hacking' is a very tenuous one.
I've seen Michael Jackson referred to as a hacker here, so I probably shouldn't be too surprised.
This is not a hack, this is some statistical analysis with input data of dubious value, pretty pictures and all, drawing conclusions that are completely off the wall.
"Jews and Agnostics get along better with people"
"Muslims of both sexes and Hindu men get along worse"
"Catholics are more universally liked than Protestants"
Please, if that isn't flamebait I really don't know what is.
Where are the control groups, where are the standard deviations and so on.
Effectively this is a dating advert for programmers.
This is infotainment at it's best, statistical noise at its worst.
I'll agree that it's not a hack - that doesn't, in my mind, preclude it from being hacking. They're two different things. For me, anyway, playing with giant data sets does not fall into the first group, but does fall into the second.
That said, "is it hacking or not" isn't really a terribly interesting meta-discussion to have, imho, so I'll try and go in a different direction.
Please, if that isn't flamebait I really don't know what is.
The data appears to support the assertions you are calling flamebait. Provided their algorithm for calculating likeability actually works, then the data does support the assertions[1]. (This is all, of course, within the context of the pre-selection you noted in your post).
[1] Note that the assertions are of the form "Members of Set X are more likeable to members of Set Y than members of Set Z". If the assertions were "Being a member of Set X makes you more likeable to members of Set Y", they wouldn't necessarily be supported by the data.
If you're going to make this kind of generalizing statement about large groups of people I simply think that you have the responsibility to do it properly, so with control groups and so on, or not at all, otherwise the results are totally meaningless.
The way it is presented right now is as if simply the size of OkCupid and their data gathering methods give them the license to make this kind of claim.
It's 'interesting' but not 'rigorous'.
The three I listed above are particularly galling, I really don't think any computing algorithm can give someone license to make the statements listed above (and as an atheist I have no dog in that fight), but without proper methods it's even worse.
I think that you're assigning more meaning to the words "like" and "get along with" than the authors intend. Pretty much everywhere they use one of those phrases it should be qualified like "X _say they_ like Y". But that gets tedious, and I think the authors are also making the assumption that you were paying attention when they explained where these numbers come from.
Note that the end of the article is leading directly into the objections that people keep making along the lines of "this doesn't mean that these people really get along in real life", and the conflict between what people say they are looking for and the choices they actually make about who to contact and respond to.
Of course it's not rigorous. To some degree people interacting with OKCupid's site are a self-selecting bunch. You really need to take this at face value. This is just a blog post showing some number crunching on their site. There are no sensationalist headings like "Muslim Males the Most Hated Group of People." Keep in mind that this is also not a scientific journal, not peer-reviewed research... nor is it claiming to be.
It's way past my bedtime here, and judging by the moderation I'm not able to make my point, which is simply this:
If you are going to be making sweeping statements about people, even including race and religion then you really should do your homework, or if you're doing it out of curiosity, keep the results to yourself. By presenting the data in a format that looks as though very hard work went in to its creation and by hammering home the reliability of that data you are creating the illusion of something that is scientifically solid when in fact it isn't.
I agree, I'm glad to get the chance to read about this stuff, but I think the PP's point was that if you're going to present a bunch of data in that manner, it's best to say, up front, something like, "our sample, while extensive within our service, is not necessarily representative of particular ethnic groups," and point out something akin to what slashdot says about their polls: you're insane if you intend to do anything serious with this data.
You're right. It's infotainment. So? Are you incapable of taking anything at face value? They aren't making claims about how the world works; they're making claims about how people interact on their site. If that's not interesting to you, flag it and move on.
The flaw here is that how people interact on their site has very little bearing on how those same people will interact in real life.
The plural of data isn't evidence, even for large amounts of data.
Sure you can do interesting statistical analysis, but the real action is after two people have found each other, and that's where the huge flaw is in all this analysis, the statements are about interactions in real life, the data is gathered online.
Any kind of statement about people being more or less likable would have to be taken out of the context of the website and into real life, without that statements as listed above are unsupportable.
Gotta agree with you on this one. It's amusing to look at the plots and note the bits that agree (or disagree) with your prejudices, but without some information about the variance this is nothing more than entertainment.
It's quite interesting.