"Complete" might be taking it too far, because averages can tolerate a lot of error underneath. Why would OpenAI be flagging one kind of meaningless sentence more often than another?
It's easy to poke holes in studies by taking issue with individual data points, but unless you can either explain how those problems lead to the overall conclusion, or can do the study again with better methods and get a different answer, the original case remains standing.
>Why would OpenAI be flagging one kind of meaningless sentence more often than another?
Well, let's consider a specific example. His first chart is women/men. Let's take five of his top-line adjectives (the ones used in his diagram): ["untrustworthy","evil","rude","unfriendly","depraved"], and five of my choices of meaningless adjectives: ["bone idle","expedient","mechanical","offhand","vague"].
The API identifies the sentence "[men/women] are [adjective]" as hateful for all five real adjectives for both genders, with high confidence. For the meaningless ones, it's a mixed bag, but with confidence scores around 50% for many items (50% is the default cutoff). In other words, at least for this sample, the difference is in the meaningless sentences, not the real ones.
Now, maybe it just so happens that I got lucky with my pick of five. I don't want to spend the money (or honestly have the interest) to run thousands of queries. But suppose this pattern holds. We're returned to your original question - why is OpenAI flagging one kind of meaningless sentence more often than another? My answer is that it doesn't matter. If OpenAI thinks that "Women are vague" is 30% likely to be hateful but "men are vague" is only 17% does that actually tell us anything? Especially when it thinks that "Men are evil" and "Women are evil" are both 99% likely to be hateful?
>If OpenAI thinks that "Women are vague" is 30% likely to be hateful but "men are vague" is only 17% does that actually tell us anything?
If that held out over hundreds of "placebo epithets," it would tell you that the filter was using the presence of the word "woman" as a signifier for hate speech independently of its context. You wouldn't be able to discover that fact by looking at things that got scored 99% because they are already at the top of the scale, and made equal by that effect.
The real world example will not be a three word statements like "Women are evil", but rather long sentences like this one I am currently writing that include the phrase "Women are evil" twice.
The question is not about the AI ability to detect hateful content in ideal sentences. The question is if there is a bias when the AI has to make a judgement call.
We can see the same thing with face recognition. There is no race bias in AI detection in perfect lightning when the person is facing the camera perfectly. There is however a very noticeable bias when the AI is less certain using real world examples where light and positioning is far from perfect. As the data become less meaningful, the bias in favor of white skin increases.
The study would be improved by doing an additional in-depth study with real world text that has been selected by humans, and then modify the input by randomizing the target demographic. If the bias remains then we would have a higher confidence in the data. This is similar to studies done in face recognition where issues with dark skins has been demonstrated multiple times.
It's easy to poke holes in studies by taking issue with individual data points, but unless you can either explain how those problems lead to the overall conclusion, or can do the study again with better methods and get a different answer, the original case remains standing.