Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Isn’t this likely from bias in the training data? The system is more sensitive to label something as hate if that group is more likely to experience hate on the internet. How the system responds to “Blacks” vs “African-Americans” is a perfect example of this. The latter has historically been perceived as more respectful so it won’t be used as often in the hate speech in the training data. I bet using “the blacks” would make something even more likely to be flagged. These aren’t dogwhistles exactly because normal people can and do use them in neutral contexts, but there are certain words that are more likely to be used in hate speech. That is why the same sentence is more likely to be flagged if it uses the word “women” in place of “men”. That isn’t a intentional bias towards women. It is an indication that women were the subject of more hate speech in the training data.


Of course it is a bias in the training data, but it's probably not the dataset that you're thinking of. So far as we can tell, the filtering part doesn't come from the main corpus, but rather from human-guided moderation - basically, people voting on whether any given answer is "hateful" or not. ChatGPT filters reflect the biases of that later group (or, perhaps, the biases of the people who instructed them).


It seems like the obvious solution would be to replace words that refer to specific races, nationalities, genders, religions, etc. with symbols that only indicate the category so if someone ranks "the Danes are rude" as hateful that is translated to "the [nationality] are rude" and is applied equally to a comment that says "the Swedes are rude". That way they can continue using the same system while eliminating the most obvious sources of discrimination.


But that is the same problem. There aren't any human reviewing the requests as they come in. That human guided moderation gets converted into some type of model that isn't any smarter than ChatGPT. That model is just as susceptible to this issue.


I mean, how could they? The model responds faster than a human can read. Should they introduce minute long waits and make it 100x more expensive?


This is addressed in the article!

The general point is that if this theory were true, we wouldn't expect significantly more bias against Republicans than against Democrats. Hence ChatGPT having a general left-wing bias (which was also confirmed in other tests, linked in the article) is the simpler explanation. People on the left generally judge hatred against majority groups and Republicans as less bad.


It’s quite possible there’s more hateful speech against Democrats than Republicans.


Likewise the classic “racist bans affect Republicans disproportionately because more racists are Republican”.


That doesn't completely settle the issue because Democrats and Republicans use different words and styles of phrase in their polemics.


Exactly. Take the "Democrat Party" versus the "Democratic Party". They both look neutral, but a language model will code the former as more of an insult. This is because it correlates with insults because it resembles language more often used by conservative who are more likely to insult Democrats than liberals are.

Liberals will put in more effort to avoid language that looks like hate speech because they are more concerned about being politically correct. Therefore language and topics that are more frequently discussed by conservatives will be more likely to resemble language that is used in hate speech because there is less effort put in to avoid that.


Yes, that one leapt out at me in the comparison between 'Democrat voter' and 'Republican voter'. The term 'Democrat' is almost always used as a slur, while 'Republican' is just the mainstream term.


I would suggest you read some of reddit.


Generally, the equivalent slur of the left is wantonly labeling things as "conservative". Democrat is kind of odd manifestation as a slur, but I guess represents the parties disdain more directly. The left usually isn't that direct.


Yes, I agree this is the most likely explanation. Just the inclusion of "fat" or "poor" (and slurs or any mention of race, obviously), or any word which is often used to insult someone will make text more likely to be flagged as hateful by a moderator in a training set.

I do still think that the men/women bias and the Democrat/Republican bias both make more sense as originating in moderators favoring one group over the other, since none of these are typically used as an insult by themselves.


>I do still think that the men/women bias and the Democrat/Republican bias both make more sense as originating in moderators favoring one group over the other, since none of these are typically used as an insult by themselves.

They may not be used as insults themselves, but they are used more in hate speech.

Republicans are generally more opposed to the idea of "hate speech" as a category of speech and are therefore less likely to identify any speech as hate speech. Democrats have embraced it more as a concept, are more likely to label something as hate speech, and are more likely to think that type of speech is bad. It therefore seems likely that Democrats would use what a neutral observer would categorize as hate speech less than Republicans due to self-censorship. That would result in the word "Republicans" appearing in hate speech less often than "Democrats" because the hate speech infused insults will be targeting the opposite party.


[flagged]


How am I blaming the victim? I am simply pointing out that language can indicate bias without necessarily being biased itself.

I am guessing that if we applied a similar model to Russian and English, the model would indicate there is an inherent bias against the west in Russian and a bias against Russia in English. That is all were seeing here. It isn't actually indicating anything about the language. It is telling us about who uses the language and how they use it.

Words, phrases, and linguistic approaches that are generally coded as conservative will be more likely to denigrate liberals and vice versa. Conservative speech will be more likely to be flagged for hate speech because conservatives by and large care less about being PC. It is important to reiterate that does not mean conservatives are necessarily any more racist. Their speech just correlates more with racists speech because less effort is put into avoiding that correlation.


> I am guessing that if we applied a similar model to Russian and English, the model would indicate there is an inherent bias against the west in Russian and a bias against Russia in English. That is all were seeing here. It isn't actually indicating anything about the language. It is telling us about who uses the language and how they use it.

So if your argument is that the model has been trained on a collection of information about what is hate speech assembled by liberals, then I can see it might be possible.

But if your argument is that (speculatively) republicans engage more in hate speech, then bad things said about republicans is not detected as hate speech by the model, the jump is rather far.


>So if your argument is that the model has been trained on a collection of information about what is hate speech assembled by liberals, then I can see it might be possible.

It doesn't specifically need to be "assembled by liberals" to have a liberal bias. Liberal people are more likely to categorize anything as hate speech than conservatives. Liberals think being PC is important. Conservatives are generally dismissive of being PC. Even if there is no inherent bias in the makeup of this hypothetical review panel, the panel will result in ruling that are more in line with liberal thought because conservatives are less likely to take an active lead in labeling hate speech.

>But if your argument is that (speculatively) republicans engage more in hate speech, then bad things said about republicans is not detected as hate speech by the model, the jump is rather far.

My argument is that these systems can't actually identify hate speech. The question isn't whether Republicans engage in hate speech more frequently. They likely engage in speech that resembles hate speech more frequently because they don't care about being PC.

Usage of the term Latinx is an example. I have heard valid arguments why people should are shouldn't use that term, however its usage is currently much more common in liberal circles. Therefore a phrase using "latinx" instead of "latino" is going to be less correlated with hate speech because racists just aren't using "latinx".


It's changed drastically since release. There are lots of people out there who have noticed this and many have saved examples of before and after responses to prompts.


How did it change?


The "content moderation system" is new, so I don't think it changed. What, however, changed during the time ChatGPT is live, is what kind of prompts it refuses to answer, because the topic is offensive/inappropriate. It had hilarious versions where it would tell you a joke about men about not about women or one ethnicity but not the other.


> The latter has historically been perceived as more respectful

Maybe if you only consider Americans. But rest assured, many black people do not want to be called African or American. Because they are neither.


Yes, I agree. I thought the double qualifiers of "historically" and "perceived" would indicate that I don't personally agree with the notion, but American society at large has agreed with that for most of the last 50 or so years.


My point is that it's perceived that way within America. But the internet is larger than just America. And presumably/hopefully ChatGPT gobbled up all kinds of data originating from other countries.


I'm astonished so many commenters assume the bias originates from the training data and nobody seems to scrutinize the adjectives that are used in this test.


> The system is more sensitive to label something as hate if that group is more likely to experience hate on the internet

No it is more sensitive to things that have already been labeled as hate in the training data. So much "hate" (whatever that may be) against unfavorited groups goes by online without anyone batting an eye.


That's exactly what I'm thinking. Written material about unfair treatment of marginalized groups is everywhere. The reverse is not true. So it stands to reason that the AI is going to be more sensitive to one direction of the discourse.

One of the priors that is going unspoken here is that "The Truth Is Politically Neutral". And... that's not always correct. I mean, to borrow the libertarian angle here: do we want the AI to tell us what we want to hear or do we want it to tell us the truth?


While this is fair, there is no excuse for the difference in the flagging of republicans vs. democrats.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: