Disclaimer: I do not like to read LLM-generated text any more than anyone else.
IMHO a big problem with Pangram in particular is that they market it as a reliable tool that can be used to catch students cheating. This can obviously have disastrous effects on young lives, because it is not as reliable as they suggest.
Per their own benchmarks, they do not achieve 100% accuracy even on text that is published on the Internet, and which is likely encoded into the models themselves.
There is validity to their goals, but that is overshadowed by the irresponsible way in which it is marketed.
(All of this, swirling in a context where students are being told that they absolutely must become proficient at using LLMs to do exactly this kind of work by the highest levels of state and federal governments, faculty leadership, as well as the leaders of the workforce into which they hope to graduate. The message to youth is extremely muddled at best.)
The false positive rate for Pangram 4 is something like one in 24,000.[0] To put that in perspective, the wrongful-conviction (false positive rate) for death-sentenced defendants in the US is estimated conservatively to be around 4.1%.[1] The FP rate for death-sentence convictions is 1,000 times bigger than Pangram’s FP rate.
Now, the US criminal system is not a great yardstick for justice. But it goes to show you Pangram is really good evidence that something was LLM generated. It can be an amazing tool for enforcing AI policies in schools, and there ought to be ways to use it with caveats for the rare but inevitable false positives (appeals, etc).
There are a couple of statistical errors in your argument here.
First is frequency. Even using Pangram's claimed numbers, the University of Georgia should expect to see several false positives every week. Remember that the metric is # of assignments run through Pangram, not number of students. A campus of 40k students will see many more than 40k assignments every week, and so should expect honest students to be accused of cheating with some high degree of frequency. You're comparing infrequent events (death penalty sentences) to high-frequency events (students submitting assignments).
And obviously, you are citing a company marketing document as fact, of which we should all be suspicious. (There are also obvious problems with the eval dataset that the paper does not address.)
Second, you're using the upper bound for Pangram's claimed numbers and the lower bound cited in the NIH publication.
> at least 4.1% would be exonerated. We conclude that this is a conservative estimate of the proportion of false conviction among death sentences in the United States.
> The false positive rate for Pangram 4 is something like one in 24,000.
Gotta suck to be one of the 8B/24k=~300k people in the world whose writing pattern is falsely labelled as slop by this tool that people say is so accurate so customers are going to feel really sure about your alleged dishonesty about writing your own texts
This false positive rate is a double-edged sword. Please still be careful when accusing people
I don't think 100% accuracy is logically possible. Because it's entirely possible that someone would just naturally write the exact same thing as an LLM would write. And after the fact there is no way to distinguish the two. But pangram does have an extremely low false positive rate, which I think does make it useful for detecting cheating students. Assuming the base rate of cheating students is 1%, and assuming pangram has a false positive rate of 1 in 10,000 and a true positive rate of 7,000 in 10,000, that means ~98% of students flagged by pangram actually cheated. Combined with a teacher's familiarity with that student's previous work, which should rule out many more false positives, it should be a very useful tool.
> ~98% of students flagged by pangram actually cheated
That 2% is a large number! Of people who will have their integrity impugned for no good reason! That's not okay!
Your calculations also are mixing assignments and students. The rate of false positives of 1/10k is of corpuses, not students. 10k students might each submit 2-3 written assignments per week. Obviously, this greatly increases the impact of the false positive rate.
And all of these numbers are dependent on lab conditions for usage, which are not the case in the real world.
> I don't think 100% accuracy is logically possible.
Yes. Which is why marketing this product as it currently is, is a deeply irresponsible endeavor.
IMHO a big problem with Pangram in particular is that they market it as a reliable tool that can be used to catch students cheating. This can obviously have disastrous effects on young lives, because it is not as reliable as they suggest.
Per their own benchmarks, they do not achieve 100% accuracy even on text that is published on the Internet, and which is likely encoded into the models themselves.
There is validity to their goals, but that is overshadowed by the irresponsible way in which it is marketed.
(All of this, swirling in a context where students are being told that they absolutely must become proficient at using LLMs to do exactly this kind of work by the highest levels of state and federal governments, faculty leadership, as well as the leaders of the workforce into which they hope to graduate. The message to youth is extremely muddled at best.)