Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You could think of what they did in the first study as constructing an exam to test how well various LLM's do as an advice columnist. They wanted a lot of personal advice questions where the LLM should not affirm by default. If a few questions with wrong answers got in there, it probably wouldn't affect the results all that much?

Unfortunately they didn't test anything newer than GPT4o, so we don't know how much GPT-5 improved. It would be nice if someone turn their list of questions into a benchmark.



They actually did test GPT-5: https://www.science.org/doi/10.1126/science.aec8352 (see the figure under Conclusion). Its rate of endorsement of user action, 52%, was the same as GPT-4o. So based on their setup it seems that the newer model didn't reduce affirmation.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: