Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I just find it insane that we're bootstrapping reinforcement learning and world planning on top of basic next token prediction.

I'm amazed that it works, but also amazed that this is the approach being prioritized.



Pure RL NN 'solved' simple games like Pokémon years ago. I think added challenge of seeing how well LLMs can generalize is a noble pursuit. I think games are a fun problem as well.

Look how poorly Claude 3.7 is doing on Pokemon on Twitch right now.


> Pure RL NN 'solved' simple games like Pokémon years ago.

Please link to said project. From my search of Google filtered to 2010-2020, it returns nothing outside of proofs-of-concept (e.g. https://github.com/poke-AI/poke.AI) that do not perform any better, or instead trying to solve Pokemon battles which are an order of magnitude easier.


There is this amazing video [1] of some guy training a pure RL neural network to play Pokémon Red. It's not that old and the problem was certainly never completely solved.

[1] https://youtu.be/DcYLT37ImBY


Maybe they are conflating the Starcraft success that Deepmind had with AlphaStar?


And AlphaStar or OpenAI Five were playing games knowing internal game states and variables

Playing games from pixels only is still is a pretty hard problem


Codebullet on YouTube remakes games and then makes the computer beat the game.

Because pixels are hard.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: