Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Probably not; you could try gluing the models together but it's unlikely it'd improve them and I'm not sure they have the same architecture. Also, even the small/medium GPT3 models are much much worse than the largest one (DaVinci).

On the other hand, the lottery ticket hypothesis says that every good big network contains a good or better smaller network you could extract, if only you knew where it was. (So the reason big models are good is there's more opportunities for the smaller models to appear.)



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: