Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

A consumer CPU like a 285k caps out around 130 GB/s of memory bandwidth.

Each of its 24 cores can do two 8 wide FMA ops per cycle. Lets say holding a continuous 4 GHz clock speed.

This works out to over 1.5 trillion 32 bit floating point multiplies per cycle.

If you are doing vector matrix multiplies (like in single token no batching). `xW` then each weight loaded sort of gets used in 1 multiplication and 1 addition.

Doing the math you can clearly see even if each weight were just 1 byte you can at most load 130 billion of them in a second from memory.

But in the same timespan you could have done over 1.5 trillion multiplications.

So you are still memory bound.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: