Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I really like these set of lecture notes for optimizing matrix multiplication: https://ppc.cs.aalto.fi/ch2/v7/ (The transpose trick is used in v1)


I find it surprising that, even after using all those tricks, they are still only to achieve around 50% of the theoretical peak performance of the chip in terms of GFLOPS. And that's for matrix multiplication, which is a nearly ideal case for these techniques.


Hmm where did you see that? In the graph at the bottom of v7, it got "93% of the theoretical maximum performance"


This deserves it's own submission, wonderful resource!




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: