I find it surprising that, even after using all those tricks, they are still only to achieve around 50% of the theoretical peak performance of the chip in terms of GFLOPS. And that's for matrix multiplication, which is a nearly ideal case for these techniques.