I go to some pain in my talks to say "instruction count is not a good proxy for performance", but unfortunately folks do still use it. It's handy to say "hey, there's no loop in this output" or "this loop does 3 multiplies; the alternative does two and an add" or similar. It's a mixed blessing to have brought the compiler output to the masses, I can only hope it starts a useful learning process!
Now any load from the L3 cache memory or from the main memory takes much more time than any other instruction (not counting exceptions generated by instructions, which include many memory accesses that slow them down, or deprecated instructions that are kept for backwards compatibility and that are executed by long microcode sequences).