Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is the sort of thing that shows why the whole "pure RISC philosophy" of implementing only the simplest instructions is ultimately a pretty dead-end in processor design. CISC-style dedicated instructions and hardware which utilises them will always be more efficient than a pure software implementation using just "simple RISC instructions", because hardware can easily express specialised computations like the extreme parallelism of algorithms such as SHA while software implementations essentially rely on the instruction-scheduling mechanisms to arrange potentially hundreds or even thousands of instructions in an optimal order.

In addition, when there is dedicated hardware, it is also significantly easier to add a SHA instruction than attempt to recognise sequences of many regular instructions that could be "fused together" into a single operation using that hardware. When there isn't, adding such an instruction that gets internally expanded into multiple uops for the equivalent using the existing functional units is still beneficial, since it leaves open room for an immediate performance gain on all existing software when a future revision does add that dedicated hardware, or optimises its microarchitecture to allow those (possibly different) uops to execute faster.

Despite the name, ARM is definitely not very RISC anymore, and that's what has kept it competitive.



> CISC-style dedicated instructions and hardware which utilises them will always be more efficient than a pure software implementation using just "simple RISC instructions"

For a counterexample, see the x86 string instructions. "Hardware" implementations like `repne scasb` are routinely outperformed by software implementations using SSE2.

Another problem is that these instructions don't die. SHA will some day be replaced, but the instruction will live on. The x86 BCD instructions illustrate this.

From what I remember, RISC is less about "simple" instructions, and more about regular instructions which can execute with predictable throughput, ideally one per clock cycle. "Do this particular arithmetic" is in the RISC philosophy, while looping instructions (rep, lswi, etc.) are not.


For a counterexample, see the x86 string instructions. "Hardware" implementations like `repne scasb` are routinely outperformed by software implementations using SSE2.

That's only because SCAS (and CMPS) has not (yet) received quite the same amount of attention as MOVS and STOS. For a counterexample to your counterexample, look up "enhanced REP MOVSB". REP MOVS/STOS can operation on cacheline-sized blocks since at least the P6, when it was introduced as the "fast strings" feature, and its performance has been steadily improved over the processor generations.

Another problem is that these instructions don't die. SHA will some day be replaced, but the instruction will live on. The x86 BCD instructions illustrate this.

Replaced for secure crypto, yes, but there are plenty of other applications like (nonmalicious) data corruption detection where a reasonably fast yet far more collision-resistant algorithm than regular CRC is very useful.

On the topic of BCD instructions, there's this: https://news.ycombinator.com/item?id=8477254


I do agree, but I wouldn't include vector instructions in "simple RISC instructions".


I think RISC-V is showing that RISC ISAs can still compete in this space: https://arxiv.org/pdf/1607.02318v1.pdf. RISC-V has extensibility of the ISA as a core feature, allowing for special-purpose accelerators and coprocessors as needed.


Or there is the middle ground of retaining the general purpose nature by making the "dedicated" hardware an FPGA. Potentially the best of both worlds. Xlinx has the Zynq (ARM+FPGA), Intel has bought Altera and is now doing CPUs with FPGAs in the package [2].

[1] http://www.xilinx.com/products/silicon-devices/soc.html

[2] http://www.pcworld.com/article/3055526/intel-starts-baking-s...


This can work for special-purpose projects, but might become problematic in general: reprogramming the FPGA can be slow (on the order of 10s of milliseconds) and you don't necessarily want to include FPGA state as part of task state in the OS. You also need to be careful about thermal limits and power draw (especially when reconfiguring). I doubt current-generation FPGAs will land in anything smaller than a server for a long time. There, you can have semi-permanent accelerators according to the server's role, and a larger thermal budget.


This article is referring specifically to ARM64 assembly which is actually quite a bit more RISC-y than previous versions of the ARM ISA. In favor of simplicity, many of the more complex ARMv7 instructions (such as the block loads/stores) were removed in ARMv8. I'd still agree that ARM is not very RISC anymore, but the ongoing move to ARM64 takes it quite a bit back in that direction.


Bingo.

If and when speed matters, ASIC beats a general-purpose CPU hands down, every time.


Also in sync with the micro-coded CPUs of Burroughs, Xerox PARC, ETHZ, Genera and so on, even if they eventually lost to more general purpose CISC instructions.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: