Reminds me of the R4300 multiply bug, and the lazy compiler fix that just always emitted a nop between them if it saw two mults in a row. It was a real, hardware bug tho.
Reminds me of a lab in college where we were doing MIPS assembler. This particular lab involved output to the console using trap handlers (basically a glorified hello world). I was the only one in the lab to get it to work at the time, and I could explain how. I had a random assignment of a strange hex number to, I think, a register (it may have been a memory location) preceded by a comment to the effect of: "Do not touch this. I'm drunk. I don't know why this works or how, but it does. Sober you will not understand." And, true enough, I didn't. And this was before google was a thing, so I'm still not sure how/where I found that magic number to make things work.
Very often hardware bugs can be worked around by well-placed NOP.
Last time I did that was on MSP430, where it seems that some integrated peripherals set completion interrupt flag one cycle before the operation really completes. (I'm not sure whether it really is bug, I didn't find this documented anywhere, but for TI's example this does not matter and produces marginally better performance)
Which MSPs, if I may? I use some of their 2xxx hobbyist line (since I'm a hobbyist), and this sounds like the kind of PITA that I'm not nearly enough of an embedded dev to avoid tearing my hair out over.
The MSP430x5xx and MSP430x6xx families have this problem. From the user-guide [1]:
> Include at least one instruction between the clear of an interrupt enable or interrupt flag and the EINT instruction. For example: Insert a NOP instruction in
front of the EINT instruction.
> Include at least one instruction between DINT and the start of an code
sequence that requires protection from interrupts. For example: Insert a NOP
instruction after the DINT
Luckily if you use the MSP compiler's intrinsic functions to enable/disable interrupts it takes care of this for you.
What I described happened on G2553 with USCI_A in SPI mode. Given the fact that various code samples on the net do not handle this behavior it seems that it is either happening only for some configurations or only on some versions of the HW. But: when you see even bytes on SPI dropped or replaced with previous ones, add NOP.
SPARC and other classic RISC pipelined architectures require you to put a nop "Bubble" in the pipeline after a branch if you can't use the instruction slot that's always executed right after the branch.
The democoder in me thinks of "that's not a bug, it's a feature" and wonders if, like the branch delay slot, the delayed memory effect was used to squeeze in an extra instruction and save a cycle or two in optimised code.
A similar trick can be found in this interesting article: