Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

And not just on super fancy computers, 128GB of RAM is like $500 and is supported by many low-to-mid-range desktop systems.


Without ECC that's a risky way to compute.


Eyeballing the numbers [0][1] I would expect one bitflip between once every year and once every 100 years on 128GB of memory. I'm willing to take those odds.

Of course with some planning you can get an AMD system with ECC support on the mainboard; ECC RAM is about the same price as consumer RAM.

0: https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf

1: https://www.cs.virginia.edu/~gurumurthi/papers/asplos15.pdf


When Craig Silverstein left Google somebody asked him "What was the biggest mistake you made while at Google?" His response was "Not using ECC memory."

The problem is that both the consequences and debugging time of a single random bitflip are potentially unbounded. That time that GMail lost 10% of their accounts and had to restore from tape was because of a single bitflip (due to a software bug, not a cosmic ray). Google Search lost months of engineering time in the early days from cosmic rays, back when they really could've used those engineers on other stuff.

The odds may be low, but they do happen when you have multiple computers, and the consequences are high. Not really odds I'd want to take, when ECC RAM isn't that much more expensive than non-ECC RAM.


Since Google was building their own servers it seems like they might have been able to sidestep the ECC tax too. Normally ECC costs so much more because it's the "serious user" solution so vendors feel free to mark it up.


I kinda wonder if we’ll come out the other side of this DRY era and rediscover redundant systems. Turning really important decisions into a single bit in memory is overoptimization. Especially when all of the less important bits are consuming gigabytes.


How does ECC RAM protect against software bugs?


To be clear, that sentence was in there to illustrate the effect that single-bit errors could cause. That incident happened after Google had already long since switched to ECC and obviously wouldn't have been prevented by it.

However, we have a large number of other processes (testing, type systems, formal verification, code reviews, release processes, etc.) to protect against software bugs. There is no protection against cosmic rays. You don't want to be in a situation where all of the defect-mitigation work that the last 50 years of computer science has accomplished is rendered useless by a random freak occurrence.

(The bug in question was actually in a migration script, and made it into production because people thought that migration scripts were one-off throwaways that didn't need the same amount of testing, code review, verification, and general carefulness that the production code does. Lesson learned. The postmortem for it actually had the lesson of "Treat your migration code as permanent, and apply all the same standards of maintainability and reliability of it that you do to production code.")


It's also why ZFS requires the use of ECC memory in the official documentation - ZFS spent great efforts building redundancy and error-checking capabilities as part of the filesystem, especially for guarding against silent data corruption, even at the expense of performance. But it would be useless and greatly decrease the benefits of these features if the memory can fail silently.

Also, hitting by a beam of cosmic ray is not the only way that the bits in RAM can be flipped, dynamic RAM has inherent instabilities like row hammering, or can fail early due to manufacturing defects.


Could explain how hardware would have prevented a software bug?


> Eyeballing the numbers [0][1] I would expect one bitflip between once every year and once every 100 years on 128GB of memory. I'm willing to take those odds.

Throw in the likely hood of the flipped bit being consequential and the odds look even better. To potentially do major damage the flipped bit would have to be in an area of memory of something being executed and it has to flip after being read and before being executed, which probably ads at least 2 more orders of magnitude. Even for general bitrot of data has to be in memory and get flipped between reads and writes. Those odds are vanishingly small compared to programmer error.


IIRC ryzen in general can run with ecc, though motherboard support is a little shaky


The low to midrange Xeons with this kind of memory capacity have ECC.

If you're not buying for your particular specifications, you're doing it wrong anyway... There are plenty of workloads tolerant to ECC errors (e.g. just about any kind of simulation).




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: