Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Technology

Understanding Computer Reliability at Scale

Quick fact

A typical computer has a failure rate of about one undetected error per several hundred years of operation, thanks to built-in error detection and correction mechanisms.

Why this is interesting

Your computer performs billions of operations every second, yet it almost never crashes. How does it manage to be so reliable when its components are made of atoms and can fail?

Read the full explanation

Understanding Understanding Computer Reliability at Scale

Imagine you're copying a book by hand. You might make a typo. To catch it, you could check your work against the original letter by letter. Computers face a similar problem, but at a massive scale: every bit of data (a 0 or 1) can be accidentally flipped by electromagnetic interference, cosmic rays, or manufacturing defects. Since computers process billions of bits per second, errors are inevitable. However, they are also rare—only about one bit flip per 10^14 operations. The trick is to detect and correct these errors before they cause a crash or corrupt data. Computers use clever tricks like parity checks and error-correcting codes to do this. A simple example: in a parity check, each group of data bits is given an extra bit that makes the total number of 1s even. If a single bit flips, the parity becomes odd, and the error is detected. But detection isn't enough for critical systems; they also correct the error by storing redundant copies or using more sophisticated codes.

A deeper explanation

The core principle is redundancy: adding extra information or hardware to detect and recover from errors. There are two levels: detection and correction. Detection methods like parity checks or checksums (e.g., CRC) can identify that an error occurred, but not which bit is wrong. Correction methods, such as Error-Correcting Codes (ECC), not only detect but also pinpoint the faulty bit. ECC memory, used in servers, stores extra bits that form a mathematical code allowing the system to detect and correct single-bit errors and detect double-bit errors. This is done automatically by the memory controller without halting the processor. For even higher reliability, systems use redundancy at the component level: Triple Modular Redundancy (TMR) runs the same computation on three separate processors and votes on the result. If one fails, the other two produce the correct answer. This is used in spacecraft and critical flight controls. Reliability also involves design practices like error handling in software (checking for invalid values, retrying failed operations) and hardware design that avoids single points of failure. By layering these mechanisms, computers achieve an astonishingly low failure rate: a typical server has an expected failure rate of less than 0.01% per year for a single memory bit, and even if a hard drive dies, RAID systems can rebuild data from redundant copies. This robustness is why we trust computers for tasks from email to controlling power grids.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.