Engineering
Fault-Tolerant Design of a Microprocessor Using Triple Modular Redundancy
Quick fact
Triple modular redundancy (TMR) is a fault-tolerant design technique that uses three identical processors and a majority voter to mask faults. If one processor produces an incorrect result, the other two agree on the correct output, so the system continues operating correctly without any downtime.
Why this is interesting
Have you ever wondered how airplanes and spacecraft keep flying even when their onboard computers face hardware failures? In such systems, a single electronic failure cannot be allowed to cause a crash, and so designers use a clever trick: they replicate the entire computer three times and let a vote decide what to do.
Read the full explanation
Understanding Fault-Tolerant Design of a Microprocessor Using Triple Modular Redundancy
Imagine you need to build a computer that must never stop working even if one part breaks. One basic idea is to have two computers and compare their outputs; if they disagree, you know something is wrong, but you don't know which one is correct. The next step is to use three computers. Now, if one disagrees with the other two, you can be fairly confident that the majority is right because they both produce the same output. This is the principle behind triple modular redundancy, or TMR. In TMR, you replace a single, vulnerable processing unit with three identical copies of that unit. All three copies receive the same inputs and execute the same instructions in lockstep. Their outputs are fed into a 'majority voter' (also called a 'voter' or 'majority circuit'). The voter simply compares the three outputs and selects the one that appears on at least two of the three lines. If one module produces a wrong output due to a hardware fault, the other two will still agree, and the voter will pass the correct output to the rest of the system.
A deeper explanation
The engine of TMR is fault masking. The voter hides a single module's failure from the system, so the overall system behaves as if no fault occurred, and there is no need to detect or recover from the error. This is in contrast to error detection and reconfiguration, where the system stops, identifies the fault, and then reconfigures to bypass the faulty unit. The underlying principle is majority voting. The voter is a simple logic circuit that outputs the majority value among its three inputs. For binary signals, it is a straightforward combinational logic function: if two or more inputs are 1, the output is 1; if two or more are 0, the output is 0. This can be implemented with AND-OR logic or with a majority gate. The three modules must be designed to fail independently. That is, if one module fails due to a defective transistor, the other two should not fail in the same way from the same cause (e.g., at the same location on the chip). This is achieved through physical separation, independent power supplies, and different manufacturing processes, though perfect independence is impossible. TMR is a powerful method for ensuring high reliability, but it comes with a significant cost: it triples the hardware area, power consumption, and cost. So it is used primarily in safety-critical systems like avionics, aerospace, and medical devices, where the cost is justified by the need for extreme reliability. In practice, TMR is often applied not to the entire microprocessor but to critical sub-components, such as registers or ALUs, to balance cost and fault coverage.