Chemistry
DNA Replication Fidelity: Base Pairing and Proofreading
Quick fact
DNA polymerase is so accurate that it introduces an error only about once per 10^5 base pairs, and with proofreading that rate drops to roughly one in 10^7 to 10^9.
Why this is interesting
When a cell copies its DNA, it makes only about one mistake for every ten billion letters copied. How does a molecular machine achieve such precision?
Read the full explanation
Understanding DNA Replication Fidelity: Base Pairing and Proofreading
Imagine copying a long document by hand. You might make a typo, but if you also read back what you wrote while writing, you can catch and correct most slips. DNA replication works similarly. DNA is a double helix composed of two strands. Each strand is a sequence of nucleotides, each with a base: adenine (A), thymine (T), cytosine (C), or guanine (G). A always pairs with T, and C with G, via hydrogen bonds. When the cell divides, an enzyme called DNA polymerase moves along the original strand, reading each base and adding the complementary base to the new strand. The first layer of fidelity is base pairing: the chemical shapes and hydrogen bonding patterns strongly favor the correct match. For example, A pairs with T via two hydrogen bonds, but pairing A with C would introduce mismatched hydrogen bond donors/acceptors. This geometric and chemical complementarity provides a selectivity of roughly one error per 10^5 nucleotides. But that's still too many errors for a large genome. So DNA polymerase has a second 'spellchecker' function: as it adds a nucleotide, it checks whether the newly added base is properly paired. If not, it removes it with a proofreading exonuclease activity. This 'edit while writing' reduces the error rate to around one per billion.
A deeper explanation
The mechanism of replication fidelity relies on two sequential kinetic checkpoints. The first is during nucleotide incorporation. DNA polymerase selects the incoming nucleotide triphosphate based on geometric complementarity with the template base. The active site of the enzyme is shaped to fit Watson-Crick base pairs, but mismatched pairs distort the DNA double helix, and the enzyme detects this distortion. However, even with this selection, occasional misincorporations occur (about 1 in 10^5). The second checkpoint is the proofreading step. DNA polymerase has a 3'→5' exonuclease active site that is separate from the polymerase active site. When a mismatched base is incorporated, the 3' end of the new strand becomes frayed and unstable because the mismatched pair does not form proper hydrogen bonds. This fraying causes the new strand to slip back into the exonuclease site, where the incorrect base is cleaved off. The polymerase then resumes synthesis in the forward direction. This proofreading step further reduces the error rate by a factor of 10^2 to 10^3. Additional post-replication mismatch repair systems catch errors that escape both checkpoints, bringing the total error rate to ~10^-10 per base pair. The energy for proofreading comes from the hydrolysis of the phosphodiester bond, not from ATP, and the process is driven by the thermodynamic instability of the mismatched base pair. This two-layer system is a beautiful example of how biological processes achieve high precision through kinetic proofreading.