In the previous two articles, we took a close look at the nitrogenous bases—the real heart of DNA, where genetic information is actually encoded. We explored why bases with aromatic structures were selected, and what invisible forces allow these flat, ring-shaped molecules to build long nucleic acid structures: pairing with the bases across from them while stacking one above another, all without the structure falling apart. Now my curiosity turns back to the question of how these bases pair with one another. Why do the four bases always pair with particular complementary partners? Let us take a closer look at how bases find their match.
Base Pairing Between Pyrimidines and Purines
A base pair consists of one pyrimidine paired with one purine. This combination keeps the overall dimensions of the base pairs consistent, helping the DNA double helix maintain a nearly uniform width. We can see just how similar the two base pairs are by comparing their dimensions. If we measure the distance between the C1′ atoms of the two sugars—the positions where each base forms its glycosidic bond with the sugar—both A-T and G-C base pairs span roughly 11 Å.[1] (An ångström, Å, is 10⁻¹⁰ m, or one ten-billionth of a meter.) If different base pairs differed greatly in size, the double helix would not be able to maintain a consistent width, making it difficult to form a stable helical structure. In other words, complementary base pairing plays an important role in maintaining the geometric stability of DNA.
Watson–Crick Base Pairs
This complementary pairing is known as Watson–Crick base pairing. In the model proposed by Watson and Crick for the DNA double helix, adenine (A) selectively pairs with thymine (T), while guanine (G) pairs with cytosine (C), with each pair following a specific pattern of hydrogen bonding. A and T are connected by two hydrogen bonds, whereas G and C form three. Within each base pair, highly electronegative nitrogen and oxygen atoms participate in hydrogen bonding with hydrogen atoms. The figure below shows the partial charge distributions produced by differences in electronegativity within the base pairs and how these give hydrogen bonds their directionality. The dotted lines represent hydrogen bonds. Notice how the hydrogen-bond donor, hydrogen, and acceptor tend to align in an almost straight line, giving these interactions their strong directional character.
In an A-T base pair, adenine and thymine each provide one hydrogen-bond donor and one acceptor, forming two hydrogen bonds in total. In a G-C base pair, guanine provides two donors and one acceptor, while cytosine provides one donor and two acceptors, producing three hydrogen bonds in total. With one additional hydrogen bond, a G-C base pair binds somewhat more strongly than an A-T base pair.
This made me wonder about something. Since hydrogen bonds are highly directional interactions and require the right geometric arrangement to form, perhaps a G-C base pair, with its extra hydrogen bond, would be more difficult to form correctly than an A-T pair. And once formed, perhaps it would also be more stable—and therefore require more energy to break apart. That led me to another question: could G-C base pairs be more common in regions where especially high accuracy of DNA replication is required? I went looking for evidence, but I could not find strong support for such an idea.
The fact that G-C forms one more hydrogen bond than A-T certainly matters. But when it comes to the accuracy of DNA replication, something more important than this difference of a single hydrogen bond is that correctly formed A-T and G-C Watson–Crick base pairs have nearly the same overall dimensions and geometry. The strong directionality of hydrogen bonding also plays an important role in producing this proper geometry. When a DNA polymerase accepts and incorporates an incoming nucleotide, it is known to evaluate not only its chemical interactions but also whether the resulting base pair has the correct geometry to fit precisely within the active site—in other words, whether the shape of the base pair is right.
G-C base pairs also have one more hydrogen bond, and together with the effects of base stacking and the surrounding sequence, this means that DNA with a higher G-C content generally tends to be more thermally stable than DNA with a higher A-T content. Conversely, A-T-rich DNA tends to separate into its two strands at somewhat lower temperatures. This brings to mind certain promoter sequences involved in transcriptional regulation, such as the TATA box, where A-T-rich regions are relatively easier to open because their base pairing is weaker. But this difference cannot be explained simply by counting hydrogen bonds. As we saw in the previous article, base stacking contributes just as importantly to the stability of DNA structure.
In proteins, a backbone of linked amino acids provides the structural framework while the various side chains carry out many of the functions. DNA and RNA have a somewhat similar organization: the sugar-phosphate backbone provides the basic structural framework, with the bases attached along its side. There is, however, an important difference. Whereas protein secondary structure is formed largely through hydrogen bonding involving the backbone, the structures of DNA and RNA depend heavily on hydrogen bonding between bases and on base stacking. RNA in particular can form a remarkable variety of secondary structures through interactions between its bases, including stems (or helices), loops, and pseudoknots.
So the base pairs of DNA sit securely inside the double helix, held together by hydrogen bonds. But then how can DNA-binding proteins, such as transcription factors and regulatory proteins, recognize a particular DNA sequence without first opening the double helix? The answer lies in two structural grooves that run along the surface of the DNA double helix: the major groove and the minor groove.
Major and Minor Grooves: Gateways for Protein Recognition of DNA Sequences
Purine bases—adenine and guanine—form glycosidic bonds with the C1′ atom of the sugar through their N9 position, while pyrimidine bases—thymine and cytosine—do so through their N1 position. This arrangement allows the bases to attach to the sugar-phosphate backbone without disrupting the hydrogen-bonding patterns between base pairs, while preserving the aromatic π-electron systems of the base rings and allowing a stable spatial arrangement without steric clashes.
In simplified drawings of DNA, we often place a base pair in the center and draw the two sugar-phosphate backbones as though they were positioned symmetrically on either side. But the real structure is a little different. Imagine looking down onto the plane of a base pair and drawing an imaginary reference line through its center. The two C1′ atoms of the sugars—the points connecting the bases to the sugar-phosphate backbones—are not positioned symmetrically with respect to this central line. Instead, the line connecting the two C1′ atoms is shifted toward one side of the base pair. In this sense, a base pair is inherently geometrically asymmetric.
Because of this asymmetry, the sugar-phosphate backbones come closer together along one edge of the base pair, leaving less space there, while the opposite edge is more widely exposed. Imagine the base pair as a flat plate. The two sugar-phosphate backbones are not attached symmetrically at the middle of opposite sides of that plate; instead, their attachment points are shifted toward one side. As a result, one side becomes crowded, with the sugars relatively close together, while the other side opens into a wider space. The figure below illustrates this geometry.
As these asymmetric base pairs stack one after another, oriented nearly perpendicular to the helix axis, they form the double helix and the sugar-phosphate backbones wind around the axis in a spiral. Along the side where the two backbones lie closer together, this repeating arrangement creates a deep, narrow channel: the minor groove. On the opposite side, the backbones remain farther apart, producing a wider, more open channel: the major groove. In other words, these two grooves are the three-dimensional consequence of the asymmetric positions at which the sugars attach to each individual base pair, repeated continuously along the axis of the double helix.
The figure above shows the C1′–C1′ asymmetry within the plane of an individual base pair, while the figure below shows how this asymmetry appears as two different grooves in the three-dimensional double helix. The measured widths of these grooves can vary depending on exactly how the measurement is defined—for example, whether it is measured from sugar to sugar or from phosphate to phosphate. For standard B-form DNA, the major groove is commonly described as being about 22 Å wide and the minor groove about 12 Å wide.
So what do these two grooves actually mean, and what do they do? They play a crucial role when the many proteins that bind DNA search for their specific targets along the molecule. The bases are tucked away inside the DNA double helix, making the inner pairing surfaces of the bases difficult to access without opening the helix and breaking the hydrogen bonds between the strands. The major and minor grooves, however, provide important gateways through which DNA-binding proteins—including transcription factors and regulatory proteins—can read information from the bases without having to open the DNA.
Through these grooves, the outer edges of the base rings become exposed, revealing the identity of the base pairs within. Different base pairs present different arrangements of hydrogen-bond donors and acceptors, as well as different substituent groups along their exposed edges. A protein can therefore feel its way along these differences and determine which base pair occupies a particular position. For example, if a methyl group (CH₃) is encountered in the major groove, it can indicate the C5 methyl group of thymine and therefore identify an A-T base pair. In this way, proteins can use a kind of chemical fingerprint to read DNA sequences and regulate transcription.
The wide, information-rich major groove provides detailed clues about which base-pair combination is present. The narrower minor groove contains less chemical information, making precise discrimination between base pairs more difficult. For example, A-T and T-A base pairs present nearly identical chemical patterns in the minor groove, as do G-C and C-G pairs, making it difficult to distinguish the orientation of the base pair precisely. Because the minor groove is more confined by the sugar-phosphate backbones, the variety and amount of chemical information exposed there are more limited. Proteins that bind in the minor groove therefore often recognize their binding sites not only by directly reading the chemistry of the bases, but also through structural features of the DNA itself, such as groove width, curvature, bendability, and the local electrostatic environment.
Base Readout and Shape Readout
When a protein interacts with and binds to DNA, directly recognizing the chemical groups of the bases is called base readout, while recognizing the geometric shape of the DNA rather than directly reading the bases is called shape readout. In other words, DNA-binding proteins can use these two strategies—base readout and shape readout—to locate particular binding sites. In reality, the two do not necessarily operate as completely separate mechanisms; many DNA-binding proteins use information from both.
Let us look at an example of base readout. The figure below was generated using the DNA Readout Viewer (DRV), software designed to analyze and visualize how much a protein relies on base readout versus shape readout when recognizing DNA. Here, it shows how the estrogen receptor interacts with DNA.[2] The estrogen receptor (ER) is a transcription factor. When the hormone estrogen binds to ER, the receptor undergoes a conformational change and then binds to an estrogen response element (ERE) in DNA to regulate transcription. The image below visualizes this interaction. Residues of the ER protein form hydrogen bonds with specific atoms of the bases exposed in the major groove. This is a good example of a protein literally reading the bases and binding to them—a classic case of base readout.
Now let us turn to shape readout. Instead of reading the bases themselves, a protein reads the three-dimensional structure of the DNA. It recognizes physical features such as the overall width of the DNA, the width and depth of its grooves, its curvature, and its electrostatic environment—in effect, reading the kind of structure that a particular base sequence creates and using that information to bind.
The Narrow Minor Groove Found in Certain AT-Rich Regions
Let us see how the base sequence itself can influence the structure of DNA. Certain AT-rich sequences—particularly A-tracts—have a characteristically narrow minor groove. If we look back at the base-pair diagram above, the atoms exposed by an A-T base pair toward the minor groove include adenine N3 and thymine O2. Both carry lone pairs of electrons, allowing them to act as hydrogen-bond acceptors and giving them a partially negative character. When these hydrogen-bond acceptors line the floor of a narrow minor groove, water molecules can become arranged between them in an ordered orientation, forming a structured network of water molecules that hydrogen-bond with the bases. This spine of hydration is closely associated with the narrow minor groove found in A-tracts and helps stabilize an already narrow structure. In other words, rather than the hydration spine being the sole cause that makes the minor groove narrow, the sequence-dependent narrow groove and the ordered water molecules work together to create a stable structure.
An Example of Shape Readout: TBP, the TATA-Binding Protein
One representative protein that uses such structural features of DNA to recognize a particular sequence is TBP, the TATA-binding protein. Whereas many transcription factors recognize base-pair information through the major groove, TBP—as its name suggests—binds to the TATA box through the minor groove, and in doing so dramatically deforms the DNA.
This structural deformation—the bending of DNA—is an important preparatory step for the initiation of transcription. Before transcription can begin, the pre-initiation complex (PIC) must assemble on the DNA. This is a substantial molecular complex composed of many proteins, including TBP, several general transcription factors, and RNA polymerase II. As illustrated in the image below, the structure created when TBP bends the DNA provides an important platform that helps these transcription factors and RNA polymerase II come together in the proper spatial arrangement to assemble the transcription initiation machinery.
But how does TBP actually bend the DNA? As its name suggests, the TATA box is rich in A and T bases, and TBP targets the minor groove of this sequence.
At this site, TBP reorganizes the hydration network while simultaneously inserting four phenylalanine (Phe) side chains—one pair near each end of the TATA box—between the bases.[3] Like wedges driven in from both sides, these side chains slip between the bases, disrupting base stacking at two sites and sharply kinking the DNA.
But why phenylalanine? Because phenylalanine has an aromatic side chain containing a benzene ring. Once inserted between the bases, its aromatic ring can come into close contact with the aromatic rings of the DNA bases and participate in various noncovalent interactions, while disrupting the π–π stacking that originally existed between neighboring bases. This produces two kinks, and together they can bend the DNA sharply—by around 80°—toward the major-groove side. Ultimately, the structural properties created by the TATA-box sequence, together with the binding of TBP, dramatically reshape the DNA and provide an important foundation for the transcription initiation process that follows.
[References]
[1] New insights into Hoogsteen base pairs in DNA duplexes from a structure-based survey
https://doi.org/10.1093/nar/gkv241
[2] DNA Readout Viewer (DRV): visualization of specificity determining patterns of protein-binding DNA segments
https://doi.org/10.1093/bioinformatics/btz906
[3] The 2.1-Å crystal structure of an archaeal preinitiation complex: TATA-box-binding protein /transcription factor (II)B core/TATA-box
https://doi.org/10.1073/pnas.94.12.6042
[Image sources]
Figure 1 1-1. DNA chemical structure — CC0 1.0
Figure 3
3-1. DNA — CC BY-SA 3.0
Figure 4
4-1. DNA Readout Viewer (DRV): visualization of specificity determining patterns of protein-binding DNA segments — CC BY-NC 4.0
Figure 5
5-1. Molecule of the Month: TATA-Binding Protein — CC BY 4.0
5-2. TATA Binding Protein Structure — CC BY-SA 3.0
Figure 3
3-1. DNA — CC BY-SA 3.0
Figure 4
4-1. DNA Readout Viewer (DRV): visualization of specificity determining patterns of protein-binding DNA segments — CC BY-NC 4.0
Figure 5
5-1. Molecule of the Month: TATA-Binding Protein — CC BY 4.0
5-2. TATA Binding Protein Structure — CC BY-SA 3.0




