In the previous article, we took a close look at the three major forms of DNA, their characteristics, and how they differ from one another. Depending on the geometry adopted by the same two strands as they form a double helix, DNA can be broadly classified into three forms: A-, B-, and Z-DNA. Among these, the A-form is found mainly in double-stranded RNA and RNA:DNA hybrids, although DNA can also adopt an A-form structure under special conditions such as dehydration. Under physiological conditions, however, the form normally adopted by DNA in our cells is the right-handed B-form double helix, built from Watson–Crick (WC) base pairs. This is the standard DNA structure with which we are most familiar. When torsional stress develops in DNA during processes such as transcription, some regions can locally switch into Z-DNA, a left-handed helix with a characteristic zigzag backbone. In other words, DNA generally exists in the B-form, but under particular conditions and environmental changes inside the cell, local regions can shift into other forms and then return to the standard B-form. DNA, it turns out, is a much more flexible and dynamic molecule than we might have imagined.
But the structural differences among A-, B-, and Z-DNA that we looked at in the previous article have one thing in common. In all three cases, the basic framework of a double helix, with two strands paired together, remains intact. But DNA’s transformations do not stop there. When a situation cannot easily be accommodated within the framework of the double helix, DNA can boldly step outside this familiar structure and take on much more unconventional forms.
What kind of situation could make DNA abandon such a stable double helix and choose an entirely different structure? And what new possibilities might these unfamiliar structures open up for DNA? There is a Korean saying that “a woman is innocent of changing her look,” but for DNA, transformation is essential. In this article, let us explore some of DNA’s bolder—and well-motivated—transformations beyond the familiar B-form double helix.
Non-B DNA structures: non-canonical DNA structures
The DNA structure most familiar to us, and the one most commonly described in textbooks, is the B-form double helix based on WC base pairs. In the actual cellular environment, however, DNA can form a variety of structures with different geometries, collectively referred to as non-B DNA structures. Rather than simply being abnormally distorted DNA, these non-B DNA structures can be thought of as alternative structural states that DNA can adopt under particular sequence and environmental conditions. Physical stresses and different chemical environments arising during complex processes such as replication, transcription, and recombination can provide conditions for these structures to form. This structural diversity also expands the ways in which DNA and proteins can recognize and interact with one another and plays important roles in regulating DNA function. Known non-B DNA structures include hairpins, cruciforms, triplex DNA, and G4 (quadruplex) structures, which form selectively under particular sequence and physical conditions.
Before looking at why DNA might leave the stable B-form double helix and adopt a non-B structure, and what consequences this can have, let us first look at the conditions that make these structures possible. The standard base pairing familiar to us is based on WC hydrogen bonding, in which A pairs with T and G with C. These complementary base pairs help maintain a uniform helix width and stabilize DNA. We have already looked closely in an earlier article at which atoms in each base act as hydrogen-bond acceptors and donors in WC base pairing. What is interesting, however, is that WC pairing is not the only hydrogen-bonding pattern available to the bases. Around the flat perimeter of each base are different bonding edges, and depending on which edge meets which partner, base pairs with entirely different geometries and stabilities can form. In addition to the familiar standard WC pairing, there are many other modes of interaction, including reverse WC, Hoogsteen, reverse Hoogsteen, and wobble base pairing. These alternative pairing modes depart from the standard WC double-helix arrangement and create new structural possibilities. When such non-canonical interactions are repeatedly used in particular sequences, they can make higher-order alternative structures such as triple helices and quadruplexes possible rather than a single ordinary double helix. The reason a single base can participate in so many different modes of interaction can be found in the different hydrogen-bonding edges around its perimeter.
Three hydrogen-bonding edges of the bases
A base is a planar aromatic ring, and around its perimeter, atoms capable of participating in hydrogen bonding are exposed in slightly different directions. These are broadly classified into three hydrogen-bonding edges: the Watson–Crick edge, the Hoogsteen edge, and the sugar edge. Rather than representing three spatially separate surfaces, these edges are better understood as a way of describing the directions in which bases interact when they form hydrogen bonds with one another. As shown in the image below, where the atoms participating in the three hydrogen-bonding edges are distinguished by color, the same atom can participate in two different edges. This shows that the edges are not spatially isolated compartments, but rather a classification based on the direction from which a base is viewed or on the partner with which it interacts. The N6 atom of adenine, for example, can act as a hydrogen-bond donor in both the WC and Hoogsteen edges. The difference is that in a WC base pair it participates from the interior of the double helix, whereas in a Hoogsteen base pair it participates through the major-groove side. The sugar edge, for reference, is the hydrogen-bonding edge located on the sugar side of the base. In RNA, it is particularly important in the formation of tertiary structures because the 2′-OH group provides additional opportunities for interaction. We will return to this when we discuss RNA.
Many combinations of these three edges are possible when bases pair with one another. WC-WC, WC-H, WC-S, H-S, and S-S are just some of the many possible hydrogen-bonding combinations. But base-pair geometry depends not only on which edges meet; it also depends on the relative orientation of the sugars attached to the two bases, described as cis or trans. The standard A-T and G-C base pairs familiar to us are cis Watson–Crick/Watson–Crick pairs, in which both bases use their WC edges.
Reverse Watson–Crick base pairing
In a reverse WC base pair, both bases still use their WC edges, but the sugars attached to the two bases are positioned on opposite sides of the base pair, giving a trans arrangement. As a result, the combination of hydrogen-bond donors and acceptors is the same as in a standard WC pair, but the spatial orientation of the base pair is reversed, changing the direction of the sugar-phosphate backbones and the C1′–C1′ distance. Such base pairs are observed more stably in non-canonical environments, such as internal loops and folded structures in RNA, than in a regular double helix such as B-DNA. Even the same pair of bases can therefore form completely different hydrogen-bonding geometries depending on which edges face one another and whether the sugars attached to the two bases point in the same or opposite directions. The image below shows several different hydrogen-bonding geometries that can be formed by an A-T base pair.
Hoogsteen base pairing
Whereas WC base pairing occurs toward the interior of the double helix, Hoogsteen base pairing involves a purine—adenine or guanine—shifting its orientation from anti to syn and forming hydrogen bonds through its Hoogsteen edge. When the orientation between the sugar and base becomes syn, the base moves from beside the sugar to a position more directly over it. This shortens the C1′–C1′ distance, makes the base pair more compact, and tilts it toward the major groove, exposing certain atoms on the major-groove side. Hoogsteen pairing may sound quite unfamiliar, but Hoogsteen and WC base pairs actually exist in a dynamic equilibrium and have the potential to interconvert in either direction depending on the circumstances. Under ordinary conditions, more than 99% remain in the WC state, while Hoogsteen states are very short-lived and occur at very low abundance (<1%). Under mechanical stress, however—for example, when DNA must bend sharply or fit into a confined space as it wraps around histone proteins during DNA compaction—the more compact geometry of a Hoogsteen base pair can make it less unfavorable than the WC geometry. WC base pairs can therefore transiently switch into Hoogsteen geometry as an alternative way of accommodating changes in their environment. These fleeting transitions provide yet another example of the structural flexibility of DNA. There can, however, be a price for this transition. Unlike WC base pairs, in which the bases fit tightly together inside the double helix and reactive atoms are buried within it, Hoogsteen pairing exposes some atoms on the purine WC edge to the surrounding solvent through the major groove. This can make them more accessible to chemical reactions such as methylation, oxidation, and alkylation, potentially increasing their susceptibility to damage. Such damage can interfere with the activity of DNA polymerases during replication, causing replication errors that may ultimately lead to mutations.
Through these additional bonding edges, DNA can form a variety of non-B DNA structures rather than being limited to the B-form double helix. Interestingly, when we examine the regions in which these non-canonical structures form, we find that they are commonly rich in repetitive sequences rather than random sequences. Because identical or very similar base sequences are arranged in repeating patterns, they can have a kind of physical symmetry. This symmetry makes DNA more prone to deformation by twisting or folding in particular directions, facilitating the formation of non-B structures. Z-DNA is also sometimes included in the category of non-B DNA structures because repetitive sequences are involved in its formation. By understanding repetitive sequences, we can better understand how they are related to non-B structures and how these structures form.
Repetitive DNA sequences
Non-B DNA structures are closely associated with repetitive sequences because alternative structures can form readily in repeat-rich regions of DNA. DNA is often described as a blueprint for making proteins, but in reality, protein-coding sequences account for only about 1–2% of the human genome, and even when sequences coding for functional RNAs are included, the proportion remains relatively small. Repetitive regions, by contrast, make up more than 50% of the human genome. Although they were once often dismissed as nonfunctional “junk DNA,” it has become increasingly clear that repetitive sequences can play important roles in many different aspects of genome biology.[1]
Let us look at some of the general functions and roles of repetitive sequences. First, repetitive DNA is involved in transcriptional regulation. When identical or similar sequences are repeated, they can provide more binding sites for particular transcription factors or DNA-binding proteins, allowing the strength and timing of gene expression to be finely regulated. Repetitive sequences can also influence the structural state of chromatin. Certain repeats can help maintain particular regions in a more loosely packed, open euchromatin state or, conversely, promote a more condensed and closed heterochromatin state, thereby helping regulate gene accessibility. They play particularly important roles in forming structural regions of chromosomes and maintaining stability at sites such as centromeres and telomeres. Perhaps most importantly, repetitive sequences also contribute to genome evolution. Their length can change relatively easily through processes such as replication slippage, which we will examine below, or unequal crossing-over. (During homologous recombination in meiosis, corresponding regions of homologous chromosomes normally align so that genetic information is exchanged between matching positions. But because repetitive sequences create identical or similar sequences at multiple positions, the chromosomes can sometimes pair with the wrong copy and become misaligned. Crossing-over in this misaligned state can produce structural variations such as gene duplications or deletions.) This variability makes repetitive sequences hotspots for mutation and recombination. Such properties can contribute to evolution by promoting the appearance of new traits and accelerating genomic change that can help organisms adapt to new environments. At the same time, however, this flexibility and variability can increase genomic instability and also make the formation of non-B DNA structures more likely.
Tandem repeats and interspersed repeats
Repetitive sequences can be broadly classified according to their arrangement into tandem repeats and interspersed repeats. Tandem repeats consist of the same sequence repeated continuously next to itself and are often referred to as satellite DNA. Depending on the length of the repeat unit, they can be divided into microsatellites, minisatellites, and macrosatellites. In terms of repeat-motif length, microsatellites contain repeat units of about 1–4 base pairs (bp), minisatellites about 5–64 bp, and macrosatellites can contain repeating units thousands of base pairs long. Among tandem repeats, microsatellites—also known as short tandem repeats (STRs)—have high mutation rates because they consist of many copies of short repeat units. Repetition of the same sequence also increases opportunities for these sequences to interact with one another, so when the DNA becomes exposed as a single strand, microsatellites can provide physical conditions that favor the formation of various non-B DNA structures.[2]
Interspersed repeats, by contrast, are produced largely by transposable elements (TEs), often nicknamed “jumping genes.” Transposable elements are DNA elements capable of spreading their sequences within the genome. LINEs, SINEs, and LTR elements are representative examples. Among these, retrotransposons such as LINEs, SINEs, and LTR elements spread copies of their sequences throughout the genome through a “copy-and-paste” process involving reverse transcription. As a result, identical or similar sequences become scattered across many genomic locations. Importantly, transposable elements do more than simply increase the amount of repetitive DNA. Because they can carry regulatory elements such as promoters, enhancers, and transcription-factor binding sites as they spread, they can also reshape gene-regulatory networks. Many transposable elements also contain G- or C-rich sequences, making them prone to forming non-B DNA structures such as the G-quadruplexes and i-motifs that we will discuss later. When inserted near genes, they can therefore become directly involved in structure-based transcriptional regulation.
Repetitive sequences, then, are more than simply DNA sequences that happen to repeat. They play important roles connecting gene regulation, chromatin structure, evolution, and the formation of non-B DNA structures. But why are non-B structures particularly likely to form in repetitive regions? Repeated sequences and their structural flexibility increase the number of ways in which DNA can fold and rearrange itself. Such rearrangements can in turn create structure-based regulatory mechanisms that allow gene expression to be adjusted rapidly and reversibly. From this perspective, non-B DNA structures formed in repetitive regions have a dual nature: on one hand, they can increase genomic instability; on the other, they can contribute to the functional flexibility and regulation of the genome. To understand how different non-B structures arise from repetitive sequences, let us first look at replication slippage, a phenomenon that clearly illustrates how the number of repeats in a short repetitive sequence can change.
Direct repeats and replication slippage
As helicase unwinds the DNA double helix and DNA polymerase moves along the template strand during replication, the polymerase can sometimes pause, temporarily dissociate from the DNA, and then rebind to continue synthesis. If this happens within a repetitive region, it may become difficult to tell exactly which repeat unit synthesis should resume from. As a result, the strands can realign incorrectly at an identical repeat sequence, with synthesis resuming at a position ahead of or behind the point where it originally stopped. In a sense, the polymerase has become confused by a row of identical-looking repeats. This misalignment of the strands within a repetitive region is called replication slippage, or slipped-strand mispairing, and it occurs particularly often in short repetitive sequences such as microsatellites. Repeats in which the same sequence is repeated in the same orientation are called direct repeats.
If the newly synthesized nascent strand slips out and forms a loop, and the polymerase then rebinds to a repeat behind it, an additional repeat unit can be copied, resulting in an insertion. If, on the other hand, the template strand slips and forms a loop, the polymerase can skip over that region without reading it, so the repeat is not copied, resulting in a deletion. These misaligned structures can create a bulge within the double helix. In most cases, however, the DNA mismatch repair (MMR) system detects this structural change and repairs and corrects it, so it does not become a lasting problem. If the repair system fails and the error remains uncorrected into the next round of replication, however, the change in repeat length can become fixed as a mutation.
Hairpins and cruciform structures
Repetitive sequences not only promote slippage during replication; they can also allow DNA to pair with itself when it is temporarily exposed as a single strand. In particular, when an inverted repeat is present, the exposed single strand does not have to wait for its complementary opposite strand. Instead, it can fold back on itself through intramolecular base pairing and form a stable stem-loop secondary structure. In this process, complementary sequence regions form WC base pairs to create a double-stranded stem, while the non-complementary region between the two repeats remains as a single-stranded loop. This entire structure is called a hairpin. The non-complementary connecting sequence between the inverted repeats is called a spacer or intervening sequence, and when the hairpin forms, this region becomes the loop.
An inverted repeat consists of two reverse-complementary sequences arranged in opposite orientations within the same strand, with a non-complementary spacer sometimes present between them—for example, 5′-CTAG … GATC-3′. A perfect inverted repeat with no spacer is called a palindrome. In DNA, this means that the sequence read 5′→3′ on one strand is identical to the sequence read 5′→3′ on its complementary strand. This is somewhat like a palindrome in everyday language—a word or phrase that reads the same forward and backward—but with the important difference that a DNA palindrome must take both complementarity between the two strands and their directionality into account.
As the length of a repetitive sequence increases through slippage, this kind of self-folding can occur more readily. In addition, when negatively supercoiled DNA locally unwinds, both strands can form hairpin structures at the same time. When this happens, the double helix can be converted into a cross-shaped structure known as a cruciform.
These structures are better understood not as structures deliberately created for a particular purpose, but as alternative DNA structures that can arise naturally under physical conditions created by repetitive sequences and supercoiling. Their presence can affect the rate at which polymerases and transcription machinery move along DNA and can alter patterns of protein binding, allowing them to act as subtle modulators of gene expression. As is so often the case, however, problems can arise when repetitive regions expand too far and these structures become excessively stabilized, making the genome unstable and potentially contributing to diseases associated with disrupted gene regulation. RNA is somewhat different. Because its diverse shapes are directly tied to its functions, hairpins serve as fundamental building blocks for functional RNA folding and play key roles in both structure formation and catalytic activity. By making use of hairpins, RNA can form remarkably complex and diverse structures.
[References]
[1] Repetitive DNA sequence detection and its role in the human genome
[2] Structural and functional significance of microsatellites
[Image sources]
Image 22-1. Non-canonical DNA structures: Diversity and disease association — CC BY 4.0
Image 3
3-1. Structures and stability of simple DNA repeats from bacteria — CC BY 4.0
Image 4
4-1. Insertions and deletions in protein evolution and engineering — CC BY-NC-ND 4.0
4-2. https://academic.oup.com/gpb/article/5/1/7/721067
Image 5
5-1. Structures and stability of simple DNA repeats from bacteria — CC BY 4.0




