For DNA to move beyond its familiar double helix and take on various non-canonical structures, it cannot simply fold and bind in any sequence it wants. Just as remodeling a house requires the right structure and conditions to create new spaces, DNA also needs certain basic conditions that allow it to form structures different from the ordinary double helix.
So, before taking a closer look at non-canonical DNA structures themselves, the previous article first explored these basic conditions. In repetitive sequences, base slippage can occur relatively easily during replication, and inverted repeats in particular can form hairpins when complementary regions within a single strand pair with each other, or cruciform structures when such structures form in both strands. We also saw that DNA bases have additional hydrogen-bonding edges, such as the Hoogsteen edge, besides the familiar Watson–Crick edge, allowing bases to interact with one another in ways other than standard base pairing.
With this background in place, let us now take a closer look at the non-canonical structures DNA can form, one by one. In this article, we will explore the DNA triplex, in which a third strand joins the existing two strands, and in the next article, we will look at four-stranded structures formed when four strands come together.
DNA triplex
A DNA triplex forms when a third strand binds to an existing double helix. The original duplex maintains its standard Watson–Crick base pairing, while the third strand binds through the additional hydrogen-bonding surface of the DNA bases—the Hoogsteen edge that we looked at in the previous article. In other words, both Watson–Crick and Hoogsteen base pairing can be seen within a single structure.
Intermolecular vs intramolecular DNA triplexes
Depending on where the third strand comes from, DNA triplexes can be divided into intermolecular and intramolecular triplexes. An intermolecular triplex forms when an independent strand from outside—the aptly named triplex-forming oligonucleotide (TFO)—binds to a purine-rich repetitive region of double-stranded DNA. An intramolecular triplex, on the other hand, forms when a repetitive sequence within the same DNA molecule folds back on itself, bending into a structure known as H-DNA. Typical H-DNA readily forms in mirror-repeat regions composed of polypurine and polypyrimidine sequences.
The key to triplex formation is the presence of a continuous stretch of purines in one strand of the existing duplex. This is especially important for an intramolecular triplex because the third strand enters along the major groove and must continuously recognize the Hoogsteen edges of the existing Watson–Crick base pairs. Purines are particularly suited for this because their double-ring structure presents a consistent arrangement of hydrogen-bond acceptors and donors toward the major groove, making them more favorable than pyrimidines. As the third strand binds continuously along one strand of the duplex, the same geometric pattern needs to repeat for a continuous triplex to form. If the sequence is not composed entirely of purines—a homopurine stretch—but instead alternates between different bases, the hydrogen-bonding pattern exposed in the major groove changes, making stable binding of the third strand difficult. After all, when putting LEGO blocks together, wouldn’t they hold more securely if all the holes you need were there rather than missing here and there?
Direction of the third strand: parallel vs antiparallel and the anti conformation
When the third strand binds within the major groove of an existing Watson–Crick duplex, its arrangement can be classified as parallel or antiparallel according to the direction in which it runs. In a parallel triplex, the third strand runs in the same direction as the purine strand it recognizes and forms Hoogsteen hydrogen bonds. This arrangement is mainly seen in the pyrimidine motif, in which the third strand consists of a continuous stretch of pyrimidines. Here, a motif refers to a recurring sequence pattern with a particular base composition. Triplets such as T·A–T and C⁺·G–C are formed in this arrangement. In an antiparallel triplex, by contrast, the third strand runs in the opposite direction to the purine strand and forms reverse Hoogsteen hydrogen bonds. This arrangement is mainly observed in the purine motif, where purines occur consecutively, forming triplets such as A·A–T and G·G–C. In addition, although thymine (T) is not a purine, triplets containing T can also form reverse Hoogsteen bonds in the antiparallel arrangement.
In this way, the orientation and stability of a triplex are determined by the base composition of the third strand and the repetitive pattern of the target sequence. In addition, the bases of the third strand generally adopt an anti conformation relative to the sugar-phosphate backbone. This allows the bases to be properly exposed toward the major groove so that Hoogsteen or reverse Hoogsteen bonds can form. Maintaining the anti conformation in the third strand is therefore a representative structural principle of triplex formation. The image below uses G as an example to show structurally why the third strand adopts the anti conformation in a DNA triplex.
pH dependence of the pyrimidine-motif TFO
When the third strand consists of pyrimidines—the pyrimidine motif—it enters the major groove and binds parallel to the purine strand of the duplex. An important point here is that the N3 of C (cytidine) in the third strand must be protonated (H⁺) to form a stable bond with the Hoogsteen edge of G (guanine). In its unprotonated state, C has two acceptors for Hoogsteen hydrogen bonding but lacks a donor. Therefore, as pH decreases and protonation of C increases, formation of the C⁺·G–C triplet becomes more favorable. Intermolecular pyrimidine-motif triplexes in particular tend to become unstable at physiological pH, whereas intramolecular triplexes can remain stable even at neutral pH. The purine motif, on the other hand, does not depend on protonation of C, forms reverse Hoogsteen bonds, and can be stabilized by divalent cations such as Mg²⁺.
A prerequisite for intramolecular triplex H-DNA formation: mirror repeats
An intramolecular triplex, in which the third strand comes from within the same DNA molecule, forms naturally when part of the DNA folds back on itself. Unlike a triplex formed by a TFO, it typically forms in mirror-repeat sequences composed of polypurine and polypyrimidine tracts. A mirror repeat is a sequence in which the two sides are symmetrical around a central point, as if reflected in a mirror. In other words, a region gives the same sequence when read from 5′→3′ in one direction and from 3′→5′ in the other, rather like folding it exactly in half. The difference from the inverted repeats we looked at earlier is that mirror repeats have symmetry but do not require complementarity. These mirror-repeat regions are also composed of a polypurine tract and its complementary polypyrimidine tract. When negative superhelical stress loosens the duplex, or when the duplex opens for transcription or replication and exposes single-stranded DNA, such mirror-repeat regions can fold back on themselves to form a triplex. One half of the mirror repeat maintains Watson–Crick pairing to form a duplex, while the other half folds into the major groove of its own duplex and forms Hoogsteen bonds, creating the third strand. The remaining strand is left single-stranded. This is easier to understand when seen in an image. Because this intramolecular folding bends the DNA helix into a hinge-like shape, the resulting structure is called H-DNA.
Depending on the base composition of the third strand—whether it is a pyrimidine motif (Hy) composed of pyrimidines or a purine motif (Hu) composed of purines—and whether the third strand comes from the 3′ or 5′ side of the repeat, four different topologies can form: Hy3, Hy5, Hu3, and Hu5. The remaining part of the repeat that does not participate in triplex formation is left as a single-stranded loop. Thus, even with the same sequence information, completely different non-canonical structures can form depending on how the strands are rearranged.
DNA triplexes and disease
When DNA replication or transcription begins, double-stranded DNA locally unwinds, exposing regions of single-stranded DNA. As DNA polymerase or RNA polymerase moves forward, torsional stress gradually builds up in the DNA, producing positive supercoiling ahead of the moving polymerase and negative supercoiling behind it. In regions of negative supercoiling in particular, the DNA duplex is more prone to local unwinding, creating a physical environment that favors the formation of non-canonical DNA structures at certain sequences. Under these conditions, DNA containing mirror repeats or polypurine and polypyrimidine tracts can fold internally and form triplex (H-DNA) structures. Rather than being functional structures deliberately formed by the cell, these triplexes often arise spontaneously when the right sequence and physical conditions come together. Once formed, however, they can physically interfere with DNA replication and transcription, acting as obstacles that slow or even stop the progression of DNA polymerase or RNA polymerase. This can lead to delayed replication, interrupted transcription, and increased instability of the replication fork. Even without fully understanding every detail of replication and transcription or the mechanisms of all the enzymes involved, it is easy to imagine how such an obstacle could disrupt processes that normally need to proceed smoothly.
DNA triplex structures are in fact known to be associated with various structural alterations, including DNA deletions, duplications, inversions, and translocations. These structures can interfere with replication fork progression, causing fork stalling or collapse, and double-strand breaks (DSBs) can also occur in the process. If errors continue to accumulate as DNA repair proceeds, genome instability can increase.
Of course, cells also actively intervene to reduce these structural risks by recruiting various proteins. Helicases such as WRN, BLM, and FANCJ can unwind triplex structures and resolve the structural problem, topoisomerases relieve superhelical stress, and DNA repair systems repair damaged regions. The problem arises when repetitive sequences have expanded excessively or when structural resolution systems, including helicases, fail to arrive and act at the right time. In such cases, triplex structures can continue to form, increasing genome instability and eventually contributing to disease.
In particular, during DNA replication, triplex structures can form more readily when the purine-rich strand serves as the lagging-strand template or in regions where single-stranded DNA remains exposed for a relatively long time. This can cause the replication fork to stall or collapse, potentially resulting in double-strand breaks. In addition, DNA damage-sensing systems may recognize these abnormal structures as damage and activate repair responses. Because this process itself carries the possibility of errors, it can ultimately contribute to the accumulation of mutations and chromosomal instability.
Diseases associated with DNA triplexes
Representative diseases associated with DNA triplex structures include autosomal dominant polycystic kidney disease (ADPKD) and Friedreich’s ataxia (FRDA). [1] ADPKD is a genetic disorder caused mainly by mutations in the PKD1 or PKD2 gene. Among these, triplex formation within polypurine–polypyrimidine repeat sequences (Pu•Py tracts) in the PKD1 gene has been proposed to play an important role in increasing replication instability and promoting the occurrence of gene mutations. As replication forks repeatedly stall and restart in this process, DNA breaks and errors during repair can accumulate, leading to pathogenic mutations.
Friedreich’s ataxia (FRDA) is an autosomal recessive hereditary neurodegenerative disorder that affects the peripheral nerves, spinal cord, and cerebellum. FRDA is caused by mutations in the FXN gene, which encodes the mitochondrial protein frataxin. Abnormal expansion of GAA repeats in the first intron of the FXN gene is known to form non-canonical structures such as triplexes or R-loops, interfering with transcription and consequently reducing expression of the frataxin protein. In this case, the significance of the triplex goes beyond simply promoting mutations: it can directly contribute to the suppression of gene expression.
DNA triplexes as therapeutic tools
The ability of triplex structures to inhibit transcription can, in turn, be exploited for therapeutic purposes. A TFO (triplex-forming oligonucleotide), a third strand designed to bind to a specific DNA sequence and artificially form a triplex, can selectively suppress the expression of a particular gene. This approach has developed into an antigene-based therapeutic strategy. Unlike the targeted gene-correction techniques we will look at next, this strategy does not cut a gene and correct its sequence. Instead, it functionally switches the gene off by blocking transcription itself at the DNA level. In an antigene strategy, a TFO binds to the promoter or transcriptional regulatory region of a target gene and forms a triplex, physically interfering with transcription factor binding or obstructing the progression of RNA polymerase. Even after transcription has already begun, a TFO can block continued elongation by RNA polymerase and thereby suppress mRNA production itself.
This strategy can target genes and proteins that play key roles in the survival and proliferation of cancer cells. For example, TFOs have been directed to the promoters of genes such as the oncogene c-MYC, which strongly promotes cell proliferation, and the gene encoding high mobility group box 1 (HMGB1), a transcription-related protein whose overexpression has been associated with cancer. By suppressing transcription of these genes, TFOs were shown to inhibit cell proliferation. Survivin, an anti-apoptotic protein, is barely detectable in most normal cells but is expressed in many cancer cells. In human lung cancer cells, a TFO targeting survivin was reported to reduce its expression, inhibit cell proliferation, and induce apoptosis. [2] In another study, delivery of a TFO targeting the human epidermal growth factor receptor 2 (HER2) gene region to HER2-overexpressing breast cancer cells markedly inhibited cell growth and induced apoptosis even in the absence of p53. Interestingly, HER2 gene expression itself was not significantly reduced in this study, suggesting that a DNA damage response triggered by triplex formation at the target site, rather than simple transcriptional suppression, played an important role. [3]
However, this strategy has not yet come into widespread clinical use. Because stable triplex formation by TFOs depends on specific purine-rich repetitive sequences, the range of available targets is limited, and delivering TFOs stably into the cell nucleus remains a difficult challenge. Transcriptional inhibition may also not be completely specific, potentially reducing efficiency through off-target effects, while distortion of DNA structure may trigger unintended DNA damage responses and increase cytotoxicity. Although nanoparticle delivery systems and chemical modifications such as peptide nucleic acids (PNAs) are being developed, efficient intracellular delivery of TFOs still remains an important challenge.
TFOs and targeted gene-correction strategies
TFOs can also be used in targeted gene-correction strategies. Early approaches to targeted gene correction introduced externally designed donor DNA into cells to induce homologous recombination, but their efficiency was very low because the donor DNA often failed to find the target site. Targeting strategies using TFOs were proposed to address this problem. Because TFOs bind specific DNA sequences through the major groove with relatively high sequence specificity, they can bind precisely to desired sequences, giving them considerable potential as gene-targeting agents for personalized therapies. Another important advantage is that sequences that can be targeted by TFOs, known as triplex target sites (TTSs), are abundant in mammalian genomes. [4] A triplex formed when a TFO binds to its target site can be recognized as a structural distortion of DNA and stimulate DNA repair systems, including nucleotide excision repair (NER). These repair responses can increase the likelihood of recombination and gene conversion around the target site. The aim of this strategy is to take advantage of this response to deliver the desired normal sequence information to the target DNA. Going one step further, researchers proposed physically linking donor DNA containing the “correct answer sequence,” which could serve as a template during homologous recombination, directly to the TFO. The effectiveness of this “tethered donor-TFO” strategy, which combines target recognition and sequence-information delivery in a single molecule, was later tested experimentally. [5] Researchers introduced the supF reporter gene, a useful tool that makes it possible to measure whether gene correction has occurred, into cells and compared the frequency of gene correction when TFO alone, donor DNA alone, or the two physically linked together were introduced. The results showed that gene conversion efficiency increased when the donor DNA was physically tethered near the target site.
The strategy was taken even further by linking TFOs to photo-reactive DNA crosslinking agents such as psoralen, with the aim of generating a stronger DNA damage signal and thereby activating the cell’s DNA repair systems more strongly. [6] It is a rather unusual strategy: cause stronger DNA damage in order to trigger a stronger repair response. Psoralen is activated by ultraviolet light and forms covalent bonds with certain DNA bases, particularly thymine. It first inserts itself between the two strands of the DNA duplex and, upon exposure to light, can form a crosslink connecting the two strands. This is a severe form of damage that effectively ties the two DNA strands together with covalent bonds. Unlike a simple structural distortion, such damage physically blocks DNA replication and transcription, so from the cell’s point of view it must be removed, triggering a strong repair response. In particular, several DNA repair pathways, including homologous recombination, can be recruited during the processing of these crosslinks. Although the psoralen-TFO strategy can further improve gene-correction efficiency, crosslinks can impose a substantial burden on cells and increase the possibility of side effects such as nonspecific DNA damage and impaired genome stability, limiting their practical therapeutic application.
These targeted gene therapies based on “target-site recognition” and the induction of strong DNA damage eventually reached a major turning point with the arrival of the CRISPR-Cas9 system. CRISPR-Cas9 can locate a specific DNA sequence, introduce a double-strand break directly at the desired site, and then make use of the cell’s DNA repair systems to modify the genome. This made targeted gene correction far more efficient and broadly applicable. As a result, the importance of TFOs as gene-correction tools has declined considerably. Even so, TFOs remain unique research tools and potential therapeutic strategies because they can directly recognize specific DNA sequences without requiring an external protein enzyme, create a structural change in the form of a triplex, and thereby regulate transcription or induce DNA repair responses.
[References]
[1] Non-canonical DNA structures: Diversity and disease association
[2] Triplex-forming oligodeoxynucleotides targeting survivin inhibit proliferation and induce apoptosis of human lung carcinoma cells
[3] Nanoparticle delivery of TFOs is a novel targeted therapy for HER2 amplified breast cancer
[4] Triplex technology in studies of DNA damage, DNA repair, and mutagenesis
[5] Targeted Correction of an Episomal Gene in Mammalian Cells by a Short DNA Fragment Tethered to a Triplex-forming Oligonucleotide
[6] Repair of DNA lesions associated with triplex-forming oligonucleotides
[Image sources]
Image 11-1. Intrastrand triplex DNA repeats in bacteria: a source of genomic instability — CC BY 4.0
1-2. Insights into the Molecular Structure, Stability, and Biological Significance of Non-Canonical DNA Forms, with a Focus on G-Quadruplexes and i-Motifs — CC BY 4.0
Image 2
2-1. Three- and four-stranded nucleic acid structures and their ligands — CC BY-NC 3.0
2-2. Triplex-forming oligonucleotides as an anti-gene technique for cancer therapy — CC BY 4.0
Image 3
3-1. Three- and four-stranded nucleic acid structures and their ligands — CC BY-NC 3.0
Image 4
4-1. Structures and stability of simple DNA repeats from bacteria — CC BY 4.0
Image 5
5-1. Non-canonical DNA structures: Diversity and disease association — CC BY 4.0
Image 6
6-1. Non-canonical DNA structures: Diversity and disease association — CC BY 4.0
1-2. Insights into the Molecular Structure, Stability, and Biological Significance of Non-Canonical DNA Forms, with a Focus on G-Quadruplexes and i-Motifs — CC BY 4.0
Image 2
2-1. Three- and four-stranded nucleic acid structures and their ligands — CC BY-NC 3.0
2-2. Triplex-forming oligonucleotides as an anti-gene technique for cancer therapy — CC BY 4.0
Image 3
3-1. Three- and four-stranded nucleic acid structures and their ligands — CC BY-NC 3.0
Image 4
4-1. Structures and stability of simple DNA repeats from bacteria — CC BY 4.0
Image 5
5-1. Non-canonical DNA structures: Diversity and disease association — CC BY 4.0
Image 6
6-1. Non-canonical DNA structures: Diversity and disease association — CC BY 4.0





