In the previous article, we looked at how DNA exists inside the cell nucleus. DNA is first wrapped around histone proteins, and the packaged DNA then bends into loops, allowing regions that are brought close together in space to interact more frequently. On a larger scale, some regions are relatively densely packed, while others are more loosely organized. Each chromosome also tends to occupy its own territory within the nucleus. In other words, DNA does not exist as a long, fully extended chain. Instead, it forms an extremely complex three-dimensional structure within the tiny space of the nucleus, and we saw that this spatial organization—and the accessibility it creates—also affects transcription during gene expression. Having taken a step back to look at the overall hierarchical organization of DNA, let us now zoom all the way in and turn our attention to its smallest building blocks.
Nucleotides: The Basic Units of DNA
DNA is more than just a molecule that passes genetic information on to the next generation. It contains the information needed to produce proteins that directly carry out the activities of life, as well as functional RNAs that assist in protein synthesis. It also contains regulatory information that helps determine when, where, and to what extent this information can be expressed. In addition, DNA carries information encoding proteins that repair DNA when it is damaged, information involved in regulating cell proliferation and differentiation, and important regulatory information involved in suppressing the expression of genes associated with cancer development.
This enormous amount of critical information must remain stably stored across countless generations and be copied with high fidelity, with errors kept to a minimum. Yet alongside this stability and accuracy, a carefully limited degree of variation must also be possible—enough to allow evolution to occur. The molecules must also be capable of forming long polymer chains and of being efficiently packaged so that all of this material can fit within the extremely limited space of the cell nucleus. The chemical structures that most effectively satisfy all of these important requirements at the same time are nucleic acids, such as DNA and RNA, whose basic units are nucleotides.
A nucleotide, the basic unit that makes up DNA, consists of a nitrogenous base, a deoxyribose sugar, and phosphate, and there is a clear division of labor among these components. The four nitrogenous bases carry the information, while the sugars and phosphates attached to them form a repeating backbone that helps provide structural stability to the DNA helix. Information is encoded simply by arranging four types of bases in different sequences, while the basic structure itself does not depend on changes in that sequence. The backbone remains repetitive and uniform. This separation of information from structure was likely one of the key features that allowed DNA to become a stable medium for storing information. DNA is chemically quite stable, but it is not a molecule that never changes. Replication maintains a high degree of accuracy, yet rare changes in the base sequence can become mutations, which in turn have provided material for evolution. This coexistence of stability and variability is one of the features that makes DNA so well suited to storing genetic information and transmitting it across generations.
DNA encodes information using only four types of bases. Each base pairs complementarily with only a specific partner, and because of this simple rule of complementarity, one strand of DNA alone can be used to accurately reconstruct the complementary strand on the other side. Complementarity is what makes accurate replication possible. One way to think about why proteins, despite being capable of far greater chemical diversity than nucleic acids, did not become the genetic material is that they lack a simple rule by which one chain can serve as a template for directly producing a complementary second chain. Nucleic acids follow the physicochemical rule of base complementarity: with A pairing with T and G pairing with C, the sequence of one strand automatically specifies the other, making template-based replication possible. Proteins, in contrast, have no such one-to-one complementary pairing rule, and protein-specific features such as the distinct side chains of 20 different amino acids and protein folding are far from simple. In DNA, which uses only four types of bases and forms a double helix, pairing bases of different sizes also makes it easier to maintain structural stability by keeping the width of the helix relatively constant.
Definition of Nucleic Acids
The name “nucleic acid” reflects both where these molecules were first discovered and their chemical properties. In 1869, while studying pus rich in white blood cells, the Swiss biochemist Friedrich Miescher discovered a new substance whose chemical properties differed from those of the proteins known at the time. He found that this phosphorus-rich material was concentrated in the cell nucleus and called it “nuclein,” referring to its origin in the nucleus. It was later discovered that the molecule contained large amounts of phosphate, which contributes negatively charged acidic groups and gives the molecule an overall acidic character. It therefore came to be called “nucleic acid,” meaning an acidic polymeric substance discovered in the nucleus. As research progressed in the early twentieth century, scientists found that nucleic acids could be divided into two structurally similar types that differ in one of their components—the type of sugar they contain. Nucleic acid containing ribose is ribonucleic acid (RNA), whereas nucleic acid containing deoxyribose is deoxyribonucleic acid (DNA). Despite the name, nucleic acids are not actually found only in the nucleus, but the name has remained in use.
Perhaps the most precise way to define a nucleic acid is as a polymer made of repeatedly linked nucleotides. A short chain consisting of several to a few dozen nucleotide monomers joined by phosphodiester bonds is called an oligonucleotide. A longer chain is called a polynucleotide, and when such long chains form functional macromolecules such as RNA or DNA, they are referred to as nucleic acids.
Nucleotides
Let us take a closer look at the structure of the nucleotide monomer, the basic unit of nucleic acids. A nucleotide is a molecule in which a sugar is linked to a nitrogenous base and one or more phosphate groups. It can also be defined as a nucleoside—a structure consisting of a nitrogenous base linked to a sugar—with one or more additional phosphate groups attached. This stepwise expansion is illustrated in the figure at the end of this article. Nucleotides can be further classified according to the number of phosphate groups attached to the nucleoside.
A nucleoside is formed when a sugar and a nitrogenous base are joined by an N-glycosidic bond. When phosphate is then linked to the sugar of the nucleoside at its 5′ OH through an ester bond—in this case, a phosphoester bond—the molecule becomes a nucleotide. (To distinguish the carbon numbers of the sugar from those of the nitrogenous base, the sugar carbons are labeled with a prime symbol, ′.) As an individual monomer, a nucleotide can contain from one to three phosphate groups. When DNA polymerase or RNA polymerase synthesizes a nucleic acid, it extends the chain using these free nucleotide molecules as building blocks. The substrates used by the polymerases are always nucleoside triphosphates: RNA polymerase uses NTPs (ATP, GTP, CTP, and UTP), whereas DNA polymerase uses dNTPs (dATP, dGTP, dCTP, and dTTP), where the “d” indicates the deoxy sugar used in DNA. The reason is that energy is required for a polymerase to add a new nucleotide to the existing chain. To put it very simply, a triphosphate contains several negatively charged phosphate groups packed close together, creating electrostatic repulsion and a relatively high-energy state. In a sense, the nucleotide arrives carrying the energy needed to attach itself to the nucleic acid chain. When the new nucleotide is incorporated into the chain, two phosphate groups leave as pyrophosphate (PPi), and this reaction helps make formation of the new bond possible.
Because a phosphate already forms one phosphoester bond with the sugar of a nucleotide and then forms another ester bond with the 3′ OH of the neighboring sugar, a single phosphate ultimately forms two ester bonds with two sugars. This is called a phosphodiester bond. The fact that a new nucleotide is added to the 3′ OH of the existing sugar is an extremely important feature, because it determines the directionality of nucleic acid synthesis, so it is worth remembering. During ordinary DNA and RNA synthesis, a polymerase uses the 3′ OH at the end of the existing chain to form a new bond with the phosphate of the incoming nucleotide. The nucleic acid chain therefore grows in the 5′→3′ direction.
Structure of Nucleic Acids
When phosphate groups link neighboring sugars through covalent phosphodiester bonds, they form strong and robust sugar-phosphate backbones along the outside of the nucleic acid structure. The nitrogenous base pairs on the inside, however, are connected by hydrogen bonds, which are weaker and more reversible than covalent bonds. They are therefore not permanently locked in place, allowing the molecule to combine stability with flexibility. Having these two different types of interactions provides two levels of bonding strength, allowing the structure to remain stable while also responding flexibly to thermal fluctuations.
The nitrogenous bases on the inside are nitrogen-containing aromatic ring compounds. Based on their structures, they are divided into purines, which contain two rings, and pyrimidines, which contain a single ring. A purine, consisting of a fused six-membered and five-membered ring, is relatively larger than a pyrimidine, which has a single six-membered ring. Pairing bases from these two size classes therefore keeps the distance between the two sides similar and helps maintain a consistent width of the helix. Adenine (A) and guanine (G) are purines, while cytosine (C), thymine (T), and uracil (U), which replaces T in RNA, are pyrimidines. We will examine the nitrogenous bases in much greater detail in a separate article.
The Roles of Nucleosides and Phosphat
The fact that adding one or more phosphate groups to a nucleoside turns it into a nucleotide means that this seemingly small change—the addition of phosphate—marks a fundamental shift in the biological function of the molecule. So what changes when phosphate is added? To begin with, although it is not chemically impossible for nucleoside units to link together, they are not well suited to efficient polymerization under biological conditions. Once a phosphate group is attached at the 5′ position of the pentose sugar, however, the molecule becomes a nucleotide, and in nucleic acid synthesis it is the triphosphate form of the nucleotide that is used as the building block. When this nucleoside triphosphate reacts with the 3′ OH of the existing chain, a 3′–5′ phosphodiester bond forms between the two nucleotides and the chain becomes longer. Several nucleotides polymerized in this way form an oligonucleotide, and when many more nucleotides are polymerized into a long chain, the result is a nucleic acid such as DNA or RNA. By covalently linking one sugar to the next, nucleotide phosphates form the backbone that allows a long nucleic acid chain to remain stable. At the same time, the strongly negatively charged phosphate groups give the nucleic acid as a whole a negative charge, which also plays an important role in allowing it to exist stably in polar, aqueous environments. In short, phosphate plays a central role in allowing nucleic acids to become long, stable polymers.
Why, then, is it difficult for nucleosides to polymerize with one another without phosphate? As we saw in detail in the articles on carbohydrates, the many -OH groups of monosaccharides can, in principle, form glycosidic bonds with OH groups on other sugars. Nucleosides also contain a sugar, but carbon 1 of that sugar—the anomeric carbon—is already joined to the nitrogenous base by an N-glycosidic bond. The remaining OH groups at the 2′, 3′, and 5′ positions do not readily react directly with one another under biological conditions to produce a long, regular chain. In nucleic acids, phosphate serves as a bridge connecting two sugars, making it possible to form long, stable chains. In addition, because sugars and phosphates are repeatedly linked in a consistent pattern, the nucleic acid chain acquires regularity and a defined 5′-to-3′ directionality. The significance of this directionality and regularity in nucleic acids is profound.
Another important change that occurs when phosphate is added to a nucleoside has to do with energy. As we saw above, the nucleotides used for nucleic acid synthesis enter the reaction in their triphosphate form. To think about it in a highly simplified way, when several negatively charged phosphate groups are packed close together, they repel one another. The molecule is therefore in a relatively high-energy state, somewhat like a tightly compressed spring. When a new nucleotide is added to a nucleic acid chain, two phosphate groups leave as pyrophosphate (PPi), bringing the system to a more stable state, and the difference in energy between these states is used in forming the new bond. Of course, this simple picture does not explain everything that is actually happening, but it is a useful analogy for understanding why nucleoside triphosphates can provide the energy required to form a nucleic acid chain. In a sense, a nucleotide arrives carrying not only the material that will become part of the nucleic acid, but also the energy needed to connect itself to the chain. The familiar molecule ATP likewise uses its phosphate bonds to transfer energy for many cellular reactions.
3′→5′ Directionality
When the OH group on the 3′ carbon of one nucleotide sugar is connected through a 3′–5′ phosphodiester bond to the phosphate attached to the 5′ carbon of the next nucleotide, the nucleic acid chain naturally acquires a 5′ end at one side and a 3′ end at the other. Structurally, the chain therefore has a defined direction. During ordinary DNA and RNA synthesis, each new nucleotide is added to the 3′ OH at the end of the existing chain. This directionality plays a central role in allowing nucleic acids to store and replicate genetic information reliably. Because the sugar-phosphate backbone of a nucleic acid repeats in the same regular pattern, DNA and RNA can form orderly linear chains, and this regularity allows the enzymes involved in replication and transcription to move along the chain in a consistent direction while producing a new strand. For example, an enzyme replicating DNA reads the existing template strand in the 3′→5′ direction while synthesizing the new strand in the 5′→3′ direction. Both DNA polymerase and RNA polymerase move along a template strand while adding new nucleotides to the 3′ OH of the growing chain. Giving the nucleic acid backbone a defined direction therefore provides an important structural basis that allows the enzymes involved in replication and transcription to read their templates and synthesize new chains in an orderly and consistent manner. Thus, the regular 3′–5′ linkage pattern provides the structural foundation that gives nucleic acids an ordered directionality and allows replication and transcription to proceed in a consistent way.
The rule that nucleotides must be added to a 3′ OH creates a problem for DNA polymerase when it tries to add the very first nucleotide and initiate polymerization. DNA polymerase cannot initiate polymerization on its own, because adding a new nucleotide requires a pre-existing 3′ OH. To solve this problem, an enzyme called primase produces a short RNA primer. DNA polymerase can then begin synthesis by adding nucleotides to the 3′ OH at the end of that primer.
Going back a little, however, why was a 2′–5′ linkage not selected instead, given the 2′, 3′, and 5′ OH groups of the sugar? Because RNA also has a 2′ OH, forming a 2′–5′ phosphodiester bond is not chemically impossible. In fact, 2′–5′ linkages are found in certain specialized RNAs and biological reactions. However, a 2′–5′ linkage creates a spatial arrangement different from that of the usual 3′–5′ linkage, and when such linkages are incorporated into an RNA chain, they tend to reduce the stability of the double helix. Experimental studies have shown that 2′–5′ linkages alter the local structure of an RNA duplex and that thermal stability decreases as the number of these linkages increases.[1] It is difficult to explain with any single reason why the 3′–5′ linkage became the standard form during evolution, but the structural stability and regularity of a backbone built from 3′–5′ linkages can be considered one important factor.
The absence of a 2′ OH in DNA, meanwhile, is a somewhat different issue. The 2′ OH in RNA can attack a neighboring phosphodiester bond and cause cleavage of the RNA chain, making RNA more susceptible to hydrolysis than DNA. Deoxyribose in DNA lacks this 2′ OH and is therefore chemically much more stable, an important property that makes DNA well suited to the long-term storage of genetic information. In the next article, we will look more closely at the 2′ OH of RNA, its chemical instability, and how these features relate to the functional differences between RNA and DNA.
Nucleoside Analogue Drugs
What role, then, does the nucleoside—the molecule that may appear to be a precursor to a nucleotide—play? If we turn the explanation above around, nucleosides, which lack phosphate and are close to electrically neutral, can cross the cell membrane using specific membrane transporters. Nucleotides, in contrast, have bulky, charged phosphate groups and therefore have much more difficulty crossing the membrane. Bases and nucleosides produced by the breakdown of nucleic acids and nucleotides can be recovered by the cell and converted back into nucleotides. These recovered materials can then be reused to synthesize the nucleotides needed for DNA replication or RNA transcription. This recycling route is known as the salvage pathway.
This principle is also used as a drug strategy by employing analogues that mimic the structure of nucleosides to interfere with nucleic acid synthesis. Viruses must replicate their own DNA or RNA in order to multiply, and rapidly proliferating cancer cells likewise need to synthesize large amounts of new nucleic acids. Nucleoside analogue drugs take advantage of this requirement. After entering a cell, many of these drugs are phosphorylated by host-cell or viral enzymes and activated into nucleotide analogue forms. Once activated, the analogues can compete with normal nucleotides or become incorporated into a nucleic acid chain, interfering with polymerase activity or normal nucleic acid synthesis. Some drugs lack a 3′ OH and can directly terminate chain elongation, while others interfere with polymerase progression in different ways. Put simply, it is a clever strategy that introduces fake building blocks resembling the real ones, thereby disrupting the ability of viruses or cancer cells to synthesize nucleic acids. Acyclovir, used to treat herpesvirus infections; zidovudine, used to treat HIV infection; and remdesivir, used to treat RNA virus infections, are representative drugs that exploit this general principle to interfere with viral nucleic acid synthesis. However, the ways in which these drugs are activated and the specific mechanisms by which they interfere with polymerases differ.
The Multiple Roles of Nucleotides
Nucleotides, however, are far more than simply the building blocks of nucleic acids that store genetic information. They are also extremely important metabolic molecules that perform a wide range of essential functions in cellular life. ATP, the familiar energy currency of the cell; GTP, which acts as a molecular switch for G proteins in cell signaling; cAMP and cGMP, which serve as intracellular second messengers and help coordinate physiological responses; redox cofactors such as NAD+, NADP+, and FAD, which play essential roles in transferring electrons during processes such as glycolysis, the TCA cycle, and fatty acid synthesis; CoA (coenzyme A), which carries acetyl groups and fatty acids and plays important roles in energy metabolism and lipid synthesis; and UDP-glucose, an intermediate that converts sugars into activated chemical units that can be transferred—all of these are either nucleotides themselves or molecules containing structures derived from nucleotides. Nucleotides and their analogues also provide an important foundation for drug design in modern medicine. Thus, beyond information storage, nucleotides are essential biomolecules used throughout living systems for energy transfer, cell signaling, metabolic reactions, and medical applications.
Finally, the figure below illustrates the stepwise progression in which the addition of a sugar to a nitrogenous base forms a nucleoside, and the addition of one or more phosphate groups to the nucleoside forms a nucleotide. Depending on the number of phosphate groups attached, a nucleotide can exist in monophosphate, diphosphate, or triphosphate form.
[Reference]
[1] Structural insights into the effects of 2′-5′ linkages on the RNA duplex


