Ramachandran plot: There are limits to the structures a protein can adopt
In the previous article, we looked at how the Ramachandran plot developed historically and how it is still used today. Above all, Ramachandran showed, by calculating the possible φ and ψ combinations one by one, that the protein backbone is in fact restricted to a surprisingly limited range of these two dihedral angles because of steric hindrance. In other words, only a very small fraction of all conceivable structures are physically possible for a protein in the first place, and because many sterically impossible structures are excluded from the beginning, this provided an important clue that a protein does not randomly explore every possible three-dimensional structure when it folds. But this alone is nowhere near enough to explain protein folding completely. Even if we consider only the physically allowed structures that do not cause steric clashes, the number of possible conformations is still beyond imagination. So why do real proteins, among all these possibilities, accurately find and fold into their own unique native structure, and how can they find that structure in such an incredibly short time? It was Christian Anfinsen who provided an important clue to this question.
Anfinsen’s discovery: The folding blueprint is contained in the amino acid sequence itself
In 1961, Anfinsen made a landmark discovery in protein-folding research through experiments using the nuclease ribonuclease A (RNase A). RNase A is an enzyme made up of 124 amino acids and contains four disulfide bonds within its structure. To completely unfold the already folded protein, he first used urea to disrupt non-covalent interactions within the protein, including hydrogen bonds and hydrophobic interactions, and then treated it with β-mercaptoethanol to break all four disulfide bonds as well. As a result, the enzyme protein became completely unfolded and denatured, and it lost all of its enzymatic activity. What happened in the next experiment was even more remarkable. When urea and β-mercaptoethanol were removed from the completely ‘wrecked(?)’ RNase A, the unfolded protein began to fold itself back into its original three-dimensional structure even though there was no external assistance or energy supply. It did not simply recover a roughly similar shape. All four disulfide bonds were re-formed in exactly their original positions, and eventually the original enzymatic activity was restored almost completely. From this experiment, Anfinsen drew an important conclusion. The information required for a protein to fold into a particular structure is not outside the protein, but is already contained in the amino acid sequence itself, its primary structure. This idea is known as Anfinsen’s dogma or the thermodynamic hypothesis, and he was awarded the 1972 Nobel Prize in Chemistry for this work.
Why is it called a thermodynamic hypothesis?: Gibbs free energy and the global minimum
So why was Anfinsen’s idea called a thermodynamic hypothesis? What particularly caught Anfinsen’s attention in these experiments was that, no matter how many times the protein was unfolded and then left alone again, it eventually returned to the same structure. Although an unfolded protein can adopt many different conformations, it folded by choosing almost the same final structure each time. Anfinsen explained the reason thermodynamically. In thermodynamics, spontaneous changes under given conditions proceed in the direction of decreasing free energy and eventually reach the most stable equilibrium state. Anfinsen, in effect, rediscovered the concept of Gibbs free energy—a thermodynamic concept established in the nineteenth century, more than a century earlier, to explain steam engines and the direction of chemical reactions—as a biochemical solution to the mystery of proteins, one of the central materials of life. He thought that, like other forms of matter, a protein naturally folds towards the stable state with the lowest free energy among its many possible conformations under a given environment. That stable free-energy region corresponds to the protein’s native structure. Here, native structure means the characteristic three-dimensional structure in which a particular protein exists most stably under physiological conditions.
Being in the state of lowest free energy means more than simply being stable; it means occupying the thermodynamically most favourable position among the possible states. This is commonly described, quite literally, as the global minimum. The native structure of a protein is therefore not formed by chance, but corresponds to the lowest and most stable region—the global minimum—of the free-energy landscape created by its amino acid sequence. This idea later became an important starting point for modern theories of protein folding, including energy-landscape theory and the folding funnel.
Levinthal’s paradox: A protein cannot fold by randomly searching through structures
Anfinsen had proposed that the destination of protein folding was a stable state of low free energy, but this immediately raised another important question. Among the enormous number of possible structures, how can a protein find that destination so quickly? This question was posed in a particularly dramatic way by the famous Levinthal’s paradox, named after Cyrus Levinthal.
In 1969, Cyrus Levinthal posed this question about protein folding using a very simple calculation. He roughly estimated how long it would take a polypeptide chain to randomly examine every possible conformation one by one, choose the most energetically stable one, and finally fold into its native structure. Imagine, for example, a small protein chain made of 100 amino acids. Excluding the first and last residues, each amino acid residue has two dihedral angles, φ and ψ, giving roughly 200 rotational degrees of freedom. Now let us make the highly simplified assumption that each dihedral angle can occupy only three stable states. The number of possible three-dimensional conformations would then be as large as 3²⁰⁰. Even this gives the astonishing figure of about 10⁹⁵. If we further assume that checking a single conformation takes roughly 10⁻¹³ seconds (0.1 ps), then exploring every possible conformation one by one would require 3²⁰⁰ / 10¹³ seconds. That is about 10⁷⁵ years. It may be difficult even to grasp how enormous this number is. Considering that the age of the universe is about 13.8 billion years, or 1.38 × 10¹⁰ years, is this not vastly longer than the age of the universe itself? In practical terms, it is impossible.
But what happens in reality? In living systems, most proteins take only a few milliseconds to a few seconds to fold into their complete structures. If the many proteins in our bodies took anything like that enormous amount of time to form, how could living organisms ever function or even exist? Levinthal’s paradox points to this extraordinary gap between theoretical calculation and reality. Levinthal was not arguing that proteins actually explore every possible structure. On the contrary, he was showing that if they did, reality could not be explained.
In other words, the central point of Levinthal’s argument is that a protein does not randomly explore every possible structure one by one. The word paradox originally implies a contradiction, and here the contradiction is really pointing out that our assumed model of folding—the idea that a protein folds through a random search—is itself wrong. There must therefore be some other principle that gives protein folding a certain direction from the beginning and allows a protein to find its native structure far more efficiently.
A review and reflection on theories of protein folding
Before we move on to energy-landscape theory, let us turn the clock a little further forward and briefly review the questions through which protein-folding research developed historically. In 1951, Linus Pauling explained that the peptide bond has a planar character because of resonance and, on that basis, accurately predicted the major secondary structures known as the α-helix and β-sheet. In 1963, Ramachandran went one step further and showed that the combinations of φ and ψ available to the protein backbone are themselves highly restricted because of steric hindrance. He was the first to demonstrate that proteins cannot adopt every conceivable structure in the first place, but are confined to a physically allowed range of conformations. This then led researchers to ask why, among the many allowed structures, a protein chooses one particular structure. In 1972, through his experiments with RNase A, Anfinsen proposed the thermodynamic hypothesis: all the information required for protein folding is already contained in the amino acid sequence, and under a given environment a protein folds towards its stable native structure, where free energy is lowest. This, in turn, raised yet another question. How does a protein find its native structure so quickly among the enormous number of possible conformations? Levinthal posed this question and showed through calculation that, if a protein randomly explored every structure one by one, it would take longer than the age of the universe. Given that real proteins in the body fold in the blink of an eye, he argued that they clearly cannot be exploring every possible structure.
After the questions raised by Anfinsen and Levinthal, various theories were proposed to explain protein folding. Until the 1980s and 1990s, a common view in the field was that proteins folded sequentially along relatively fixed folding pathways. In other words, certain structures were thought to form first, providing the basis for subsequent structures to develop step by step. Gradually, however, this view began to be questioned. Do all proteins really fold by following one predetermined route? In 1997, José Onuchic, Peter Wolynes, Luthey-Schulten and others published a paper entitled “Theory of Protein Folding: The Energy Landscape Perspective”, systematically presenting a new way of looking at protein folding. Unlike the earlier idea that a protein follows one fixed, predetermined pathway, they made the striking proposal that a protein can use not one but many possible pathways and still ultimately reach its low-free-energy native structure. Rather than behaving like someone searching for treasure by randomly examining every possible structure one by one until finding a single path, it is more like a stone placed on top of a mountain, rolling down along many possible steep slopes until it eventually reaches the lowest point. A protein continually moves back and forth among countless microscopic states, yet across the overall free-energy landscape it can gradually move towards more stable regions.
Thermodynamics and statistical physics: Statistics of the invisible microscopic world
To understand this new paradigm more deeply, we need to turn the clock back by about 150 years, to the time when thermodynamics and statistical physics were born. Most of the scientists who proposed this new theory actually came from backgrounds in physics rather than biology, and the key concepts needed to understand energy-landscape theory—thermodynamics, statistical mechanics, enthalpy, entropy and free energy—were not originally developed to explain proteins at all. These ideas were first developed by physicists trying to understand the macroscopic behaviour of matter and energy, and only later were they extended to biological phenomena such as proteins. So, personally, this became an opportunity for me to study thermodynamics in depth—something I had long tried to avoid while keeping my eyes half closed and looking the other way whenever possible. To trace the historical background in which the laws of thermodynamics emerged and how they developed, I placed immediately before this article two articles covering everything from Carnot’s engine, Clausius and Boltzmann to, finally, Gibbs free energy. Please refer to those articles for the details. In particular, they explain in depth what entropy is and why nature appears to move in the direction of increasing entropy, so I recommend reading them and then returning to this article to continue the discussion. In this article, however, I will keep the thermodynamics very brief so that we do not lose the original thread of protein folding, and then look at how these concepts connect back to protein folding.
Physics is the study of natural phenomena around us using mathematics and physical laws. In the late eighteenth century, the Industrial Revolution brought dramatic changes as machines powered by heat began to replace human labour. Coal was burned to boil water, and the expanding steam pushed pistons to perform work, but even the best steam engines could not convert all the heat energy supplied to them into useful work. Research aimed at reducing heat loss and improving thermal efficiency therefore became a major concern for physicists. Thermodynamics is the field that studies relationships among macroscopic physical quantities that we can observe, such as temperature, volume and mass, and it actively investigated questions such as how heat moves, how energy is converted, why some processes occur naturally and why others do not. At the time, the existence of atoms had not yet been proven. Thermodynamic research could describe macroscopic phenomena and outcomes such as the expansion of gases or the melting of ice, but it could not explain the more fundamental question of why nature always seems to move in a particular direction—for example, why, when hot and cold water are mixed, heat always flows from the hotter side to the colder one. Around this time, research into the motion of gases became increasingly active, and attempts began to explain gas molecules from a microscopic perspective. The emerging idea was that the macroscopic states we observe are actually statistical outcomes produced by the states of countless invisible atoms and molecules. Imagine that a small box contains several gas molecules. Each molecule can occupy countless positions within the box and can also move at different speeds. If we specify where every molecule is and how fast it is moving, we define one microstate. For example, pressure can vary depending on how and at what speeds the molecules strike the walls. In other words, the temperature or pressure we observe can be understood as the statistical result of these enormous numbers of microstates. A macroscopic state such as pressure or temperature therefore depends on an unimaginably large number of microscopic states, including the arrangements and motions of the molecules within it.
Gibbs free energy
Eventually, physicists expanded thermodynamics into a vast theoretical framework by statistically calculating the probabilistic behaviour of these countless microstates and using it to explain which processes occur spontaneously in nature. These are the thermodynamic laws of the macroscopic world that we can observe. And these laws of thermodynamics can also be applied directly to understanding complex biological molecules such as proteins. In particular, Gibbs free energy—perhaps one of the most elegant ideas offered by thermodynamics—is a key concept for explaining the directionality of protein folding proposed by Anfinsen and Levinthal. Inside the cells of the human body, temperature and pressure are essentially constant, and under these conditions molecular ensembles spontaneously move towards states of lower free energy (G). Free energy is an indicator of how stable a state is: a state with high free energy is unstable, while lower free energy corresponds to greater structural stability. Naturally, nature favours more stable states, and so changes tend to proceed in the direction of lower free energy. Let us understand Gibbs free energy in the context of protein folding. The equation is shown below, and free energy is determined by a fierce tug-of-war between two physical quantities: enthalpy (H) and entropy (S). Δ means the amount of change.
\(\Delta G = \Delta H - T\Delta S\)
Let us first look at the basic idea. Enthalpy ($H$) comes from the Greek words En (‘within’) and Thalpos (‘heat’) and represents changes in the energy, or heat, exchanged by a substance. Interactions involved in protein folding, such as hydrogen bonds, disulfide bonds, salt bridges and van der Waals forces, release heat when they form, so as folding progresses, enthalpy decreases. (This is the opposite of breaking bonds, which requires an input of energy.)
Entropy ($S$) is a physical quantity related to the number of different microstates that molecules can occupy. Entropy is often described as a measure of disorder, but personally I think this interpretation can make entropy more confusing rather than easier to understand. I find it much clearer to think of entropy in terms of the number of microstates, or possible arrangements, that can produce a particular macroscopic state. An initially extended chain can occupy an enormous number of microstates. Because it has such a vast range of possibilities, its entropy is high. As folding progresses, however, the number of available microstates decreases, and once the chain becomes confined to a single native structure, its entropy becomes extremely low. In the equation, the change in entropy is multiplied by the absolute temperature. Therefore, as temperature rises, both the free-energy-lowering effect of an increase in entropy and, conversely, the free-energy-raising effect of a decrease in entropy become greater.
So how should we interpret this equation? When thinking about protein folding, we cannot simply consider movement towards the lowest energy. We must take into account both the enthalpy gained through molecular bonds and interactions within the protein and the overall change in entropy of the protein together with the surrounding water. As hydrogen bonds and other non-covalent interactions form during folding, the process becomes favourable in terms of enthalpy. At the same time, however, the freely moving polypeptide chain becomes increasingly restricted to a particular structure, reducing the number of microstates available to the protein—in other words, reducing its entropy. The competition between these two effects determines the free-energy value. There is, however, one very important factor that must not be forgotten: the role of water as the solvent in the aqueous environment of the body where protein folding takes place. This is extremely important. We often identify the hydrophobic interaction as one of the strongest driving factors in protein folding, and the reason is also related to entropy. When hydrophobic parts of a protein are exposed to water, the surrounding water molecules must give up some of their usual freedom and arrange themselves in particular ways. In other words, the entropy of the water decreases considerably. But when hydrophobic residues cluster together and become buried inside the protein, the surrounding water molecules are released and regain their freedom, increasing their entropy. From the protein’s point of view, burying hydrophobic residues causes further folding and reduces the protein’s own entropy. But if the increase in the entropy of the water is much greater than this decrease, then the total entropy of the protein-plus-water system actually increases. This is why the hydrophobic effect is such a central driving force in protein folding: according to the second law of thermodynamics, nature moves in the direction of increasing entropy. Protein folding therefore proceeds in the direction of lower free energy when all these factors—the protein, temperature, water and so on—are considered together.
A three-dimensional free-energy landscape and dialanine peptide
What exactly are the mountains and valleys we talk about in an energy landscape? And how can free energy actually be represented as a landscape? Gibbs free energy, calculated through the fierce tug-of-war between enthalpy and entropy, provides the final criterion that determines the spontaneous direction of movement of molecular ensembles in the constant-temperature environment of a cell, and it therefore becomes one axis of the energy landscape. To understand this idea, let us look at the free-energy surface (FES) of an alanine dipeptide. If structural variables such as a protein’s dihedral angles or atomic positions are placed on the x- and y-axes, and the Gibbs free-energy value associated with each structural coordinate is plotted as height on the z-axis, the result is a vast three-dimensional mountain range. As a note, the molecule used here is an actual dipeptide composed of two alanine residues, Ala–Ala, and is different from the peptide-mimicking dipeptide Ace–Ala–Nme used earlier by Ramachandran, so the two should not be confused.
How is this FES constructed? Even an alanine dipeptide, which can be regarded as one of the simplest protein models, has many different factors—or degrees of freedom—that can determine its three-dimensional structure. These include bond lengths, bond angles, the φ and ψ dihedral angles, rotations of methyl groups, side-chain movements, small vibrations and the positions of water molecules. Among all these variables, however, only the φ and ψ dihedral angles are retained explicitly, while the other degrees of freedom are statistically averaged at each φ, ψ combination and incorporated into the free-energy calculation. In other words, all the other degrees of freedom are averaged out, leaving only φ and ψ as coordinate axes and allowing us to create a two-dimensional map. φ and ψ are placed on the two axes, and, for example, if a state with φ = −60° and ψ = −40° is held fixed, all the other degrees of freedom continue to vary. Their microscopic contributions are averaged to obtain the free energy at that particular point, which is then plotted on the z-axis. A single point in the three-dimensional landscape therefore does not represent one individual structure, but the average free energy of the many microstates that share those coordinates. Put another way, whereas Ramachandran used the Ramachandran plot to show which combinations of φ and ψ are sterically allowed, the FES shows, even within those allowed regions, which φ, ψ combinations correspond to more stable states and which correspond to less stable ones, based on their free-energy values.
When free energy is represented in three dimensions like this, the result looks like a landscape of mountains and valleys. This is presumably where the term energy landscape comes from. Let us remember again that, thermodynamically, a system tends to change spontaneously in the direction of decreasing free energy and eventually reach its most stable equilibrium state. Deep valleys with low free energy are shown in blue and represent stable structures in which a protein that has moved down the slope can remain for a relatively long time. The red peaks rising between one valley and another represent unstable regions of high free energy—the free-energy barriers that the protein must cross in order to move from one state to another. Stable conformations in which the molecule can comfortably remain for longer periods because their free energy is low, such as the angles associated with α-helices or β-sheets in proteins, correspond to the blue valleys. Like water or a rock naturally moving from a high mountain ridge towards a valley, a protein moves in the direction of lower free energy and ultimately reaches the deepest valley, the native state.
This also shows us why protein folding cannot be explained simply as falling to the lowest point. Thermodynamics can tell us where the lowest valley lies, but how quickly an actual molecule can reach that valley is a different question. That is a problem of kinetics.
The folding-funnel model: Compressing an ultra-high-dimensional free-energy space
For a very simple molecule such as alanine dipeptide, an energy landscape can be represented relatively easily through actual calculations using only a single pair of dihedral angles, φ and ψ. But what about a real protein? The principle is of course the same. Unlike a dialanine peptide, however, a real protein may contain hundreds or even thousands of φ and ψ angles. A protein made of 100 amino acids, for example, has roughly 100 φ angles and 100 ψ angles. If we tried to represent them all as coordinate axes—φ1, φ2, φ3 ... φ100 and ψ1, ψ2, ψ3 ... ψ100—and calculate the free energy, the landscape would expand into a 200-dimensional space. In reality, once we also take into account many additional degrees of freedom, including side-chain rotations and bond vibrations, together with various interactions such as the hydrophobic effect, hydrogen bonding and electrostatic interactions, the free-energy landscape can effectively become a space of hundreds of dimensions and, in some cases, far more. It is obviously impossible to draw this vast high-dimensional space directly. Energy-landscape theory therefore compresses the overall trends and key features of this enormously complex space into the shape of a funnel rather than attempting to show every detail. The folding-funnel diagram is not a literal representation of the actual free-energy surface. Instead, it is a schematic illustration of the overall tendency for free energy to decrease as folding progresses. In other words, the funnel is a conceptual diagram that compresses the essential features into a simplified form that humans can understand. It is therefore important to recognise that it is not a direct drawing of the actual free-energy landscape.
At the top of the funnel are numerous states with relatively high energy, and free energy becomes lower as we move downwards. But the funnel is not simply a slide. As a protein moves down, it does not travel straight in one direction. Because of thermal fluctuations, it continually moves, backtracks and explores many different pathways. Even so, its motion has a bias. Because the overall direction favours decreasing free energy, the protein gradually moves into regions of lower and lower free energy. And what becomes important in this process is the interplay between thermodynamics and kinetics.
The interplay between thermodynamics and kinetics in the energy landscape
Here we arrive at two of the most important questions for understanding protein folding. Where is the stable state? And how quickly can the protein get there? The first is a thermodynamic question, and the second is a question of kinetics.
In an energy landscape, thermodynamics is represented by the depth of the valleys. The deeper the valley, the lower the free energy and the more stable the state. Kinetics, by contrast, is related to the height of the mountains separating the valleys—that is, the height of the free-energy barriers. No matter how deep and stable a valley may be, if a very high barrier lies on the way to it, the protein may not be able to reach that valley quickly. Conversely, if the free-energy landscape leading towards the native state slopes down relatively smoothly overall, the protein can reach that region comparatively quickly by using multiple pathways. Thermodynamics determines which valley is deepest—in other words, where it is most favourable to go—whereas kinetics determines how quickly the protein can reach it. If there are deep intermediate valleys, or local minima, the protein may become temporarily trapped there. These are called kinetic traps. Naturally, the more such traps there are, the slower the folding process becomes. It is therefore thought that natural proteins have evolved towards relatively smooth, minimally frustrated free-energy landscapes that minimise such traps.
The evolution of protein folding
Taking this idea further, the concept of an efficient free-energy landscape also gives us a new way to think about protein evolution. It is difficult to imagine that, throughout evolution, proteins developed by randomly inventing completely new structures and folding mechanisms from scratch each time. Over long evolutionary timescales, rather than creating and testing an infinite number of folding strategies from the beginning whenever a new protein arose, life may have evolved in a smarter way by repeatedly reusing folding structures that had already proved capable of folding quickly and stably while performing their functions well. As a result, the protein world contains many examples in which proteins with very different amino acid sequences nevertheless share remarkably similar three-dimensional folds. Motifs and domains are representative examples of this repeated use of structural principles across different proteins. Domains, in particular, can fold as independent structural and functional units and can be reused in different proteins or freely combined with other domains, rather like Lego blocks, to create proteins with new functions. A good example is the SH3 domain. This domain is found in humans, yeast and plants, and although their amino acid sequences differ considerably, they nevertheless share remarkably similar three-dimensional structures. This suggests that what evolution preserves is not necessarily only the amino acid sequence itself, but may also be the structural design—the fold or topology—that allows a protein to fold stably and function effectively.
This leads to an even more interesting thought. If structural patterns in proteins have been repeatedly reused throughout evolution, could this repetition itself provide a very important clue for understanding and predicting protein folding? Over billions of years, nature must have gone through endless trial and error, creating and selecting vast numbers of proteins. In that process, stable and functionally successful structures would of course have survived and been used repeatedly. If so, perhaps we can work backwards from the enormous amounts of protein sequence and structural data left behind by that process and learn the rules of nature.
In fact, as science and technology have advanced rapidly in recent years, enormous amounts of protein sequence and structural data have accumulated. With the emergence of artificial intelligence, we can now analyse and learn from these vast datasets with a speed and accuracy that human ability simply cannot match. One of the most representative results is AlphaFold. Through this learning process, AlphaFold has reached the point where, given only an amino acid sequence, it can predict which three-dimensional structure that protein is most likely to adopt. AI unquestionably played a decisive role in making such prediction possible. That cannot be denied. But perhaps the more important question is the fundamental reason why such prediction could be possible in the first place. The traces of repetition and reuse that nature has accumulated layer by layer. And what explains the successful folding that made this possible is the free-energy-based energy landscape we have been exploring here. We will look at this in more detail in the next article.
[Image sources]
Image 1
1-2. Folding Funnel — CC BY-SA 4.0
Image 2
2-1. Folding funnel schematic — CC BY-SA 3.0
2-3. Funnel-shaped energy landscape — CC BY 4.0

