Modern hierarchy of the levels of protein molecule structural organization
14 min read
In 1959, the Dane K. Linderstrøm-Lang proposed distinguishing four levels of protein structural organization: primary, secondary, tertiary, and quaternary structures, which denote, respectively, the amino acid sequence, the ordered structure of the polypeptide backbone, the three-dimensional structure of the protein, and the structure of protein aggregates (Fig. 10; Elliott W., Elliott D., 2000).
(b) Secondary structure Individual segments of the polypeptide chain exist in the form of an α-helix, a β-pleated sheet, or a random coil
α-helix β-pleated sheets random coil or loop
Secondary structures are folded into a compact globular protein
The molecule of such a protein can be represented in this way:
(d) Quaternary structure Individual protein molecules united by non-covalent interaction into a multimeric protein
Fig. 10. Levels of structural organization of a protein.
This classification prevailed until the early 1980s, when G. Schulz and R. Schirmer proposed supplementing it with two more levels of organization: supersecondary structures and domains. As a result, the concept of six levels of protein structural organization with a specific hierarchy, expressed by G. Schulz and R. Schirmer in the form of a diagram, was established.
Fig. 11. Modern concept of the structural organization of proteins
As can be seen from the diagram, the structural organization of proteins is based on a specific genetically determined amino acid sequence of the polypeptide chain, i.e., the primary structure, which determines all subsequent higher levels of organization.
The primary structure of a protein characterizes the sequence of amino acid residues of a polypeptide chain linked by covalent bonds. It is encoded in the nucleotide sequence of messenger RNA, which, in turn, is determined by genes. The primary structure of a protein is determined by: 1) the nature of the amino acids included in the molecule; 2) the relative amount of amino acids; 3) a strictly defined amino acid sequence. In 1953, Professor Fred Sanger of Cambridge University was the first to determine the amino acid sequence of a protein (insulin) and was awarded the Nobel Prize for this in 1958.
Fig. 12. Primary structure of a protein: A – amino acid chain;
B – fragment of the primary structure of the insulin molecule.
The polypeptide chain contains a free amino group at one end (N-terminus) and a carboxyl group at the other (C-terminus). The beginning of the chain is taken to be its N-terminus; it is from here that the counting of amino acids begins. This coincides with the direction of polypeptide chain synthesis. The amino group at the N-terminus of a polypeptide chain can sometimes be acetylated, i.e., have an acetic acid residue attached (CH3 – CO), as for example in cytochrome C1. In a number of proteins, the N-terminus is a pyroglutamic acid residue that does not contain a free amino group.
At the C-terminus, either a free carboxyl group or an amidated one is found. Modifications of the C-terminus are rarer compared to N-terminal modifications.
The main bond of the primary structure of proteins is the peptide bond. It is formed as a result of a condensation reaction between the amino group of one free amino acid and the carboxyl group of another, or between the amino group of a free amino acid and the carboxyl end of a polypeptide (water is released in the process). The polypeptide bond is quite rigid, so its conformational mobility is limited. However, in each amino acid unit, there is an α-carbon atom which accounts for the presence of two identical bonds in this unit; rotation is possible around these bonds. The rotation angles of identical bonds are called torsion angles and are designated as ψ and φ (C–Cα).
The polypeptide bond (covalent nitrogen-carbon bond) is the main type of bond determining the primary structure of a protein. However, the presence of disulfide bonds between two cysteine residues in one polypeptide chain is also possible, leading to the formation of cystine.
The primary structure of a protein predetermines the subsequent levels of organization of the protein molecule.
The secondary structure of a protein is the method of folding a polypeptide chain into an ordered structure. There are two main types of secondary structure: 1) α-helix and 2) β-configuration (β-structure). To simplify the representation of protein molecules, biochemistry uses conventional symbols for secondary structures (Fig. 13; Elliott W., Elliott D., 2004).
Fig. 13. Symbols used to represent segments of α-helix (a) and β-pleated sheets (b). On the right, their right-handed curvature, characteristic of anti-parallel sheets, is shown. Intermediate loops and disordered chain segments are depicted by a simple line.
The term "α-helix" was introduced in 1951 by L. Pauling, who discovered the folding of a polypeptide chain in the form of a right-handed helix in the protein α-keratin. Most proteins have the form of an α-helix. In shape, an α-helix resembles a rod, in which the stem is the backbone, and the branches sticking out in different directions are the side chains (R-groups). For clarity, it can be imagined as a regular helix formed on the surface of a cylinder. There are 3.6 amino acid residues per turn of the helix. This means that the C==O group of one peptide bond forms a hydrogen bond with the N—H group of another peptide bond located 4 amino acid residues away from the first. Both C==O and N––H bonds are directed parallel to the helix axis and are paired opposite each other; such an arrangement is optimal for the formation of a hydrogen bond and, consequently, for the stabilization of the α-helix. In cross-section, an α-helix looks like a disk from which the amino acid side chains point outward (Fig. 14; Rees E., Sternberg M., 2002).
Fig. 14. Secondary structure of a protein: A – right-handed α-helix; B – schematic representation of an α-helix; C – end-on view of an α-helix; D – antiparallel β-structure; E – parallel β-structure; F – detailed representation of a β-sheet.
The van der Waals radii of the atoms are such that there is no empty space inside the helix; this ensures the stability of the α-helix. The side radicals of amino acid residues do not participate in maintaining the α-helical configuration; therefore, all amino acid residues in the α-helix are equivalent.
A polypeptide chain can be positioned in space such that its individual sections are brought together and are parallel to each other. Hydrogen bonds can form between such sections of the polypeptide chain. This secondary structure is called a β-configuration or β-structure. Outwardly, a polypeptide chain with a β-configuration resembles a pleated sheet, i.e., folded like an accordion. The principle of organization of a β-structure (β-pleated sheet) is extremely simple. The polypeptide chain is in an extended state, and its C=O and N-H groups are connected by hydrogen bonds with the same groups of neighboring parallel-oriented polypeptide chains. Both chains can be independent or represent fragments of one chain common to both. The sections of the polypeptide chain forming the β-structure (pleated sheet) can be oriented in the same or opposite directions: in the first case, the pleated sheet is called parallel, and in the second – antiparallel. An antiparallel β-structure typically arises when the peptide chain turns back, forming a so-called hairpin. The turn site is called a β-bend. Along with hydrogen bonds, both forms of the secondary structure possess other bonds (Fig. 15; Müller G., Cordes U., 1970):
Fig. 15. Bonds stabilizing the secondary and tertiary structure of proteins
1) electrostatic, which arise between two oppositely charged polar groups, for example, between negatively charged side chains of aspartic or glutamic acids and positively charged protonated bases (side chains of arginine, lysine, histidine). These bonds are stronger than hydrogen bonds;
2) hydrophobic, which arise between non-polar, water-insoluble groups (CH2-CH3 groups of valine, leucine, isoleucine, the aromatic ring of phenylalanine). These radicals move closer together due to their being expelled from the water.
Electrostatic interactions participate in the stabilization of the secondary structure, but to a lesser extent than hydrogen bonds.
The secondary structure of a protein is determined by its primary structure. Amino acid residues are capable of forming hydrogen bonds to varying degrees; this influences the formation of an α-helix or a β-sheet.
Table 10 – Ability of protein amino acids to form an α-helix or β-structure
Classification of amino acids by structural organization
| Amino acid capable of forming an α-helix | Amino acid capable of forming a β-structure | Ability of the amino acid to form a secondary system |
| Glutamine, alanine, leucine | Valine, isoleucine, methionine | Actively form |
| Histidine, glutamine, valine, phenylalanine, tryptophan, methionine | Tryptophan, tyrosine, glutamine, leucine, cysteine | Prone to formation |
| Lysine, isoleucine | Alanine | Weakly form |
| Asparagine, arginine, serine, tryptophan, cysteine | Asparagine, arginine, glycine | Indifferent to this structure |
| Asparagine, tyrosine | Histidine, lysine, serine, asparagine | Counteract structure formation |
| Glycine, proline | glutamine | Disrupt this type of structure |
Helix-forming amino acids include alanine, glutamic acid, glutamine, leucine, lysine, methionine and histidine. If a protein fragment consists mainly of the amino acid residues listed above, an α-helix will be formed in this section.
Valine, isoleucine, threonine, tyrosine, and phenylalanine facilitate the formation of β-sheets in the polypeptide chain. In sections of the polypeptide chain where amino acid residues such as glycine, serine, aspartic acid, asparagine, and proline are concentrated, disordered structures arise. Many proteins contain both α-helices and β-configurations simultaneously.
Features of the formation of the supersecondary structure of proteins
The supersecondary structure of proteins consists of energetically preferred ensembles of secondary structures formed through the interaction of α-helical and β-structural sections of proteins. This level of organization is characteristic of both fibrous and globular proteins.
In fibrous proteins, two α-helices can be twisted relative to each other, forming a left-handed superhelix. Superhelical coiling is energetically favorable because additional non-covalent (van der Waals) contacts form between the side radicals of amino acids belonging to different α-helices. Such helices are found in α-keratin, tropomyosin, and other proteins, which, obviously, imparts specific functional properties to them: strength, elasticity, and the ability to contract.
In globular proteins, supersecondary structures consisting of two parallel β-sheets with various types of articulation between them are more common. They can form the following structures:
- structures in the form of a random coil (βcβ);
- α-helices
- β-structures (fig. 16; Konichev A.S., Sevastyanova G.A., 2003).
Two consecutively connected βαβ regions are called the Rossmann fold unit). A supersecondary structure in the form of an antiparallel three-stranded β-structure (βββ) is called a β-zigzag. It is common for characteristic supersecondary structures to contain metal atoms in their composition, which apparently provides them with additional stability.
Fig. 16. Protein supersecondary structure: A – random coil βcβ-unit; B – two consecutively connected βαβ regions;
C – zigzag (antiparallel three-stranded β-structure).
Protein domain structure. Supersecondary structures of a certain type are reproduced in many proteins, forming structural blocks. If a protein contains more than 200 amino acid residues, then several more or less independently formed compact regions are usually found in its structure. Their discovery formed the basis for the concepts of the modular (domain) principle of protein organization. Domains are structurally and functionally distinct regions (subregions or modules) of a molecule, usually containing from 40 to 300 amino acid residues and resembling separate small proteins in their structure. Domains are often connected to each other by short segments of the polypeptide chain, which act as hinges, providing a certain mobility of protein molecule segments relative to each other.
Domain structure is often associated with the possibility of representing the functional activity of a protein as a set of separate elementary processes. In addition, structural-functional studies of some enzymes have revealed an interesting feature of domain structure – the presence of subdomains or structural domains in their composition, each of which is capable of performing a specific catalytic function. The active center of domains is usually located in a depression between structural domains, i.e., it is located at the interface of structural domains. A mammal enzyme – fatty acid synthase – possesses a complex domain structure; within two of the six domains forming it, it was also possible to identify several subdomains. The single polypeptide chain of the enzyme contains everything necessary to catalyze seven reactions, i.e., each of the domains or subdomains catalyzes a specific reaction in a series of chemical transformations leading to the synthesis of palmitic and stearic acid molecules (fig. 17; Konichev A.S., Sevastyanova G.A., 2003).
Fig. 17. Protein domain structure: A – domain structure of immunoglobulin G; The immunoglobulin G molecule consists of two heavy (H) and two light (L) chains connected by disulfide bridges. Domains are indicated as circles: CH and CL – constant domains of heavy and light chains, respectively; VH and VL – variable domains. Interdomain regions of heavy chains possess a certain mobility: in the region of the two disulfide bridges connecting the heavy chains, there is a so-called hinge region, which confers mobility to the Fab ("arms") regions of the immunoglobulin. Variable domains form a functional domain necessary for antigen binding. B – domain structure of endonuclease. Domain I is responsible for proteolytic activity, domain II – for nuclease activity.
In all globular proteins consisting of domains, there is a high degree of affinity between amino acid residues located close to each other along the chain. On this basis, it is assumed that domains are formed independently of one another, which simplifies the macromolecule folding process.
The idea of a modular design of enzymes and other proteins allows for the frequent emergence and rapid evolution of new functional proteins. Proteins formed as a result of the association of different domains from already existing molecules can possess a wide variety of new functions. Such a mechanism for the formation of new proteins could operate much faster than random point mutations.
All domains can be subdivided into four classes or groups and α+β, depending on the mutual arrangement of α-helical and β-structural segments in the chain. α/α-Domains consist mainly of α-helices; β-segments are practically absent in them. In β/β-domains, there are several β-strands and no (or almost no) α-helices. In α/β-domains, α- and β-segments alternate along the chain. Often, β-segments form a parallel β-sheet surrounded by α-helices. In α+β domains, α- and β-segments are usually located in different segments of the polypeptide chain.
Protein tertiary structure is the way all atoms of a single chain are arranged (fig. 18; Dobrynina V.I., 1976). For a protein molecule to acquire its inherent specific properties, the polypeptide chain must fold in a certain way in space, forming a functionally active structure. Such a structure is called native – a protein with the original, natural folding of the polypeptide chain. Despite the vast number of theoretically possible spatial structures for an individual polypeptide chain, protein folding leads to the emergence of a single native configuration. The tertiary structure, just like the secondary one, is determined by the amino acid sequence of the polypeptide chain, but while the secondary structure is determined by the interaction of amino acids in the closest segments of the chain, the tertiary structure depends on the amino acid sequence of segments located far from each other in the chain. The formation of kinks in the polypeptide chain, as well as the direction and angle of the chain rotation at these kinks, are due to the number and position of certain amino acid residues, such as proline, threonine, and serine, which facilitate the formation of kinks in the helical segments of the protein. The tertiary structure of a protein is stabilized by bonds and interactions between the radicals of the amino acid residues of the polypeptide chain. Interactions between side radicals of amino acid residues are divided into strong and weak.
Fig. 18. Tertiary structure of a protein: A – schematic representation of a myoglobin molecule; B
– direction of the polypeptide chain.
Strong bonds include covalent bonds between two sulfur atoms of cysteine residues located in different parts of the polypeptide chain. Otherwise, such bonds are called disulfide bridges. A disulfide bridge is formed during the oxidation of two cysteine residues, i.e., upon the removal of hydrogen from two reactive sulfhydryl groups – SH. The new residue is called cystine.
Weak interactions arising between the side radicals of amino acid residues in different parts of the polypeptide chain are divided into polar and nonpolar. Polar interactions include ionic and hydrogen bonds. Ionic interactions are formed upon contact between positively charged groups of lysine, arginine, and cysteine side radicals and the negatively charged COOH group of aspartic and glutamic acids. Hydrogen bonds arise between functional groups of side radicals of amino acid residues. Nonpolar and van der Waals interactions between hydrocarbons and amino acid residue radicals facilitate the formation of a hydrophobic core (a fatty droplet) inside the protein globule, as hydrocarbon radicals tend to avoid contact with water. The greater the content of nonpolar amino acids in a protein, the more significant the role of van der Waals bonds in the formation of its tertiary structure. As a result of a multitude of relatively weak bonds, all parts of the protein's peptide chain become fixed relative to each other, forming a compact structure. Proteins are divided into two groups by shape: 1) globular, i.e., having a spherical shape, and 2) fibrous (thread-like), having a highly elongated shape. As a rule, globular proteins perform dynamic functions, while fibrous proteins perform structural functions.
The tertiary structure of a protein is unique, just as its primary structure is unique. Only the primary spatial folding of a protein makes it active. Various disturbances in the tertiary structure lead to changes in protein properties and the loss of biological activity.
Quaternary structure of a protein. The tertiary structure completes the description of a protein molecule's structure. However, there is a significant number of proteins whose molecules are complexes formed from several protein molecules connected by weak bonds (hydrogen, electrostatic, hydrophobic interactions) and not connected by covalent ones (peptide, disulfide). Such complexes are called oligomeric, multimeric, or subunit proteins. Their composition and stoichiometry are constant; this proves that the assembling protein subunits (protomers) "recognize" each other due to the presence of complementary-shaped areas on their surface. The arrangement of subunits in a functionally active complex is called the quaternary structure of a protein (Fig. 19; Dobrynina V.I., 1976). The areas of subunits where interactions occur are called contact sites. Protomers are connected to each other by noncovalent bonds located on the contact surface of each of them. Since these bonds are weak, dozens of bonds are formed between each pair of protomers. The process of self-assembly of the quaternary structure from protomers is highly specific; contact surfaces on one protomer correspond exactly to the contact surfaces of another protomer, so that upon contact, oppositely charged ionic R-groups or R-groups capable of forming hydrogen bonds or hydrophobic surfaces coincide. Such contact surfaces are called complementary; they fit together like a key to a lock, so erroneous connection of protomers in an oligomeric protein or connection with other proteins is excluded. Complementary interactions lie at the heart of almost all biochemical processes in cells, including enzymatic reactions, processes of compound transport across membranes, as well as the protective functions of proteins. Proteins with a molecular weight of more than 50 thousand Da possess a quaternary structure. A characteristic property of proteins with a quaternary structure is that an individual subunit does not possess biological activity. The quaternary structure of a protein is as specific and unique a characteristic of a given protein as other levels of structure. The quaternary structure plays an important role in the regulation of protein biological activity, as it is very sensitive to external conditions: small deviations in them can cause a change in the arrangement of subunits and, consequently, a change in the protein's biological activity. This phenomenon is one of the main mechanisms for regulating metabolism, as enzymes and some other metabolically active proteins have a quaternary structure.
Fig. 19. Quaternary structure of a protein: A, B. – schematic representation of a hemoglobin molecule's structure.
Read next
Agrochemistry For students
Physicochemical properties and classification of polysaccharides in agricultural chemistry
Agrochemistry For students
Classification and chemical structure of oligosaccharides in plant organisms
Agrochemistry For students