Thursday, August 09, 2007
Classification of Protein Structure
Classification of Protein Structure: The average protein consists of about 20000-30000 atoms and in order to make sense of the structure of the protein, it is necessary to simplify the protein structure. There are more than 30000 structures in the protein database [4] and to go about looking at each structure would be horrendous. Hence, one needs to come up with a classification scheme for protein structure.
One way of simplifying it is to break it into parts called the secondary structure of the protein. There are various courses/books [1,2] to explain the secondary structure, but for our discussion, it is sufficient to know that certain configurations called alpha helices and beta sheets are common structural motifs found in nearly all proteins. While these secondary structures do help in understanding the local structure of the proteins, they give very little insight about the chemistry that the protein is able to perform and about it's active site itself.
A second and more meaningful attempt at classification of protein structure would be to find certain common structural motifs that can exist independently and classify the proteins based on these structural motifs. For example, a protein can be multifunctional but each function can be carried out independently by different parts of the protein even if you split them up. It could make sense that one can split these multifunctional proteins up based on what function they perform and if you can find the same structure/function motif in different proteins, club them together as a single group. In a protein, the part of the protein that can maintain its structure and function independently is called a domain [5]. Quite often domains of one shape combine with domains of very different shapes to form quite different proteins (very much like building blocks can come together in various different configurations giving walls of various different shapes) [6].
To give meaning to the classification scheme, it would also help to know which proteins perform related function (for example, perform the same reaction on different substrates) or have active sites in the same region of the protein structure. In order to give meaning to this classification, it is better to form groups of more closely related structures that perform similarly. Great minds have always argued that evolution should be the guiding principle while studying biology and it does make sense to classify proteins which have a common evolutionary origin from those that have achieved the same structure independently (also called convergent evolution).
There are various different databases that divide proteins into individual domains and divide these domains up into evolutionarily related groups heirarchically. These databases include the SCOP (Structural Classification Of Proteins) [7](manually divided), CATH (Class,Architecture, Topology, and Homologous superfamily) [8], and FSSP (Families of Structuraly Similar Proteins) databases [9](automatically performed). However, these databases are often flawed and corrections to these databases are often suggested in literature. Part of the problem is it is very difficult to say when similarity in structures occured due to homology (evolutionarily related), or convergence (evolutionarily independent origins). The trivial relationships are those that are apparent in the sequences of the two proteins. When two proteins have very similar sequences (measured by the number of times they have the same amino acid or a slightly related amino acid in the same position of the structure), they are related and statistics based on extreme value distributions can be used to find the probability of both proteins having a common origin [11]. However, the structure remains conserved (does not vary much) much more than sequences and below a certain sequence identity, it is very difficult to prove that there is a relationship between the two proteins without a structure [10].
Other problems that come up are related to the process by which structures are obtained (X-ray crystallography or NMR spectroscopy). These methods are inherently noisy because of various problems such as Heisenberg's uncertainty principle onto the crystallization conditions and the substrates that interact with the protein. So there is never a completely correct structural alignment (that is finding one to one which residues in the structure overlap each other) that also causes minor problems in the classification procedure.
But the most important problem is the level at which to classify structures. While domains are the most commonly used level of classification (because a domain is basically independent), during the evolution process, domains might not have been the basic level at which proteins were constructed. Rather subdomain level small structural units called structural words [12] or foldons [13] (because they could be independent folding units) could also be the smallest level of proteins that had evolved from the RNA world. The theory is that these foldons could come together and form various different domains and then evolved further to form proteins with different functions.
References:
[1] - Biochemistry by Stryer.
[2] - Introduction to Protein Structure by Branden and Tooze.
[3] - A perspective on enzyme catalysis by Stephen Bankovic and Sharon Hammes-Schiffer
[4] - RCSB protein database.
[5] - Domains.
[6] - Multi-domain protein families and domain pairs: comparison with known structures and a random model of domain recombination by Gordana Apic, Wolfgang Huber & Sarah A. Teichmann.
[7] - SCOP.
[8] - CATH.
[9] - FSSP.
[10] - How far divergent evolution goes in proteins - Murzin.
[11] - Maximum Likelihood Fitting of Extreme Value Distributions - Eddy..
[12] - On the evolution of protein folds - Lupas, Ponting, and Russell.
[13] - Foldons, Protein Structural Modules, and Exons by Anna Panchenko, Z. Luthey-Schulten, and P.G. Wolynes.
Thursday, May 24, 2007
Biological Control - Doing it yourself.
RNA was considered as a step required in modern organisms to convert DNA to proteins. RNA is made up of nearly the same chemical constituents as DNA but it is more flexible and can have wide ranging 3 dimensional structures unlike DNA's double helical structure. However, this increased flexibility comes at a price - RNA is more unstable and in modern cells, a single molecule of RNA does not remain functional for long periods of time (mean life time is approx 5 minutes in E.coli).
Of course, all this changed when it was found that RNA molecules could be used as catalysts and even in modern day cells, there are some RNA catalysts also called ribozymes (and the list of ribozymes discovered keeps increasing). RNA captivated the imagination of biologists as this was a molecule that could store genetic information as well as be used as catalysts - taking on the dual role of enzymes and information storage. All of a sudden, RNA was considered to be at the origin of life as we know it. However in the RNA world hypothesis, one should take into consideration that it is not that only RNA is present. It only postulates that RNA is present and is dominant but other biochemicals such as peptides (small proteins) and DNA oligomers (small DNA molecules) are also present and aiding life (idea originally proposed in [1]).
One of the biggest controversies against the RNA world hypothesis has been that it does not play that big a role in modern cells. However, it has been found more recently that there are many RNA control elements in the cell. One such control element is the riboswitch. For a gene to be made, the DNA gets converted into a message called the mRNA (messenger RNA) which later gets converted to the protein equivalent to that message. It has increasingly been found that mRNA do not contain only the message to be read but certain control elements could also be present in the mRNA. These control elements are called riboswitch.
Lets take an example. Supposing you want to make Vitamin B1. There is an intermediate in its biochemical pathway called thiamine pyrophosphate (TPP). TPP is also important for nucleotide (the chemical constituent of RNA and DNA) and amino acid (the chemical constituent of proteins) biosynthesis and is important for the cell to have the right amount of TPP channeled into the different biochemical pathways. When too much of TPP is present in the cell, TPP binds to a certain riboswitch in it's own biochemical pathway. This causes the riboswitch [2] to suddenly have a defined 3-dimensional structure (from an earlier random or semi structured RNA element). This defined 3-dimensional structure also blocks the production of the protein for making more TPP. The switch in the mRNA turns the production of the protein that makes TPP on or off depending on whether enough TPP is present in the cell or not - hence regulating the production of TPP itself. So far, riboswitches are found more in the microbial world and are only now being found in the eukaryotic world.
Now, in the latest issue of Nature, the first riboswitch that controls splicing in higher organisms such as fungus has been found [3]. Splicing is the mechanism by which parts of the mRNA are removed before the protein is made so that parts of the DNA never translated in the protein. Alternative splicing is the mechanism by which a single gene at the DNA level can be translated into multiple protein molecules. This is done by excising different parts of the mRNA (excising the DNA only in one situation and not another) before it gets converted to protein. Splicing and alternative splicing occurs only in eukaryotes and has also been discussed here.
Anyways, the first riboswitch in the mRNA have been found to function for alternative splicing purposes. The TPP biochemical pathway discussed above is the system that they found riboswitches in. In this case, when TPP was present, the riboswitch forms a three dimensional structure that avoids splicing and the protein that is formed can not make more TPP. So the objective was again control of TPP concentration in the cell but the means used was alternative splicing instead of just blocking formation of protein. The implications of these results will only come out with time, but there is speculation that this opens up a whole pandora's box on riboswitches that could be found in eukaryotes.
[1] The Genetic Code - Carl Woese, 1968.
[2] Thiamine derivatives bind messenger RNAs directly to regulate bacterial gene expression. Wade Winkler Ali Nahvi & Ronald R. Breaker. Nature 419, 952 - 956 (2002).
[3] Control of alternative RNA splicing and gene expression by eukaryotic riboswitches. Ming T. Cheah, Andreas Wachter, Narasimhan Sudarsan & Ronald R. Breaker. Nature 447:497 (2007) and its companion discussion article - Molecular biology: RNA in control. Benjamin J. Blencowe & May Khanna. 447:391 (2007)
pdf of all cited aritcles avaiable on request
Thursday, November 02, 2006
In Living Color (Part 1)
It is often said that the 21st century will be (is) the age of biology; much like the previous century was for physics. New discoveries are occurring and biological information is growing both in size and complexity at an exponential rate. A major factor fueling this growth is the plethora of technologies available to the modern biologists in their quest to uncover the very basic molecular mechanisms of life.
green fluorescent protein (GFP).
The jellyfish, Aequorea victoria, on the left and green bioluminescence observed around the margin (note the picture on the left does not show fluorescence !)
However, it took another thirty years before the GFP became the almost ubiquitous cellular and molecular biology tool it is today. In 1987, Doug Prasher, then at the Woods Hole Oceanographic Institute, discovered and was able to make a copy of the DNA sequence within the jellyfish gene that encoded for GFP. He did not, however, succeed in making a glowing protein from the DNA sequence in the lab. Subsequently, Prasher sent his sequence to a researcher at typical trick used by biologists to make proteins. As shown in the figure on the right from Chalfie's work, the bacteria containing the genes for GFP (on the right side of the plate) emits green light under illumination with ultaviolet lamp (it was a graduate student doing rotation in Chalfie’s lab that actually performed the work and made the discovery!). This seminal work, published in 1994 in the journal Science, led to 'an explosion of color' in the biological world. Subsequently, other researchers showed that the GFP could be produced, alone or in tandem with other proteins in a variety of organisms.
Over the last decade, a great deal of research has contributed towards understanding the underlying physical mechanisms of GFP’s light emission2 and importantly, towards improving its properties through genetic manipulation. The leader in this field has been Roger Tsien, who along with co-workers demonstrated that making small changes, such as replacing a few amino acids in GFP could make it glow brighter, mature faster and prevent aggregation of the protein inside cells. His group has also succeeded in tuning the absorption and emission of the original GFP through mutagenesis, leading to a veritable palette of fluorescent proteins that absorb and emit light through the entire span of the visible light spectrum (see below). Additionally, a Russian scientist, Sergey Lukyanov, used the GFP sequence as a 'bait' to search for novel fluorescent proteins in corals and succeeded in finding several GFP-like proteins, particularly a red-emitting fluorescent protein, dsRED from Anthozoa, which is also used widely.
Panel on top shows the fluorescent protein 'palette' developed by Tsien lab - note range of colors and the fruity names. On the left, artwork with bacteria expressing various colors of fluorescent protein.
The major advantage of GFP is that inside a living cell, it can emit light on its own without the help of another protein or other chemicals. It is also possible by using molecular biology techniques, to attach the DNA of GFP to the DNA of the protein of your choice to produce a recombinant DNA. When the information from such recombinant DNA gets translated into a protein within the cell, a tandem protein is created with the GFP unit hanging from the protein. Importantly, since the size of GFP is relatively small, in most cases it does not interfere with the regular functions of the protein it is attached to.
In the simplest of applications, after shining light on the cells, the total amount of fluorescence obtained from the cells provides a measure of the level of expression of the protein tagged with GFP. However the more useful applications involve cells placed under microscopes with high magnifying power (40x to 100x) in conjunction with either arc lamps or lasers for bright illumination and high-resolution detection devices such as CCD cameras for
capturing images of the emitted light. In these cases, we can literally see where the protein of interest is located, or illuminate a particular subcellular structure. For example, the figure on the right shows the mesh of protein network that act as a 'skeleton' (in fact it is called the 'cytoskeleton') in majority of cells in higher organisms. A protein called 'actin' that is involved in this scaffold has been tagged with GFP.
It is also possible to tag two or more proteins in the cells with different fluorescent colors (see the fluorescent protein palette above) and follow their localization or movement in cells. This helps in noting where two proteins are localized during a cellular function. In the figure to the right, the protein actin is now tagged with a cyan emitting fluorescent protein (CFP). Another protein, vinculin, has been tagged with a yellow fluorescent protein (YFP). You can observe that the YFPs are localized as small elliptical structures at many places. These are called focal adhesions, which form a link between the cell cytoskeleton (in cyan) and its extra-cellular matrix. This interaction help cells to adhere and eventually move about in a tissue.
(in a future post, I will talk about a technique involving fluorescent proteins of two colors which is used to determine if two proteins interact with each other inside a cell)
Perhaps the most powerful application of fluorescent proteins is when you combine microscopy with time lapse video images. In such cases, it is possible to observe where and when the translocation of the protein in cells is taking place under a biological condition. For example, see a video here of a cell moving around with a protein involved in the focal adhesion tagged with GFP.
Visualization of protein location and dynamics in this manner enable scientists to place cells under various physiological conditions and observe the resultant phenotype of the protein behavior. Before the advent of GFP, scientist had to destroy the cells and use other tedious biochemical techniques to obtain similar information. Even then real-time data acquisition was not possible.
A quick search of the database will reveal more than ten of thousands of peer-reviewed publications where fluorescent proteins have been used to study protein functions at the cellular level. In most of these cases, research was conducted with either unicellular organisms or cells derived from tissues of mammals. However, apart from single cells, fluorescent proteins are also being used at the tissue and even the whole organism level. The picture below shows an example of a research which is investigating the movement of neurons (labeled with GFP) in the cerebral cortex.
A more well-known example of GFP in whole organisms, is the development of 'fluorescent mice' by the company Anticancer Inc . It is easy to follow tumor progression and cancer metastasis in such mice. Also, a Taiwanse reasearch group recently created 'fluorescent pigs'. Stem cells or organs from these pigs when transplanted into other organisms can be followed easily without requiring invasive techniques.
More examples of such applications of GFP can be found here. Apart from these animals, 'Alba' , the fluorescent rabbit and fluorescent aquarium fishes are two examples of more esoteric application of this scientific technology.
On a final note, betting markets for the Nobel Prize (yes they do exist !), were predicting this year’s Chemistry Nobel to go to Roger Tsein and others for their work on fluorescent proteins. It eventually went to Roger Kornberg for his work on DNA transcription. Considering the importance of fluorescence proteins and their wide-ranging revolutionary impact on biology, it is not far-fetched to think that the Nobel is not beyond the grasp of these researchers.
Notes:
1. Aequorin itself has been very useful for visualizing cellular calcium concentrations, the regulation of which is important for a number of physiological activities.
2. Without going into great details about physi-chemical mechanisms of GFP fluorescence, suffice to say that the protein has a barrel-like structure (see below); within the barrel, three critical amino acids are brought together in close spatial proximity, which forms the chromophore.
Artistic rendition of the three-dimensional structure of GFP.
4. Recommended further reading: This web-site is a very good resource for learning more about GFP's discovery, structure and applications. Also read this interview with Dr. Martin Chalfie.
Coming up: "Much to fret about ": on a technique known as fluorescence resonance energy transfer that enables biological distance measurements, detection of protein interactions, and can be used to look at protein functions at a single molecule level !
