Field of Science

Showing posts with label protein folding. Show all posts
Showing posts with label protein folding. Show all posts

In which creationists' understanding of amyloid appears...tangled

The biophysicist David Eisenberg of UCLA recently published a paper in which his group surveyed what they called the "amylome", the set of all possible proteins that can potentially form the deadly amyloid aggregate implicated in diseases like Alzheimer's. I haven't read the whole paper yet and will describe it in another post but it has some very intriguing conclusions (see the Nature News piece).

Eisenberg's group end up finding common segments possessing amyloid-forming propensity in pretty much every protein, given the right conditions. Not surprisingly, these segments are mostly kept tucked inside protein cores; if exposed the proteins are refolded with chaperones or eliminated as non-functional. But this ties in with Chris Dobson's work which I described in a recent post. Dobson's most recent paper seemed to conclude that amyloid is actually the most thermodynamically stable state of a protein, with "normal" protein states being metastable.

All this is fascinating stuff, but you can always trust creationists to put an anti-evolutionist spin on almost any scientific funding. Someone named Cornelius Hunter asserts on his blog that the fact that the most stable state of proteins seems to be amyloid and that most real proteins don't actually exist in this state seems to be a kind of miracle or at the very least indicates the enormous difficulties attendant in creating complex biomolecular structures.

Mr. Hunter seems to be indulging in a common fallacy, that of assuming that evolution somehow tends to an ideal. This is just not the way the process works. Considering the stringent constraints and time in which evolution has to work, it can only explore the available space of solutions and not the entire possible space. It can never achieve the best possible result in solution space, only one that is good enough under the given circumstances. I don't know if it's really that hard to understand or whether creationists like Mr. Hunter want to deliberately obfuscate the issue. Evolution can only work on what's already available; it can only mix and match existing motifs, and what exists need not be perfect at all. As an aside, this function of evolution reminds me of the concept of "satisficing" or "bounded rationality" in economics; lacking perfect knowledge of all solutions and unlimited time, economic actors like us can only pick solutions optimal within current constraints, not "global" maximums on the solution landscape.

There's therefore no reason to believe that biological evolution should create the most thermodynamically stable state of a protein. Creating a state that is functionally relevant is enough, even if it's thermodynamically metastable. In fact one can even make an argument that a thermodynamically superstable state might lead to an evolutionary dead end since it will be hard to tinker with. Given such constraints it's indeed impressive that evolution creates enzymes speeding up chemical reactions by twelve orders of magnitude, but even these enzymes are few and probably not the best possible in all of enzyme space.


So yes, the fact that it's amyloid and not the normal state of proteins that is the most thermodynamically stable one is fascinating, but in no way does this present a great challenge to evolution. In fact it reinforces evolution's essential character, to hunker down and make do as well as it can under the given circumstances. Don't we all?

The barrier to amyloid formation is kinetic, not thermodynamic

One of the questions I have pondered in the past is why the functional form of a protein should correspond to its most thermodynamically stable structure. Although this assumption is built into almost all experimental and theoretical studies of protein folding, it is not at all obvious since one may imagine other forms which could have improved stability. For instance, two protein forms may differ in the presence of a hydrogen bond or two. Based on the location and connectivity of these bonds, sometimes this slight rearrangement can cause a radical change in function, but there's no good reason why it should in the general case.

The answer however is most obvious in case of amyloid, that endlessly intriguing protein form that is implicated in so many devastating neurological disorders. Amyloid is a very stable state is often highly resistant to temperature, pH and high salt conditions. It's fair to ask how stable or unstable it is with respect to functional, soluble forms of the same protein.

To answer this question, a team led by Christopher Dobson who is a world expert on amyloid performed a series of thermodynamic measurements on a diverse group of proteins in which they measured the free energy differences between the soluble and the amyloid state. The proteins included everything from the Aß protein found in Alzheimer's disease to human lysozyme and insulin. The finding was that the free energy differences (ranging from about 3 kcal/mol to 6 kcal/mol) are not terribly dependent on the exact sequence, an observation which would be consistent with the striking recently uncovered fact that amyloid formation can be induced in almost any protein independent of its sequence. In fact the free energy difference seemed to depend more on the length and seemed to be optimal for a length of 100 residues for which the amyloid form was most stable. The difference also sharply tipped away from amyloid for increasing lengths.

This observation seems to suggest that one consequence of evolving larger proteins might be steer them away from the amyloid state and is consistent with the fact that almost all amyloid proteins have relatively short lengths (for instance, the Alzheimer's disease amyloid protein Aß has a length of roughly 40 residues). The propensity toward amyloid formation also depended on the concentration and the authors derived an limiting concentration beyond which amyloid formation would be rapid. This is again not surprising since the concentration-dependence of the process has also been demonstrated.

The real surprise came when they compared these limiting concentrations of the protein to the corresponding physiological concentrations of the same proteins in plasma. Remarkably, they found that in almost every case the physiological concentration was higher than that required to achieve amyloid formation. Thus
the observations clearly indicate that for many key proteins, the amyloid state is thermodynamically more stable than the native, functional state. To put it bluntly, many nicely folded and soluble proteins are actually metastable. Now, since native proteins don't constantly form amyloid and kill us all, it's clear that the barrier to amyloid formation must be kinetic. Intriguingly, the authors speculate that these barriers can be overcome when organisms are exposed to stress, mutations or aging.

This is a pretty intriguing study and seems to underscore the belief that at least for some proteins, the folded functional state is not the most stable. However in light of what we know about evolution, this should not be too surprising. Stability is just one of many factors to be optimized during natural selection and there is no reason to assume that evolution would always act to maximize this parameter at the cost of all others. It's worth always keeping in mind that evolution cannot afford to aim for the ideal but instead has to make do with what it has.

The other question in my mind is why in spite of these barriers existing in case of so many proteins like lysozyme, insulin etc. are they regularly overcome only in the case of Aß (1-42) and a select few others. Based on the speculation in the paper, this could be because these proteins are exposed to particularly harsh conditions that force them to climb past the kinetic barrier and settle into the amyloid valley of thermodynamic comfort and physiological woe.

Among many such conditions could very well be bacterial infections. A few years back I advanced a hypothesis about amyloid formation being a defense against viral and bacterial infection mediated through the production of free radicals. A kinetic barrier-surpassing mechanism of the kind speculated here might well be what allows these proteins to achieve the transition, killing the bacteria but ironically harming their owner in the process. In the context of the present study, I think there continue to be a lot of opportunities to investigate the possible infection-induced conversion of normal proteins to their amyloid form.

Hopefully someone will do the experiment.


Baldwin, A., Knowles, T., Tartaglia, G., Fitzpatrick, A., Devlin, G., Shammas, S., Waudby, C., Mossuto, M., Meehan, S., Gras, S., Christodoulou, J., Anthony-Cahill, S., Barker, P., Vendruscolo, M., & Dobson, C. (2011). Metastability of Native Proteins and the Phenomenon of Amyloid Formation Journal of the American Chemical Society DOI: 10.1021/ja2017703

The protein folding funnel and its discontents

Speaking of protein folding, here's something interesting. One of the most enduring views of protein folding from the last decade is that of an "energy funnel". The funnel was invented by the UCSD chemist Peter Wolynes in the 90s (the original paper is highly readable) and essentially depicts a plot of the configurational enthalpy (or effective energy) of the protein on the Y axis vs the configurational entropy on the X axis. In real situations this plot is multidimensional.

The funnel suggests a way out of Levinthal's paradox which contrasts the fast folding times for virtually all proteins with the vast amount of conformational space to be searched. According to the funnel viewpoint, the energy of the protein on the Y axis decreases and becomes more favorable even as the entropy on the X axis decreases, leading to fewer conformations to be searched and allowing the protein to rapidly find the native structure. The funnel has become a mainstay of descriptions of protein folding and has made its way into textbooks.

The funnel view of protein folding had always puzzled me a little for the simple reason that we usually think of the enthalpy and entropy of the protein (and in fact of any chemical system) as opposing factors. Entropy would hinder the protein even as it formed more "native" contacts and led to a favorable enthalpy. Yet the funnel seems to suggest a synergy between these two factors. Many papers have said that the funnel "guides" the protein to its correct conformational state. In this week's Nature Chemical Biology, one of the founding fathers of the field, Martin Karplus, sheds some light on this confusion and informs us that the traditional view of the funnel is indeed a little misleading.

To support his argument, Karplus illustrates two examples of protein folding studies using two kinds of systems. One is a lattice model system in which the protein is approximated by beads on a lattice. Native contacts in the protein are indicated by adjacent beads on the lattice. The other folding simulation is a standard molecular dynamics simulation of an alpha helix. In both cases the proteins are small (about 30 residues) but their behavior at low and high temperatures is intriguing.

At low temperatures, the folding landscape is more "rugged" and folding is slower. This is a well-established concept and it simply means that there is less energy for the protein to explore all the available local minima. At high temperature the landscape is "smooth" and the protein has enough energy to explore many conformational states. What is striking is that while the effective energy (enthalpy) at high temperature decreases smoothly all the way to the native state, the
free energy (which is what we should really be worrying about) has a significant barrier. Thus this barrier has to come from entropy. The crucial thing to note is that at high temperatures, the free energy is dominated by the increasing unfavorable entropy engendered by the greater number of conformations that the protein has to search.

Ultimately it's easy to forget that the protein folding "funnel" is only a theoretical construct, an intuitive model. Has anyone actually observed a funnel for a
real protein? As the article notes, for now the answer is a decided "No". Unfortunately it may be impossible to ever do so since to construct a real funnel one would need knowledge of every single conformational state that a protein visits on its way to folding. In addition since folding is a statistical phenomenon, one would also need knowledge of every starting trajectory. Needless to say, for now this is at best a pipe dream. However the funnel remains a useful construct provided we remember the subtleties and caveats that Karplus has described. Ultimately it's a model, and like other models it need not be real, but it should at least be useful.

Karplus, M. (2011). Behind the folding funnel diagram Nature Chemical Biology, 7 (7), 401-404 DOI: 10.1038/nchembio.565

The fine-tuning problem in protein folding: Is there a protein multiverse?

One of the deepest questions physicists have struggled with in the last half-decade is the so-called "fine-tuning problem". The fine-tuning problem asks why the values of the fundamental constants (Planck's constant, the speed of light, the mass of the electron etc.) are what they are.

The reason why physicists are so worried about the values of these constants is because presumably if the values were even a little different from what they are, the universe and life as we know them would not exist. For instance, even a slight weakening of the strong nuclear force that holds nucleons together would prevent the formation of atoms and thus of all complex matter. Similarly, a slight change in the electromagnetic force would fundamentally alter the interactions between atoms crucial for the formation of chemical bonds between the molecules of life.


There thus seems to be some factor during the evolution of the universe responsible for fine-tuning the values of the constants to their present values within an incredible window of accuracy. The fine-tuning problem is a real problem not least because some religious believers point to the unchangeable and precise values of the constants to be the work of some kind of intelligent designer.


In the last few decades there have been a few attempts to resolve the fine-tuning problem. Probably the most exotic and yet in some ways the most reasonable solution has been to assume the existence of multiple parallel universes. Multiple universes (or multiverses) were first proposed by Hugh Everett, a brilliant and troubled physicist who worked on nuclear weapons targeting, as a way around the so-called "measurement problem" in quantum mechanics. The measurement problem is fundamentally embedded in the quantum description of our world. The unsettling thing (and one that troubled Einstein) about quantum mechanics is that it assigns probabilities to certain events, but provides no answer as to why only one of those events materializes when we make a measurement. Everett worked around this conundrum by assuming that in fact all possible events actually do take place, but only one of them is part of our universe; the rest of the events also occur, but in parallel universes. Everett's interpretation which was regarded to be a fringe explanation for years (thus making it successfully into science fiction books) is now taken seriously by many physicists.


Being a problem associated with the most fundamental constants of nature, the fine-tuning problem makes its way into all "higher-level" sciences including chemistry and biology. In chemistry the fine-tuning problem takes on a fascinating form and entails asking why certain molecules have become fundamental to living systems while other more or less equivalent alternatives have been discarded during evolution. For instance, why alpha amino acids (and why not beta or gamma amino acids)? Why left-handed amino acids and right handed-sugars? Why phosphates and not sulfates or silicates? In retrospect one can think of answers to these questions based on factors like stability, versatility and ease of synthesis, but ultimately we may never know. However, the fine-tuning problem also manifests itself in one of the most fundamental processes in the workings of life; protein folding.


The protein folding problem is well-known; given an amino acid sequence, how can a protein fold into a single three-dimensional structure and reject the countless number of other possible structures it can fold into? What is even more remarkable about this problem is that
several thousand of those other structures are almost equienergetic with the preferred folded structure and yet they do not form. In fact it is this energetic equivalency between several structures that plagues all modern computational protein folding algorithms; the problem is not so much to generate the one correct structure as it is to distinguish it from other structures that are very close to it in energy. The fundamental assumption in all these algorithms is that the correctly folded structure is the lowest-energy structure. But that does not mean it differs in energy from the other solutions disproportionately. Therein lies the rub.

Ever since I heard about the protein folding problem this issue has bothered me as I am sure it has others. Consider that the free energy difference between two different protein structures may be only 5 kcal/mol or so, about the energy of a single hydrogen bond. Yet a protein when it folds unerringly picks only one among the two structures. How can nature manage to pick the right solution every one of millions of times when it folds proteins inside our body each second? To put it another way, here's the "fine-tuning problem" in protein folding:
why does a protein always adopt one and only one correct structure even when many other structures, very similar in energy and presumably in function, are available to it?

From a retrospective evolutionary standpoint the answer to this conundrum is perhaps not too surprising. Imagine what would happen if every time a newly synthesized copy of a given protein folded, it formed a slightly different structure. This heterogeneity and lack of quality control would play havoc with the intricate signaling networks in our body. Evolution simply cannot afford to have different three-dimensional structures for the same protein, no matter how slightly different they are. No wonder that quality control in protein folding is extreme. Of course nature does make occasional mistakes, but wrongly folded proteins are quickly degraded and destroyed.


Nonetheless, the original dilemma persists and metamorphoses into a further interesting question: isn't it possible for a protein structure that is slightly different from the one true structure to be functional? There are two possible answers here. Perhaps the alternative structure
was functional during evolution at one point, but competition from the slightly better structure weeded out the former from the gene pool. If this is the case, could there be a chance that there is some unknown form of life in which this other slightly different yet perfectly reasonable structure still exists, happily doing its job with no evolutionary pressure around to discard it? The best way to answer this question is to compare proteins from different species, something that has been extensively done for years. But such a comparison usually reveals protein homology, in which the sequences themselves are slightly different and yet perform similar functions.

That's not what we are looking for. What we are looking for is "two" proteins with
absolutely identical amino acid sequences which in two different creatures adopt slightly different three-dimensional structures and perform similar functions. Or they could even perform different functions, thus validating evolution as a force that puts slight differences to optimal use. Let us call these proteins with identical sequences but different functional folds "fold mutants". To my knowledge such fold mutants have not yet been found.

A second albeit more exotic solution to the fine-tuning problem appeals to a possible "protein multiverse". The argument here is that the kind of protein structures which we observe are indeed not the only feasible or functional ones. There are in fact other structures which are not only well-folded but also functional. For some reason, evolution, during its intricate dance of maintaining order, structure and function, chose to discard these structures in favor of ones that were more functionally relevant
in this universe. However there is no reason why they could not have been picked in a different universe, where the laws were slightly different. There is another way to think of a protein multiverse; as a set of valleys and peaks where the valleys correspond to different folded structures. Such a metaphor has also been used by physicists to argue that our universe with its own set of fundamental constants corresponds to one local minimum
in this "multiverse landscape", with other universes populating the other dips. Similarly we could imagine a protein multiverse landscape in which different protein folds occupy different valleys; we favor a particular fold only because it inhabits our own valley, but that does not stop other folds from corresponding to the others.

In a different universe, hemoglobin could have folded into a marginally different structure in which it bound not oxygen but some other small ligand like ammonia more efficiently. Such a fold mutant of hemoglobin would be useful to creatures which survive in an ammonia-rich environment (ammonia in fact has a greater temperature range as a liquid compared to water). Or one could imagine a fold mutant of carbonic anhydrase, which catalyzes the conversion of carbon dioxide to bicarbonate at a different pH or a different temperature. Fold mutants of known proteins could have every conceivable property different from their original "correctly" folded counterparts, including shape, size, polarizability and stability. The fold mutants could be exquisitely adopted to living conditions in their parents universe. Their special folds could be stabilized by environments differing
from those found on earth in ionic strengths, hydrogen bonding capabilities and hydrophobicities. For a given protein, this alternative fold could in fact be the lowest in energy and its companion fold found in our universe could be slightly higher in energy.

This kind of speculation immediately suggests two explorations. One is to look for fold mutants in other parts of the universe. This search would be part of the search for extraterrestrial life that has been going on for years. But the point is that if we happen to find fold mutants of existing proteins on other planets or in other inhospitable environments, these mutants would provide powerful support for the solution of the fine-tuning problem. They would tell us that the fine-tuning problem exists only in our narrow-minded anthropocentric imagination, that there could indeed be many folds of the same protein that are robust and functional and that we just happen to inhabit a part of the universe that stabilizes our favorite fold.


The other more readily testable experiment asks if we can produce different functional folds from the same amino acid sequence by varying the experimental conditions. It's of course well-known to crystallographers and protein chemists that slight changes in physicochemical conditions can play havoc with the structure and function of their proteins. But most of the times these slight changes in conditions produce misfolded protein junk. Is there an example of someone slightly (or even radically) varying conditions in a test-tube and producing two different folds of the same protein that are both stable and functional? If there is one I would be very eager to know about it.


On the other hand, if it turns out that it's impossible to find two different functional folds for a single protein, such an observation might well lend credence to the physicists' multiverse with differing fundamental constants. It might well be that under the present values of fundamental constants, it is impossible to stabilize a slightly different protein fold and make it functional. Perhaps only a slight albeit conceptually radical restructuring of the fundamental constants could result in a universe that is friendly to fold mutants. Such a universe would still enable the creation of complex matter through the appropriate combination of the constants, but it would indeed result in life very different from what we know.


The protein multiverse could thus help resolve the fine-tuning problem in protein folding and make biochemists and physicists part of the same multiverse fraternity. More importantly, it could once again reinforce the diversity of creation. One could have different universes with the same fundamental constants but different protein folds or different universe with entirely different combinations of the constants themselves. Take your pick.

If uncovered, such diversity would only echo J B S Haldane's quote that the "universe is not only queerer than we suppose, but it is queerer than we can suppose".

A graceful collapse

ResearchBlogging.org
Vijay Pande's group at Stanford has become well-known for using the collective force of millions of CPUs around the world for simulating protein folding in the project known as Folding@home. One of the enduring challenges in simulating folding has been to sample the long timescales that are common in real-life folding events, and recent breakthroughs have made accessing such time domains realistic. We should expect long protein folding simulations to be within the reach of many non-specialists in the next few years.

In the latest issue of JACS, Pande's group provides an example of such advances by simulating the folding of a 39 residue protein called NTL9. The actual folding time is 1.5 ms so this is a substantially long MD simulation. To achieve this, Pande's group uses Graphic Processor Units (GPUs) of the kind that are found in video game modules. Over the last few years these units have made interesting biological phenomena accessible to chemists. C & EN has a nice article on the increasing use of GPUs for biomolecular simulation.

Pande's group also uses a set of statistical tools called Markov State Models (MSMs) to identify metastable folding states and the transition trajectories between them. MSMs provide a nifty strategy to bridge the results from several short trajectories (rather than running one long one).

What is endearing about the simulation is that that the correct structure doesn't form until much later and then quickly falls in place, like a lost kid suddenly remembering his place in the marching band. As can be seen in the video below, the missing piece of the puzzle is a short C-terminal part of a beta-sheet which seems to linger as part of an alpha helix while the rest of the sheet structure forms. After comfortably waltzing around as a little helical piece for a long time, it seems to suddenly remember its correct identity and snaps and collapses into place as part of the beta sheet. Very nice!



Admittedly, a 39 residue protein is minuscule compared to most typical proteins. But the results provide a neat proof of concept. Importantly, they also show that current force fields with implicit solvent models can be accurate enough for this kind of simulation. Further validation will test these force fields more stringently.

Voelz VA, Bowman GR, Beauchamp K, & Pande VS (2010). Molecular simulation of ab initio protein folding for a millisecond folder NTL9(1-39). Journal of the American Chemical Society, 132 (5), 1526-8 PMID: 20070076

Humans beat computers in predicting protein structures



ResearchBlogging.org
I was going to first describe Rosetta in a post, but a rather cool paper related to the program which appeared in Nature yesterday makes me jump the gun.

In a nutshell, Rosetta tries to predict the structure of proteins from amino acid sequence by inserting fragments from known protein structures and doing many rounds of side chain torsional angle and rigid-body energy optimization. It uses a scoring function to rank the resulting structures that uses empirically derived hydrogen bonding, hydrophobic burial and desolvation terms. Detailed description will have to await the next post since yesterday's paper is not about Rosetta per se.

Instead the paper talks about a program named FoldIt which essentially asks relatively untrained computer gamers to address the protein folding problem. Gamers are asked to tweak, pull, freeze and rotate parts of an incorrectly folded structure to try to twist it into the correct structure. The interface looks like the picture above. Data from a total of 57,000 gamers was pooled. The gamers were driven to solve the problem by the usual incentives of competition and co-operation. Each set of movements would lead to an increase or decrease in a score, with the goal being to find the correct folded structure corresponding to the minimum score. The corresponding set of operations in Rosetta would involve hydrophobic burial, hydrogen bond formation and breakage, helix rotation and other related movements. The project essentially pitted Rosetta versus the gamers.

The results were striking. In a significant number of cases, the gamers actually outdid Rosetta. The reasons are very intriguing and- in an age where computers seem to have unlimited power over our lives- generally testify to the advantages of humans being over computers. For instance in one case, the gamers had to first unravel significant parts of the protein leading to a sharply unfavorable score and then again re-fold it, leading to a correct structure. Rosetta would not attempt the first operation because of the sharp increase in score. This is a classic example of long-term strategy. Unlike computers, humans can make seemingly bad short-term decisions that ultimately lead to good results; we observe this process in many aspects of daily life, from stock market traders taking risks because they see favorable returns later, to politicians making unpopular choices because they think these choices will eventually lead to a popular outcome. Unlike humans though, it is very difficult for a computer program to do long-term planning, and this example illustrates not only the advantages that human intuition can have but also identifies gaps in a program like Rosetta which can possibly be filled.

Another example where the humans outdid the computers was when presented with a set of 10 incorrect structures. Humans generally chose the structure closest to the given structure, whereas Rosetta picked another structure. The main point here is that simple visual clues can sometimes trump complicated decision-making (although they can also mislead). More generally, the results underscored the fact that gut feelings and mere inspection can sometimes lead to successful results.

The one case where the humans did not do as well as Rosetta was in addressing the "classic" protein folding problem, where the challenge was to predict 3D structure from sequence alone. In this case, the sheer amount of conformational space to be searched thwarts success, and there are also no visual cues to guide the process unlike before. The key value of computer approaches which can rapidly pare down the conformational space becomes evident in this example.

So since humans outdid the computer in many cases on the basis of intuition, this must be one super-smart group of biochemists, right? Au contraire! One of the most compelling facts was that most of the gamers in fact not only lacked a formal background or PhD. in biochemistry, but also lacked a formal background in science. Relatively few had college degrees, let alone more advanced ones. For instance there is a profile of a woman in the video below who works in a physical therapist's office, who says that after coming home she feels like a different person when she plays the game. This is great. The examples strikingly illustrates that even untrained humans can possess skills that may be difficult to program into a computer.

It remains to be seen if these results can be extrapolated to large-scale trials, but this very intriguing study perhaps illustrates the general principle that cracking a problem as complex as protein folding is going to require a diverse set of skills, from Monte Carlo searching to gut feelings.



Cooper, S., Khatib, F., Treuille, A., Barbero, J., Lee, J., Beenen, M., Leaver-Fay, A., Baker, D., Popović, Z., & players, F. (2010). Predicting protein structures with a multiplayer online game Nature, 466 (7307), 756-760 DOI: 10.1038/nature09304

Curbing the combinatorial catch

The 'combinatorial explosion' problem generally refers to the difficulty of locating a unique solution to a given problem when the potential space of solutions to be searched is astronomically large. It is found in many areas of science but most notably in protein folding where it takes the name of "Levinthal's Paradox". Biochemist Cyrus Levinthal pointed out in the 60s that if a given sequence of amino acids were to explore every possible conformation for each of its amino acids, even a small protein of 100 amino acid residues or so would take a time longer than the age of the universe to find the correct folded structure.

The paradox is clearly not a paradox since nature has solved the problem of protein folding countless number of times since life began on this planet (this is the protein-centric version of the anthropic principle). Thus, the combinatorial 'problem' is not a problem so far as we know that a robust and tried-and-tested solution exists and in fact has been used by nature to stunning effect. The problem is really to figure out the devilish details of this solution. In the past 30 years or so scientists have employed a battery of experimental and theoretical techniques to tackle the issue. Many important insights have revealed that understanding the factors that dictate the self-assembly of proteins can lead to great insights into the problem. Probably the foremost among these factors is the hydrophobic effect, which productively buries greasy chemical functionalities in the interior of proteins utilizing the multiple driving engines of favorable desolvation, entropic expulsion of water and weak packing-induced interactions. Other important factors ubiquitously used by nature include hydrogen bonds and salt bridges.

The key insight in tackling the problem has been to realize that protein folding or protein-protein interactions or indeed, all the myriad biomolecular interactions that occur in the cellular milieu, do not arise 'by chance'. Once we get past this stumbling block, things make a lot of sense. Chance events undoubtedly keep on happening, but nature preferentially preserves the consequences of certain events. Thus, similar motifs which have been successfully used for certain proteins are used for others. Nature does not need to keep on searching all of conformational space again and again for generating new structures. The analogy would be in designing a new house based on existing structures like bricks, arches and beams rather than designing it from scratch. A Victorian Englishman coined a word for this process of preservation of favorable elements leading to new biological entities a hundred and fifty years ago- natural selection. Thus, the protein folding problem can be immediately demystified when one realizes that natural selection keeps on using recurring motifs to build new structures. Far from being a chance event, the complexity of life can be explained by the re-use of pre-existing structures to build complexity. It may seem highly improbable and miraculous, but Darwin's genius was to provide a mechanism for precisely explaining this illusion of 'design', both on macro and molecular scales. It no longer seems improbable, but instead offers us a tool of incomparable power to peek into the heart of complex biological phenomena.

From a chemist's point of view, natural selection at the molecular level takes the form of the preservation of low-energy conformations of biomolecules that may possess other qualities such as stability, catalytic proficiency and rapid replication. Such chemical entities (think 'DNA') will persist and proliferate and they will be used in multiple designs. Consider coiled-coil structures with their typical seven-residue amino acid motifs or the catalytic triad that cleaves peptide bonds in proteases. Or think of something that's bleedingly simple- the phosphate group which, by virtue of its remarkable qualities of 'transient stability' to hydrolysis, proves to be the perfect connection for life's lego pieces. Once nature hit upon such designs, they could be easily employed in many different structures, dramatically reducing the amount of functional space to be searched. From a chemical perspective, the key property of these favored motifs is self-assembly which is driven by many well-understood physicochemical factors such as the aforementioned hydrophobic effect. Self-assembly, surely one of the greatest inventions of the laws of physics and chemistry, took the problem of the origin of life from miraculous impossibility to tantalizing possibility.

If nature can use pre-existing functionalities to solve the protein folding problem, why can't we do the same? Indeed, many theoretical approaches to protein folding have adopted this kind of approach. Probably the foremost algorithm for predicting protein folding today is a suite of programs called Rosetta which was originally developed by David Baker's group at the University of Washington. In a competition to predict protein structures in 2001, the program did so well that it was compared by a very famous computational chemist named Peter Kollman to Babe Ruth's world record, when even the second-best competitor was woefully lagging behind.

In the next post we will take a look at this program and why it works so successfully.

Impressionist thoughts on Rosetta

Here is Bosco Ho, a postdoc at UCSF comparing Rosetta to the Impressionists, my favorite cabal of artists. Along the way praise and disappointment are also exuded toward GROMOS, an earlier protein modeling program
THE GUTS OF ROSETTA

In the last two CASP meets, David Baker from the University of Washington, using his program Rosetta has come first by a hefty margin in the New Fold category. The success of Rosetta has electrified the protein-folding community.

Yet, there are theorists out there who feel slightly queasy when poking through the innards of Rosetta. Theorists such as Wilfred van Gunsteren, write programs such as GROMOS, which have the richness of 17th century Dutch paintings. Just as Vermeer was fetishistically obsessed with painting every detail of the Dutch bourgeoisie, right down to the hem-line of the chamber-maid's dress, GROMOS is obsessed with modeling every detail of 21st century atomic physics, right down to the quadruple expansion of the electron shells of polarizable atoms. The problem with programs like GROMOS is that they are lumbering giants, bloated programs that devour all the computing that you could ever offer, and still beg for more. Although GROMOS is used for many things, attempts to fold a protein have lurched to a stuttering halt, even after agos of computing time.

Programs like Rosetta, on the other hand, are more like Impressionists paintings, virtuoso dabs of paint that trick the eye into seeing a protein fold in no time at all. For instance, whereas GROMOS fastidiously models all 6 atoms in carbon rings attached to the protein and each atom in the ring is allowed to wobble, Rosetta models the carbon ring as one fat unmovable atom. Water molecules surrounding the protein? No problem, says Rosetta, we'll just ignore them. Rosetta also uses a clever trick by folding similar proteins from different species of animals, and then averaging all the structures to obtain a consensus structure. In reality, when proteins like hemoglobin fold inside your body, they don't get to watch how hemoglobin folds in rats or flies in order to come to a consensus.
Now don't get me wrong; impressionism is my favorite art style, but somehow I am always going to be a little uncomfortable about a program that relies more on statistics than physics to simulate protein folding. I already have this hang up about models in general which I have articulated before. Although modeling reality is what models are supposed to do, ultimately you can still be in for a nasty surprise if you are not paying too much attention to the actual physics and chemistry behind the molecular interactions.

As an aside, I have used Rosetta a little and it can be hideously user-unfriendly. Why the authors never sought to collaborate with a software company who would design a nice GUI for it is something I have never understood. Now in spite of the above rants let me not be misleading here; I think Rosetta is a fantastic program that has achieved some spectacular results reported in places like Nature and Science; perhaps its most stunning achievement was designing an enzyme from scratch that would catalyze a Kemp elimination reaction, a reaction that no other enzyme in nature is known to catalyze. It's just that I think that using it, at least for people who are not members of David Baker's group, might be like flying a highly sophisticated spaceship whose workings are somewhat mysterious. It could be a problem when those ill-understood cumulonimbus (or Romulans) start looming on the horizon.