Field of Science

Beware of von Neumann's elephants using bulldozers to model quarks

The other day I wrote about the late physicist Leo Kadanoff who captured one of the key caveats of models with a seriously useful piece of advice - "Do not model bulldozers with quarks". Kadanoff was talking about the problems that arise when we fail to use the right resolution and tools to model a specific system. While reading Kadanoff's warnings I also remembered one of John von Neumann's equally witty portents for flawed modeling - "With four parameters I can fit an elephant to a curve. With five I can make him wiggle his trunk".

It strikes me that between them Kadanoff and von Neumann capture almost all the cardinal sins of modeling. The other day I was having a conversation about modeling with a leading industrial molecular modeler, and he made the very cogent point that it is imperative to keep the resolution of a particular system and the data it presents in mind when modeling it. My colleague could well have been channeling Kadanoff. This point is actually simple enough to understand (although hard enough to always keep in mind when obeying institutional mandates in a shortsighted environment which thrives on unrealistic short-term goals). 

If you are doing structure based drug design for instance, it's dangerous to try to read too much atomic detail into a 3 angstrom protein-ligand structure. Divining fine details of halogen substitutions, amide flips and water molecules from such a structure can always get you in trouble. If a 3 angstrom structure is the best you have, your optimum strategy would be to try rough designs of molecules - a hydrophobic extension here, a basic amine there - without getting too fine-grained about it. What you should aim for is maximum diversity accessible with minimal synthetic effort - libraries of small peptides might be suitable candidates in such cases. After that let the chemical matter guide you. Once you have a hit, that's when you want to get more detailed, although even then the low resolution of the structure may be at odds with the high resolution of your thinking.

An equally good or even better strategy to adopt in such cases might be a purely ligand-based assault on the structure. There might be similar ligands hitting similar proteins which you might be aware of, or even in case of de novo ligand design you might want to push for purely ligand-based diversity. But this is where you now have to start listening to von Neumann. You may try to fit potential activities of ligands to a few parameters, or build a QSAR model. What you might really be doing however is building not a QSAR model but a house of cards supporting a castle in the air - in other words an overfit model with scant connection to chemically intuitive reality. In that case rest assured - von Neumann's elephant would be quite willing to crash his way in and tear apart your castle.

Kadanoff's admonition to not model bulldozers with quarks is a good admonition for structure-based design. Von Neumann's elephants are good portents to keep in mind for ligand-based drug design. Together the two can hopefully keep you from falling into the abyss and getting crushed under the elephant and the bulldozer.

Linus Pauling's last laugh? Vitamin C might be bad news for mutant colorectal cancer


Linus Pauling holding enough rope to make sure we
can hang ourselves with it if we don't run the right
statistically validated experiments
During the last few decades of his life, Linus Pauling (in) famously began a crusade to convince the general public of the miraculous benefits of Vitamin C for curing every potential malady, from the common cold to cancer. Pauling’s work on ascorbic acid resulted in many collaborations, dozens of papers and at least two best-selling books.

The general reaction to his results and studies ranged from “interesting” to “hey, where are the proper controls and statistical validation?” Over the years none of his work has been definitively validated, but vitamin C itself has continued to be interesting, partly because of its cheap availability and ubiquitous nature in our diet and partly because of its antioxidant properties that seem to many people to be “obviously” beneficial (although there’s been plenty of criticism of antioxidants in general in recent years). Personally I have always put vitamin C in the “interesting and should be further investigated” drawer, partly because oxidation and reduction are such elemental cellular phenomena that anything that seeks to perturb such fundamental events deserves to be further looked at.

Now here’s an interesting paper in Science that validates the potential benefits of ascorbic acid in a very specific but well-defined case study. It’s worth noting at the outset that the word ‘potential’ should be highlighted in giant, size 24 bold font sizes. The authors who are part of a multi-organization consortium look at the effects of high doses of the compound on colorectal cancer cells with mutations in two ubiquitous and important proteins – KRAS and BRAF. KRAS and BRAF are both part of key signaling networks in cells. Mutations in both of these proteins are seen in up to 40% of all cancers, so both proteins have unsurprisingly been very high-profile targets of interest in cancer therapy for several decades. The mutation is additionally important because it also turns out that cancers with these mutations show poor response to anti-EGFR therapies.

One of the hallmarks of cancer cells which has been teased out in fascinating detail in the last few years is their increased metabolism and especially their dependence on glucose metabolism pathways such as glycolysis that allows them to feed hungrily on this crucial substance. The current study took off from the observation that a glucose transporter protein called GLUT1 is overexpressed in these mutant cancer cells. Incidentally this transporter protein is also involved in transporting vitamin C, but in its oxidized form (dehydroascorbate – DHA). Presumably the authors put two and two together and wondered if ascorbate might be more rapidly absorbed by the mutant cancer cells and mess up the oxidation-reduction machinery inside.

It turns out that it does. Firstly, the authors confirmed by the addition of reducing agents that it’s the oxidized form of vitamin C that interferes with the cancer cells’ survival. Secondly, they looked at mutant vs wild-type cells and found that the mutant cells are indeed much more efficient at ascorbate uptake. Thirdly, they looked at various markers for cell death like apoptosis signals and found out that these were indeed more pronounced in the KRAS-BRAF mutant cells (addition of a reducing agent rescued these cells, again attesting to the function of DHA rather than reduced vitamin C). Fourthly, mice with known as well as transgenic KRAS mutations showed favorable tumor reduction when vitamin C was intravenously administered.

Fifth and most interesting, they performed protein metabolite analysis of the cells’ machinery after treatment with vitamin C and found that there was a significant accumulation of chemical intermediates which serve as substrates for the enzyme glyceraldehyde-3-phosphate dehydrogenase (GAPDH). GAPDH is a central enzyme of the glycolytic pathway and its inhibition would unsurprisingly lead to cell starvation and death. Lastly, they were able to make a statement about the mechanism of action of vitamin C on GAPDH by determining that it might interfere with post-translational modification of the protein and NAD+ depletion.

The authors end with some ruminations on the history of vitamin C therapy for cancer and the usual qualifications which should apply to any such study,. As they note, vitamin C has a checkered history in the treatment of cancer but most studies which failed to show benefits only involved large oral doses of the vitamin (Pauling himself was rumored to ingest up to 50 g of the substance a day). Intravenous administration however has suggested that far higher doses may be required for effective results. And of course, this study was done in mice, and time after time we have seen that such studies cannot be measurably extrapolated to human beings without a lot of additional work, so you should pause a bit before you rush off and try to inject yourself with Emergen-C solution.

Nonetheless, I think the detail-oriented and relatively clear nature of the study makes it a good starting point. Google searches of vitamin C and colorectal cancer bring up at least a few tantalizing clues as to its potential efficacy (along with a lot of New Age, feel-good piffle). As usual the key goal here is to separate out the wheat from the chaff, the sloppy anecdotal evidence from the careful statistical validation and the detailed mechanistic rationales from the stratospheric theorizing. When the dust settles we would hopefully have a clearer picture. And who knows, maybe the ghost of Linus Pauling might then even allow himself the last laugh, or at least an imperceptible smile.

Physicist Leo Kadanoff on reductionism and models: "Don't model bulldozers with quarks."

I have been wanting to write about Leo Kadanoff who passed away a few weeks ago. Among other things Kadanoff made seminal contributions to statistical physics, specifically the theory of phase transitions, that were undoubtedly Nobel caliber. But he should also be remembered for something else - a cogent and very interesting attack on 'strong reductionism' and a volley in support of emergence, topics about which I have written several times before.

Kadanoff introduced and clarified what can be called the "multiple platform" argument. The multiple platform argument is a response to physicists like Steven Weinberg who believe that higher-order phenomena like chemistry and biology have a strict on-on-one relationship with lower-order physics, most notably quantum mechanics. Strict reductionists like Weinberg tell us that "the explanatory arrows always point downward". But leading emergentist physicists like P W Anderson and Robert Laughlin have taken objection to this interpretation. Some of their counterarguments deal with very simple definitions of emergence; for instance a collection of gold atoms have a property (the color yellow) that does not directly flow from the quantum properties of individual gold atoms.

Kadanoff further revealed the essence of this argument by demonstrating that the Navier-Stokes equations which are the fundamental classical equations of fluid flow cannot be accounted for purely by quantum mechanics. Even today one cannot directly derive these equations from the Schrodinger equation, but what Kadanoff demonstrated is that even a simple 'toy model' in which classical particles move around on a hexagonal grid can give rise to fluid behavior described by the Navier-Stokes equations. There clearly isn't just one 'platform' (quantum mechanics) that can account for fluid flow. The complexity theorist Stuart Kauffman captures this well in his book "Reinventing the Sacred".



Others have demonstrated that a simple 'bucket brigade' toy model in which empty and filled buckets corresponding to binary 1s and 0s (which in turn can be linked to well-defined quantum properties) that are being passed around can account for computation. Thus, as real as the electrons obeying quantum mechanics which flow through the semiconducting chips of a computer are, we do not need to invoke their specific properties in order to account for a computer's behavior. A simple toy model can do equally well.

Kadanoff's explanatory device is in a way an appeal to the great utility of models which capture the essential features of a complicated phenomenon. But at a deeper level it's also a strike against strong reductionism. Note that nobody is saying that a toy model of classical particles is a more accurate and fundamental description of reality than quantum mechanics, but what Kadanoff and others are saying is that the explanatory arrows going from complex phenomena to simpler ones don't strictly flow downward; in fact the details of such a flow cannot even be truly demonstrated.

In some of his other writings Kadanoff makes a very clear appeal based on such toy models for understanding complex systems. Two of his statements provide the very model of pithiness when it comes to using and building models:

"1. Use the right level of description to catch the phenomena of interest. Don't model bulldozers with quarks.

2. Every good model starts from a question. The modeler should aways pick the right level of detail to answer the question."


"This lesson applies with equal strength to theoretical work aimed at understanding complex systems. Modeling complex systems by tractable closure schemes or complicated free-field theories in disguise does not work. These may yield a successful description of the small-scale structure, but this description is likely to be irrelevant for the large-scale features. To get these gross features, one should most often use a more phenomenological and aggregated description, aimed specifically at the higher level. 

Thus, financial markets should not be modeled by simple geometric Brownian motion based models, all of which form the basis for modern treatments of derivative markets. These models were created to be analytically tractable and derive from very crude phenomenological modeling. They cannot reproduce the observed strongly non-Gaussian probability distributions in many markets, which exhibit a feature so generic that it even has a whimsical name, fat tails. Instead, the modeling should be driven by asking what are the simplest non-linearities or non-localities that should be present, trying to separate universal scaling features from market specific features. The inclusion of too many processes and parameters will obscure the desired qualitative understanding."

This paragraph captures as well as anything else why chemistry requires its own language, rules and analytical devices for understanding its details and why biology and psychology require their own similar implements. Not everything can be understood through quantum mechanics, because as you try to get more and more fundamental, true understanding might simply slip away from between your fingers.

RIP, Leo Kadanoff.

(Ir)rational drug design and the history of 20th century science


Here is an excellent overview of the hopes and foibles of "rational" drug design by Brooke Magnanti (Hat tip: Pete Kenny) which touches on several themes and names that would be familiar to those in the field: Ant Nicholls and OpenEye, Dave Weininger and Daylight fingerprints, Barry Werth's "The Billion Dollar Molecule" and Vertex, the inflated hopes of structure-based design, cheminformatics and screening etc. 

Those who are heroic survivors of that period would probably start with looking back with dewey eyes, followed by groans of disappointment. The bottom line in that article and several similar ones is that rational drug design and all that it entails (crystallography and molecular modeling in particular) has clearly not lived up to the hype. It's also clear that the swashbuckling scientists portrayed by Werth in his book for instance were more brilliant than successful. It's a tape of hope and woe that has played before, over and over again in fact.

It's clear that much of the faith in rational drug design until now has had a healthy component of irrational exuberance to it. Looking back at the inflated expectations of the 1980s and early 90s for designing drugs atom by atom, followed by the disappointing failures and massive attrition which rapidly succeeded these expectations, makes me wonder what it was exactly that got everyone into trouble. There was a constellation of factors of course, but the historian of science in me thinks that a major part of at least the psychological (and by extension, organizational) aspects of the issue have to deal with the stupendous successes of twentieth century science in generating a mountain of optimism which skeptics are still trying to chip away at.

It's quite clear that as far as scientific progress goes, the 20th century was the mother of all centuries. Very significant scientific advances (Newton, Maxwell, Darwin, Mendel) had undoubtedly occurred in earlier times, but the sheer rate at which science advanced in the last one hundred years far outstripped scientific progress in all previous centuries. Just consider the roster of both idea-based and tool-based scientific revolutions that we witnessed in the past century: x-rays, the atomic nucleus, relativity, quantum mechanics, nuclear fission, the laws of heredity, the structure of biomolecules, particle physics, lasers, computers, organic synthesis, gene editing...and we are just getting warmed up here.

By the 1980s this amazing collection of scientific gems had reached a crescendo, especially in the biomedical sciences. The rise of recombinant DNA technology, protein structure determination, and improved hardware, software and visualization virtually ensured that scientists started feeling very good indeed about designing drugs to block particular proteins at the molecular level. Philosophically too they were highly primed by the astounding reductionist successes of the past one hundred years. After all reductionism had uncovered the cosmic microwave background radiation from the Big Bang, given us the structure of elemental life proteins like hemoglobin and the photosynthetic complex, split the atom, doubled the number of transistors on a chip in eighteen months and taught us how to copy and paste genes. Designing drugs would be a natural extension, if not a job for graduate students, after all this success.

But what happened instead was that both scientifically and philosophically we ran into a wall. What we found out scientifically was that we still understand only a fraction of the complexity of biological systems that we need to for perturbing them with the fine scalpels of small organic molecules. Philosophically we found out that biological systems are emergent and contingent, so all the reductionist success of the past century is still not enough to understand them. In fact beyond a certain point reductionism would fundamentally put us on the wrong track. The past hundred years made us believers in Moore's Law, but what we got instead was Eroom's Law. Moore's Law is what reduces my running time from 12 mins/mile to 8:30 mins/mile in a year. Eroom's Law is what keeps it from reducing much further. Exponential technological success is not axiomatic and self-fulfilling.

I thus see a very strong influence of the success of twentieth century science in steering the wildly optimistic hopes of drug discovery scientists beginning in the 1980s. Hopefully we are wiser now, but institutional forces and biases still keep us from improving on our failures. As Pete Kenny says in his post for instance, obsession with specific technologies rather than a combined application of several technologies still biases scientists and managers in biotech and pharmaceutical organizations. The rise and ebb (did you just say "rise"?) of economic forces makes the job environment unstable and discourages scientists from pushing bold ideas that promise to break free from reductionist approaches. And much of our science is still based on sloppy theorizing without proper recourse to statistics and controls, not to mention an unbiased look at what the experiments truly are and are not telling us. 

Santayana told us that we are condemned to relive history if we forget it. But when it comes to the promises of rational drug design, what we should do perhaps is to purge our minds of the successes of the 20th century and remember Francis Bacon's exhortation from the 16th century instead: "All depends upon keeping the eye steadily fixed on the facts of nature. For God forbid that we should give out a dream of our own for a pattern of the world."

Image: "Cognition enhancer" (Source: Brooke Magnanti, Garrett Vreeland)

"The Hunt for Vulcan": Theory, experiment, and the origin of scientific revolutions

Joseph Urbain La Verrier: The force of his personality
and his spectacular prediction of Neptune solidified
faith in the existence of Vulcan
In his book "The Hunt for Vulcan", MIT science writing professor Thomas Levenson tackles one of the most central questions in all of science - what do you do when a fact of nature disagrees with your theory? In this particular case the fact of nature was an anomaly in the orbit of Mercury around the sun. The theory was Newton's successful theory of gravitation which had reigned supreme for two hundred years in explaining the motion of everything from rocks to the moon. Levenson’s book looks at this question through the lens of an important case study. His writing is clear, often elegant and impressionistic, and he does a good job driving home the nature of science as a human activity with all its human triumphs and follies.

The physical entity invoked to explain the anomalies in Mercury's orbit - a small planet close to the sun which would usually be too small and intensely illuminated by the sun to be seen - was called Vulcan. The idea was that Vulcan's gravitational tug on Mercury would cause its orbit to stray from the expected path. The hypothesis had much merit to it since it was similar theorizing about the anomalies in the predicted orbit of Uranus that had resulted in the discovery of Neptune. The man who proposed the theories of both Neptune and Vulcan was Joseph Urbain Le Verrier, the most important French astronomer of his day and one of the most important of the 19th century. The successful prediction of Neptune and its dazzlingly swift observational validation was a resounding tribute to both Le Verrier’s acumen and to Newton’s understanding of the universe. Not surprisingly Le Verrier's prediction of Vulcan was taken seriously.

The book recounts how partly because of past successes of Newton's theories and partly because of the force of Le Verrier’s personality astronomers spent the next one hundred years unsuccessfully looking for Vulcan. Spectators included a host of well-known astronomers and amateurs, including Thomas Edison. The search was peppered by expeditions to exotic places like Rhodesia and Wyoming. Occasionally the newspapers would ridicule Vulcan-chasers, but none could disprove its evidence conclusively. This fact raises an important point: As far as scientific theories go Vulcan was a good theory since it was testable, but because its existence really strained the limits of astronomical technique as it existed during the time, it did not really satisfy the criteria for being a cleanly falsifiable theory. This led to the Vulcan hypothesis having enough wiggle room for people to get away with explaining away the lack of observation as bad technique or faulty equipment.

As Levenson describes in the latter half of the book, the culmination of the hunt for Vulcan came in the early half of the twentieth century with Einstein’s theory of relativity which did away with Vulcan for good. Levenson spends a good deal of time on Einstein's background and his mathematical preparation; there's a lucid description of the special theory of relativity. Vulcan was almost an afterthought in Einstein's intellectual development, but when he realized that his own theory could explain Mercury's anomalous orbit as an effect of the curvature of spacetime, the realization left him feeling like "something had snapped inside him". When finished his general theory of relativity demonstrated one of the most fascinating features of scientific discoveries – sometimes tiny anomalies in observation point not just to the reworking of an existing theory but a complete overhaul of our understanding of nature. In this case the dramatic change was an appreciation of gravity not as a force but as a curvature of spacetime itself.


It is also instructive to apply lessons from Vulcan to my own fields of drug discovery and biochemistry. Often when a drug does not work it seems convenient to invoke the existence of hitherto unobserved entities (specific proteins, artifacts, side products from organic reactions etc.) to explain the anomalies or failures. Vulcan tells us that while it is prudent to look for these entities experimentally, it's also worth giving a thought to how their existence might be explained by tweaks - or in rare cases significant overhauls - of existing theories of biological signaling or drug action. This might especially be true in case of neurological disorders like Alzheimer's disease where the causes are ill-understood and the underlying theories (the amyloid hypothesis for instance) are constantly being subjected to revision.

Levenson’s book is a tribute to how science actually works as opposed to how it's thought to work. It's also a good instruction manual for how science works when experiment disagrees with theory. In such cases the theory can then be slightly amended, radically amended or replaced. In Vulcan’s case Newtonian gravity was not really replaced, but the amendment required was so drastic that it led to a new epoch in our view of our cosmos. The story of Vulcan is a story for our scientific times.

When Skynet finally gets here it will be analog, not digital


That's historian of technology George Dyson contemplating the dangers of analog intelligence rather than digital intelligence in John Brockman's latest "big question" compilation, this time on AI. 

The real curveball however might be captured by another one of Dyson's quote: "My real worry is not that machines will become too intelligent, it's that humans will become too dumb."

The 2015 Medicine Nobel Prize is a tribute to drug discovery, chemistry and traditional medicine


It's very gratifying to see this year's Nobel Prize for Physiology or Medicine awarded to three scientists (William Campbell, Satoshi Omura and Youyou Tu) who have contributed to the discovery of novel drug substances that have decided benefited the lives of millions of human beings and animals. The prize was awarded to the discovery of avermectin and artemisinin. Avermectin cured nematode infections in millions of livestock animals including cows and pigs and its derivative ivermectin can cure river blindness, a disease primarily affecting poor people. Artemisinin can halt malaria in its tracks and contribute to substantial reduction of mortality.  It is hard to thing of a discovery which satisfies Alfred Nobel's stipulation of providing the "greatest benefit to mankind" more than this one. 

It's particularly gratifying to see pharmaceutical research being recognized with the prize. The last time the prize was awarded to drug discovery was in 1988, to Gertrude Elion and James Black. Both Black and Elion worked in the private sector. William Campbell who is one of this year's recipients also worked in the private sector at Merck in the 1980s when his team discovered avermectin. The strain of bacterium (Streptomyces) which yielded the drug had been discovered by Satoshi Omura's team in Tokyo. The collaboration was a perfect example of the public and the private sector working together to bring medical benefits to humanity. Campbell's team at Merck also discovered the even more potent derivative of avermetic, ivermectin, and this was discovered purely synthetically. It's also worth noting that Merck made avermectin available for free to treat river blindness in a generous gesture.

Youyou Tu's work on artemisinin in China during the 1960s is another lesson in pharmaceutical discovery. In the 1960s millions of poor Chinese were dying of malaria. The frontline drug, chloroquine, was being increasingly ineffective. Remarkably, Tu found an obscure reference to a document on traditional Chinese medicine written in 340 BC which described the potentially healing antimalarial powers of a herb steeped in cold water. By testing more than 200 extracts Tu discovered artemisinin. The initial paper in 1979 elicited skepticism and smug dismissal (partly because it came from communist China), but over the next thirty years artemisinin turned into a drug of choice for treating malaria.

Both artemisinin and avermectin exemplify the power of old-school chemistry and microbiology, a nexus blazed by antibiotic pioneers like Alexander Fleming and Selman Waksman, and one which has been largely forgotten in the last thirty years. In addition both compounds are natural products, and the prize underscores the value of drugs from nature (which already make up about fifty percent of all marketed drugs). Artemisinin in particular is also a vigorous validation of the potential of traditional Chinese (and other) medicine. This kind of medicine is completely different from homeopathy since it involves the use of actual chemical substances and herbal extracts. The story of artemisinin clearly indicates that we need to pay much more attention to forgotten examples from traditional Asian medicine and subject them to scrutiny.

Let's make no mistake about it: Today's Nobel Prize should thus be a resounding tribute to the power and humanity of pharmaceutical research. These days we are justifiably reluctant to associate that second adjective with anything to do with pharma; we are justifiably indignant at the price hikes, the off-label marketing and the other shenanigans which drug companies and CEOs sometimes indulge in. Yet this prize demonstrates that pharma, even Big Pharma, can do untold good, not just in discovery but in philanthropy. The distinguishing factors which characterized Merck in the 1980s were a strong focus on basic research (when my PhD advisor worked there during that time, people called it the "Merck University") and leadership under CEO Roy Vagelos - a rare, distinguished scientist at the helm who was a member of the National Academy of Sciences. Merck's success and largesse from the 1980s should be a role model for drug companies today- including Merck itself.

Interestingly, today's Nobel Prize is also great tribute to research which is decided non curiosity-driven and non-accidental. As this report on the discovery of avermectin from Campbell and his team says, "The discovery of the avermectin family of compounds was by no means serendipitous" (hat tip: Amanda Yarnell). I am as big a fan of curiosity-driven research as anyone else, but both avermectin and artemisinin were discovered with the express goal in mind of curing human disease. There is much to be said for this kind of deliberate applied research, supported by generous funding and vision at the top.

The recognition is also a tribute to the power of synthetic chemistry to create lifesaving substances that did not exist on earth before. Ivermectin is a purely synthetic derivative of avermectin created by using a chemical reaction (catalytic hydrogenation - itself awarded a chemistry Nobel Prize) that itself did not exist before.

So there it is: A Nobel Prize awarded to the benefits of private and public pharmaceutical research, awarded to the power of synthetic chemistry, awarded to the great potential of traditional medicine, and awarded to a female scientist. There's few Nobel Prizes that present such a happy constellation of qualities in one little package. Alfred Nobel would have been pleased.

Image: Nobelprize.org

George Whitesides to chemists: Move away from the molecule

This is an updated version of a post I wrote about an article by George Whitesides that exhorted chemists to move beyond the molecule. It was provoked in part by several recent discussions I have had about how chemists can have a broader impact on other disciplines as well as on the public appreciation of their own discipline.

Chemist George Whitesides probably does not consider himself a philosopher of chemistry, but he is rapidly turning into one with his thought-provoking pronouncements on the future of the field and its practitioners. His most recent rumination on the topic was a piece in the Annual Reviews of Analytical Chemistry provocatively titled "Is the Focus on Molecules Obsolete?" where he uses analytical chemistry as an excuse to really pontificate on the state and progress of chemical science. Along the way he also has some valuable words of advice for aspiring chemists.
Whitesides's main message to young chemists is to stop focusing on molecules. Given the nature of chemistry this advice may seem strange, even blasphemous. After all it's the molecule that has always been the heart and soul of chemical science. And for chemists, the focus on molecules has manifested itself through two important activities - structure determination and synthesis. The history of chemistry is essentially the history of finding out the structure of molecules and of developing new and efficient methods of making them. Putting these molecules to new uses is what underpins our modern world, but it was really a secondary goal for most of chemistry's history. Whitesides tells us that the focus of the world's foremost scientific problems is moving away from composition to use, from molecules to properties. Thus the new breed of chemists should really focus on creating properties rather on creating molecules. The vehicle for Whitesides's message is the science and art of analytical chemistry which has traditionally dealt with developing new instrumentation and methods for analyzing the structure and properties of molecules.
Of course, since properties depend on structures, Whitesides is not telling us to abandon our search for better, cleaner and more efficient techniques of synthesis. Rather, I see what he is saying as a kind of "platform independence". Let's take a minute to talk about platform independence. As the physicist Leo Kadanoff has demonstrated, you can build a computer by moving around 1s and 0s or by moving around buckets of water, with full buckets essentially representing 1s and empty ones representing 0s. Both models can give rise to computing. Just like 1s and 0s simply turn out to be convenient abstract moving parts for building computers, similarly a certain kind of molecule should be seen as no more than a convenient vehicle for creating a particular property. 
For practical applications that property can be anything from "better stability in whole blood" to "efficient capture of solar energy" to "tensile strength". The synthesis of whatever molecular material gives rise to particular properties is important, but it should be secondary; a convenient means to an end that can be easily replaced with another means. As an example from his own childhood, Whitesides describes a project carried out in his father's company in which his job was to determine the viscosities of different coal-tar blacks. The exact kind of coal-tar black was important, but what really counted was the property - viscosity - and not the molecular composition.
A focus on properties is accompanied by one on molecular systems instead of on individual moleculessince often it's a collection of different, diverse molecules rather than of a single type that gives rise to a desired property. What kind of problems will benefit from a molecular systems approach? Whitesides identifies four critical ones; health care, environmental management, national security and megacity management. We have already been living with the first three challenges, and the fourth one looms large on the horizon.
Firstly, health care. Right now most of the expenditure on health care, especially in the United States, is on end-of-life care. Preventative medicine and diagnostics are still relegated to the sidelines. One of the most important measures to drive down the cost of healthcare will be to focus on prevention, thus avoiding the expensive, all-out war that is often waged - and lost - on diseases like cancer during their end stages. Prevention and diagnostics are areas where chemistry can play key roles. We still lack methods that can quickly and comprehensively analyze disease markers in whole blood, and this is an area where analytical and other kinds of chemists can have a huge impact. And no method of diagnostics is going to be useful if it's not cheap, so it's obvious that chemistry will also have to struggle to minimize material cost, another goal which it has traditionally been good at addressing, especially in industry.
Secondly, the environment. We live in an age when the potentially devastating effects of climate change and biodiversity loss demand quick and comprehensive action. Included in this response will be the ability to monitor the environment, and to relate local monitoring parameters to global ones. Just like we still lack methods to analyze the composition of complex whole blood, we also lack methods to quickly analyze and compare the composition of the atmosphere, soil and seawater in different areas of the world. Analyzing heterogeneous systems with different phases like the atmosphere is a tricky and quintessentially chemical problem, and chemists have their work cut out in front of them to make such routine analysis a reality.
Thirdly, national security. Here chemists will face even greater challenges, since the solutions are as much political and social as they are scientific. Nonetheless, science will play an important role in the resolution of scores of challenges that have to be met to make the world more secure; these include quickly analyzing the composition of a suspicious liquid, solid or gas, unintrusively finding out whether a particular individual has spent time in certain volatile parts of the world or has been handling certain materials, and using techniques to track the movements of suspicious individuals in diverse locations. Chemistry will undoubtedly have to interface with other disciplines in addressing these problems and questions of privacy will be paramount, but there is little doubt that chemists have traditionally not participated much in such endeavors and need to step up to the plate in order to address what are obviously important security issues.
Fourthly, megacities. As we pick up speed and move into the second decade of the twenty-first century, one of the greatest social challenges confronting us is how to have very large, heterogeneous populations ranging across diverse levels of income and standards of living co-existing in peace over vast stretches of land. This is the vision of the megacity whose first stirrings we are already witnessing around the world. Among the problems that megacities will encounter will be monitoring air, water and food quality (vida supra). A task like analyzing the multiple complex components of waste effluent, preferably with a readout that quantifies each component and assesses basic qualities like carcinogenicity would be invaluable. There is no doubt that chemists could play an indispensable role in meeting such challenges.
The above discussion of major challenges makes Whitesides's words about moving away from the molecule clear. The problems encompassing health care, national security and environmental and megacity management involve molecules, but what they really are are collages resulting from the interaction of molecules with other scientific entities, and with the interaction of chemists with many other kinds of professional scientists and policy makers. In one sense Whitesides is simply asking chemists to leave the familiar environment of their provincial roots and diversify. What chemists really need to think of is molecules embedded in a broad context involving other disciplines and human problems.
Part of the challenge of addressing the above issues will be the proper training of chemists. The intersection of chemistry with social issues and public policy demands interdisciplinary and general skills, and Whitesides urges chemists to be trained in general areas rather than specialized subfields. Courses in applied mathematics and statistics, public policy, urban planning, healthcare management and environmental engineering are traditionally missing from chemistry curricula, and chemists should branch out and take as many of these as is possible within a demanding academic environment. It is no longer sufficient for chemists to limit themselves to analysis and synthesis if they want to address society's most pressing problems. And at the end of it they need not feel that a movement away from the molecule is tantamount to abandoning the molecule; rather it is an opportunity to press the molecule into interacting with the human world on a canvas bigger than ever before.

Identical ligands, unrelated proteins, similar energies - When language collides with the facts of nature

Recognition of aromatic rings by two very different
mechanisms but through similar binding energies
Over the years chemists have come up with many different ways to talk about the structure and energetics of molecules and especially to compare these parameters between various compounds. Doing this comparison is not just an academic exercise; for example, knowing which drug molecules are ‘similar’ or ‘different’ can be the deciding factor in picking one drug over another. It is also crucial for knowing the kinds of side effects that drugs can induce by interacting with off-target proteins.

Unfortunately the application of these simple descriptions to matters of molecular description is a very good example of what happens when language collides with fuzzy, ill-defined facts in nature. ‘Similarity’ is a classic example. When you are talking about two drugs being similar for instance, are you talking about their similarity purely in terms of molecular structure (which itself can be defined in many different ways), or their similarity in terms of their effects on cancer cells, or their similarity to engage a common protein target in the body, or through similar side effects? Clearly there are many different ways to define similarity and all these ways are subjective to a large extent.

But there is a problem with applying language to chemical concepts even at a very limited and basic level. A great example of this conundrum is hinted at by a paper from Brian Shoichet’s group at UCSF that just came out in the journal ACS Chemical Biology. The paper asks a very fundamental question: Do identical small molecules or ligands bind to very different proteins? The question in fact goes deeper: How do you define similarity and differences between various proteins to begin with?

To investigate this question, the authors consider 59 ligands bound to 119 different proteins in the PDB. Many bind with high affinity, ranging from low nanomolar to mid micromolar. What the study does is to classify these protein-ligand pairs into three groups. The first group consists of pairs in which the same atoms in identical ligands bind to similar or identical residues in different proteins. The second group consists of the same ligand atoms in identical ligands binding to similar kinds of residues (hydrophobic, positively charged etc.). The third group in a sense is the most interesting since it involves identical ligands binding to completely different proteins; in these cases the binding involves neither similar ligand atoms nor similar protein environments.

The authors find that a good two thirds of the set of protein-ligand pairs involve identical ligands binding to proteins with dissimilar residues. In addition, half of these involve ligands binding to proteins with completely different environments. There is thus no ‘pattern-matching code’ for the same ligand binding to different proteins.

Why do identical ligands bind in very different protein environments? The simple reason is because chemical binding is to a large extent a non-specific process, and there are many ways to skin the protein-ligand cat. Hydrophobic groups bump into hydrophobic groups, positively charged groups interact with negative charged ones and polar atoms snuggle up against other polar atoms. As the authors say:
"A reason why there is no simple code for ligand recognition among binding sites is that proteins have found multiple, at least superficially unrelated ways to recognize most common ligand groups. Thus, cationic amines can be recognized both by anionic residues such as aspartate or glutamate, but they can also be recognized by cation-Pi interactions. Nucleotide phosphates can be recognized by cationic residues such as arginines, but recognition by main chain amide nitrogens in a P-loop is also common. Ligand aromatic groups can stack with tyrosines, phenylalanines and tryptophans, but they can also form cation-Pi interactions many other variations might be mentioned."
But sometimes hydrophobic groups can also snuggle up against polar atoms or poke out into solvent and polar groups can nestle into hydrophobic pockets to various extents, simply because the other atoms in the ligand compensate for such uneasy alliances by forming favorable interactions. This can lead to the same ligands binding to very different protein atoms. As I mentioned in a previous post, atoms end up somewhere simply because they can. The differential placement of atoms in protein pockets is reflected in the different binding affinities that the authors see in their set.

From an evolutionary viewpoint this observation is very interesting. Protein-small molecule binding was constrained during evolution by the basic chemistry and physics of binding on one hand and by the damage incurred by too much non-specific binding on the other (as an extreme case, if every small molecule bound to every protein, there would be way too much noise and biological signaling networks would be effectively impossible). Thus there had to be a balance between promiscuity and specificity. Nature achieved this balance by tuning the affinity of small molecules for proteins over a wide range and by making sure that even weak affinity could translate to significant biological effects.

Unfortunately these are precisely the affinities that we ourselves want to finely tune in a drug discovery program and as the paper shows, this is always going to be an uphill battle because of the multitude of interactions and the lack of correlation between ligand and binding pocket structure (one conclusion from the paper is that you cannot always predict new targets for known ligands simply by computationally comparing binding sites).

But on another level I think this problem also speaks to the paucity of the language that we have for describing binding affinity and molecular interactions in general. Our metric for similarity in this case is the presence of similar ligand atoms binding to similar protein atoms. But nature can use another very simple measure of similarity – similarity in binding energy. It is not unreasonable to say that a ligand binds similarly to two proteins if it exhibits a similar binding affinity to both of them. And this binding affinity need not even be very different since even a few kcal/mol difference in binding energy can translate to a thousand fold difference in actual affinity (say from micromolar to nanomolar). Thus, what we call dissimilar binding may actually be judged as quite similar by nature. Consider the picture at the top of the post for instance: an aromatic ring can interact with a protein through either a stacking interaction with an aromatic amino acid or through a cation-pi interaction with a positively charged amino acid. The two interactions look very different, and yet they involve the same binding affinity. 

All this goes back to something we mentioned before: Similarity is in the eye of the beholder, and what our eye sees as squiggly lines of ligands and protein residues on a computer screen, nature sees simply as thermodynamics, kinetics and quantum mechanics, and all of it lying on a continuum. We might be dismayed to know that the same ligand is binding to very different proteins, but this is because nature may not be regarding them as very different to begin with. To figure out protein-ligand binding then, we may have to see things from the point of view of nature rather than that of our impoverished language.