Field of Science

Steering library bias toward A2A adenosine receptor ligand discovery

ResearchBlogging.org
The A2A adenosine receptor is an important GPCR, well-known for binding caffeine. Adenosine receptors are emerging as relevant drug targets for a variety of disorders including Parkinson's disease, and there is interest in discovering new ligands that bind to them. Among adenosine receptor subtypes, the A2A receptor is one of the few GPCRs whose crystal structure is available. Thus the A2A is amenable to structure-based design efforts, and virtual screening is an especially attractive endeavor in this regard.

In the present report, a team of researchers from NIH and UCSF led by Brian Shoichet, John Irwin and Kenneth Jacobson use virtual screening to discover new ligands for the A2A. There are several points to note here. The authors use the ZINC library of drug like molecules to dock about a million and a half compounds into the binding pocket of the A2A crystal structure. They pick the best-scored 500 (0.035% of the total) ligands and investigate their fit in the binding site. Using criteria like electrostatic and VdW complementarity and novelty of chemotype, they finally select 20 of these 500 hits and test them in assays. Out of these 20, 7 inhibited binding by more than 40% at 20 μM concentration, thus constituting a hit rate of 35%. While the compounds formed the same kinds of interactions as some other A2A ligands, they were also relatively diverse in structure. The ligands were also tested in aggregation-based screens to determine that their activity was not a spurious artifact of aggregation-based inhibition.

This is a pretty good hit rate. Generally virtual screening campaigns are lucky to have a hit rate of a few percent. Curiously, the authors also found a similarly high hit rate during a past VS campaign against the well-known β2 adrenergic receptor. What could be responsible for this high hit rate against GPCRs? The reasons are interesting. One reason could be that GPCRs are very well adapted to bind small molecules in compact pockets, enclosing them and forming many kinds of productive interactions. But more intriguingly, as the authors have noted earlier, there is "biogenic bias" in favor of certain target-specific chemotypes in commercial libraries that are screened, both during VS as well as HTS. This in turn reflects the biases of medicinal chemists in picking and synthesizing certain kinds of chemotypes based on the importance of drug targets and past successes in hitting these targets. GPCRs clearly are enormously important, and GPCR-friendly ligand chemotypes thus constitute a large part of screening libraries. These chemotypes are much more prevalent than those for kinases or ion channels for instance.

This observation has both positive and negative implications. The positive implication is that one is likely to keep finding high hit rates for GPCRs using VS. However, the negative implication is that one is also going to be constrained by biogenic bias, and this might preclude finding more diverse and entirely novel subtypes. Thus, while VS campaigns for GPCRs might find a good number of hits, the novelty of these hits might not always be satisfying. One other quite intriguing point emerging in this study is that the kind of hits found (agonist, inverse agonist, antagonist etc.) reflects the ligand which the target structure used for VS is co-crystallized with. Thus the A2A houses an antagonist in the binding site, leading to a preponderance of antagonists in the top docking hits. Indeed, agonists ranked abysmally low in the list.

GPCR ligand discovery is one of the most important goals in drug discovery. This and other similar studies demonstrate that, with all its caveats, VS can be productively used to mine for new GPCR drugs.

Carlsson, J., Yoo, L., Gao, Z., Irwin, J., Shoichet, B., & Jacobson, K. (2010). Structure-Based Discovery of A2A Adenosine Receptor Ligands Journal of Medicinal Chemistry DOI: 10.1021/jm100240h

It's (not) the mutation, stupid

ResearchBlogging.org
Cancer has emerged as a fundamentally genetic disease, where mutations in genes cause cells to go haywire. Yet, finding out exactly which mutations are responsible for a certain type of cancer is a daunting task. A recent report in Nature which details the cataloging of tens of thousands of mutations in tens of thousands of tumors illustrates the merits and dangerous pitfalls of such an approach.

The article talks about the International Cancer Genome Consortium (ICGC), formed in 2008, whose task is to coordinate an international effort spread across different countries, where every country has the responsibility of documenting significant mutations in certain types of cancer. For instance, the US is doing 6 types including brain and lung, China is doing gastric, India is doing oral and Australia is doing pancreatic. The process would involve sequencing tens of thousands of genes from tumors. The goal is to find out all the mutations that separate cancerous cells from normal ones.

Yet this goal immediately runs into David Hume's well-known problem of induction. If a mutation in a gene is observed in, say 5% of tumors, would it be observed in all of them? More importantly, would it be significant as a causative agent? Consider the IDH1 gene which encodes isocitrate dehydrogenase, a key enzyme in the all-important Krebs cycle. As the article notes, IDH1 was not regarded as significant when it initially showed up in a very small subset of certain kinds of tumors. But then it consistently showed up in 12% of brain tumors of a certain kind and 8% of tumors from leukemic patients. Thus, rather than being a chance occurrence, IDH1 now seems like a significant correlative factor for cancer. It is now hot cancer genomic property.

But this is just the beginning, the very beginning. Evolutionary biologists are very well familiar with the problem of determining which mutations- called "drivers"- are causative for a given phenotype, and which ones are just "passengers". One of the biggest mistakes that "adaptionists" can make is to assume that every genotypical change somehow provides a selective advantage, when the change could just be riding on the back of another significant one. Identifying and cataloging thousands of mutant genes says nothing about which of those are truly responsible for the cancer and which ones have just come along for the ride. As a researcher quoted in the article says, "it's going to take good old-fashioned biology to really determine what these mutations are doing".

And I think we can all agree that's its going to take a lot of good old-fashioned biology to accomplish this. Firstly, one has to determine the function of the mutated gene by doing knockout and other experiments, and endeavor fraught with complications. Maybe it codes for a protein, maybe it does not. Even if it does, one then has to identify the function of that protein by finding a suitable system in which it can be expressed and purified. Structure determination may be another hard obstacle on the path to success. Finally, if any kind of therapeutic intervention is going to be attempted, one would have go first find out whether the target is "druggable". And then of course, the long and wildly uncertain road towards finding a small molecule inhibitor only begins.

Even assuming that all this happens, there is no guarantee that hitting the enzyme will produce a therapeutic response. Maybe the enzyme that is mutated is part of a complex pathway of signaling, and maybe one has to really hit something upstream or downstream to actually make a difference. And of course, hitting the target may cause a difference, but it may not be therapeutically significant enough. Thus, it's pretty clear that this project is far from curing any kind of cancer at this stage. What we just described is light years ahead of what is being currently done. Plus, it's worth noting that this is data that is extremely heterogeneous, collated from a variety of populations, potentially subject to the capricious standards of individual agencies and workers. It's nothing if not a statistician's nightmare.

So is the effort worth it? Undoubtedly. Sitting among those haystacks of mutations is the valuable one that may actually be causative. We are never going to identify the culprit if we never line up the suspects. But here, much more than in a police lineup, it is easy to be seduced by statistical significance. The pursuit of the wrong gene could easily mean the loss of millions of dollars and countless hours wasted. The researchers who have descended into this quagmire need to be more careful than Ulysses on their tortuous journey toward the discovery of important cancer-causing mutations. It is all too easy to slip on a stone and chase the wrong rabbit into the wrong rabbit hole. And there are countless number of these at every step.

As Yeats might have rephrased a line from one of his enduring poems, "tread softly, because you tread on my genes".

Ledford, H. (2010). Big science: The cancer genome challenge Nature, 464 (7291), 972-974 DOI: 10.1038/464972a

You have a Ph.D.?? Who doesn't!

I just finished reading Peter Feibelman's fantastic book, "A Ph.D. is not enough", about career advice for fresh (and also slightly staler) Ph.D.s. I very highly recommend this 100 page slim little volume. There's a blurb from the great Carl "Papa" Djerassi on the cover saying that you will get from this book in one hour what it took Djerassi 40 years to learn. Djerassi may be exaggerating, but he is close. Feibelman himself was a physics professor at SUNY Buffalo and was then a member of the technical staff at Sandia National Laboratories, so he has seen the world and tasted its ugly side.

The book was written in 1993 but its contents are as relevant as ever, and probably even more relevant in this age of tight funding and layoffs. It's got the whole works; from applying for postdoc positions (be realistic, pick a project which you think you can actually finish) to picking a postdoc advisor (don't pick a flashy young professor who would be loathe to share credit), job interviews (don't be a dilettante), writing grants (be modest in your goals even as you emphasize the big picture), giving a talk (OMG he talks about slides and projectors!), to choosing between academia and industry.

The last part is particularly intriguing and Feibelman has some novel advice for wannabe professors. Unless you are hell-bent on an academic position and wouldn't want to even think of anything else, Feibelman quite emphatically discourages plunging into academia as an assistant professor. The pay is low, respect is lacking, it's one hell of a rope trick to secure funding without any significant past background, there's basically no vacation, tenure is always uncertain, and you keep wondering when you will get publishable results even as you spend most of your time explaining to pre-meds why they deserved a D on their last exam. In short, you have no life and there's lots of necessary conditions that you have to satisfy to stay afloat, none of which is sufficient.

Better than this, says Feibelman, is to start working in a goverment or industrial lab where you (hopefully) have plenty of time for research, establish a solid reputation with financial security, and then apply to a university at the tenured professor level. Of course this is easier said than done. These days it's hard to do basic research in industry and you are afraid of losing your job every day. But I think Feibelman's point is well taken; unless you have absolutely no interest in anything other than an academic position, it's definitely worthwhile considering a more indirect path to academia where you actually have a life. My old PhD advisor actually did that and it worked out well for him.

In any case, read this little book if you are a fresh, red faced, scared little new PhD. Which we all are.

The beta-amyloid hyp(e)othesis; the saga continues

As Churchill would have said, beta-amyloid is a riddle, wrapped in a mystery, inside an enigma. For years now, the "amyloid hypothesis" has been widely-accepted as somehow being importantly responsible for Alzheimer's disease. A few years back researchers were cheerfully confident that amyloid was to AD what artherosclerotic plagues were to heart disease. Rudolph Tanzi, a Harvard neurologist who is a leading authority on the disease and identified the first AD gene, had this to say in an interview in 2000:
Q. How close are we to an effective treatment for Alzheimer's disease?
A. I wouldn't be surprised if five years from now we have a pretty effective drug that can slow the disease down enough so that it will be preventable in those at risk, and significantly slow down the deterioration of people who already have it.

Q. Why do you have such optimism?

A. Because, in 15 years, we've gone from knowing little about what causes this disease to having a pretty concrete idea of which biological pathways and body proteins are involved.

If you compare Alzheimer's to heart disease where cholesterol levels must be lowered, we now have our own cholesterol equivalent, which we call the beta-amyloid. The name of the game in Alzheimer's therapy is lowering the accumulation of beta-amyloid in the brain.
It's ten years later and we are no closer to finding an AD drug. Tanzi's hope was not unwarranted given what we knew about amyloid then. But as the amyloid hypothesis matured, so did our understanding of it. First we discovered that it's not the amyloid aggregates themselves but soluble oligomers that are probably responsible for neuronal toxicity. Now it has been proposed that amyloid could have a protective antimicrobial role (I myself had an evolutionary speculation on this) in which case targeting it could even be dangerous. The fact remains that there is no proof that amyloid causes AD. It certainly seems to be related in an important way and many revealing details about it have been uncovered in the last decade, but the proof of principle has been on an increasingly slippery slope and if anything the picture gets murkier and more fascinating.

An article in the latest issue of C & EN basically says that we are targeting beta-amyloid because at least for now we cannot think of anything better to do. It's true that currently, our best bet at treating AD lies in interfering with amyloid formation. But since amyloid formation has never been shown to be causative for AD, treatments targeted at it are always going to be something of a shot in the dark. The advantage of targeting amyloid formation though, as the article says, is that there are lots of points in the mechanism where one can potentially interfere. Two key enzymes responsible for formation of amyloid are beta and gamma-secretase, which clip the amyloid precursor peptide into apparently toxic fragments. Scores of articles are published every year about new chemical agents targeting these two enzymes, and yet the jungle is thicker than we think.

Gamma-secretase is actually a multiprotein complex whose structure is not known, so finding molecules that inhibit it is like finding a black cat in a dark room. More importantly, it's also involved in a second pathway called the Notch pathway which is critical in cell-signaling. Thus blocking it may lead to one of the classic problems in drug discovery whereby eliminating a harmful function also eliminates a useful one, often fatally. Beta-secretase is much more well-studied and its crystal structure has been solved, but it poses a classic structure-based design conundrum; the enzyme's active pocket is flexible and expansive and can bind many ligands in different subpockets. Thus, developing drugs that block this moving target is admittedly challenging. Throw in the requirements for safety and an ability to cross the blood-brain barrier (BBB), and we have a pickle on our hands that's almost as dense as amyloid plaques.

But the much more serious issue is whether any of these strategies will work at all. If amyloid formation turns out to be a side-show in AD progression, then all these strategies might ultimately come to naught. Unfortunately the data so far is not promising. The last few years have seen a disappointing string of late-stage failures of amyloid blocking molecules and antibodies in clinical trials. In some cases the agents have failed to clear the plaques, but tellingly in others, clearing the plagues did not put the disease in remission. Thus the scores of pharmaceutical companies that have several pipeline dwellers focused toward amyloid may be chasing an imaginary rabbit. There are serious concerns that scientists may have to go back to the drawing board and start all over again. This would be a huge setback.

However, hope need not be completely lost. One of the most reasonable explanations for the failure of these agents is simply that they arrived too late on the scene, when the disease had progressed too far to be defeated. Perhaps these drugs would have helped had they been administered earlier. Even cancers that can be treated if detected early fail to be cured in late stages, and AD should be no different. One of the big problems in the field is that detecting AD early is still a challenge and is being addressed by many promising neural imaging initiatives. Perhaps early and focused administration of these drugs could be successful.

Yet it all hinges on putting all your eggs in the amyloid basket. Another protein implicated in AD is tau, which forms tangles in the brain. But as the article says, targeting tau may be even harder than targeting amyloid since it is ubiquitous. Other processes hypothesized to be important for AD include oxidation and other neurotoxic processes of which amyloid may simply be a side-product. And as I was thinking this morning, perhaps no drug will ever be as effective in treating AD as a balanced lifestyle that includes preventive measures; and this may especially be true if amyloid is a natural part of our body's physiology. But as of now we have to keep on trying, and the amyloid hypotheis, shaky as it is, seems to be our best bet of making a dent into this devastating disease. At the very least it will lead to novel basic insights. Perhaps it's an indication of how primitive our understanding of the disease is that we continue to cling to amyloid. Under the present circumstances it seems to be the best we can do, but as another Churchillian admonition indicated, "it is not enough that we do our best; we must do what's necessary".

New hominid fossil has 9-year-old discoverer

A boy in South Africa has stumbled upon Australopithecus sediba, a possible Homo erectus ancestor who demonstrated both upright walking traits as well as the ability to swing with apelike arms among trees. In the hunt for transitional forms in human evolution, this is clearly an important touchstone. The report will be published in Science this Friday. The report is already causing a controversy because of its suggestion that A. sediba might be a direct Homo ancestor. Such kind of classification controversy about hominid fossils has long animated anthropology and human evolution.

The 9-year-old boy named Matthew Berger was chasing his dog and accidentally came across the remarkably well preserved fossil, situated close to the site where his father Lee Berger has been excavating for years. The site also contains fossils of carnivores and is at the bottom of a cliff, indicating that both humans and animals might have lost their footing and fallen to their deaths, either when the humans were chasing the animals or vice versa.

Just one question. Since the boy discovered the fossil, shouldn't the paper bear his name as co-author? Or are 9-year-olds safely ignored for authorship on Science papers? Although he is mentioned in the paper as the discoverer, it seems a little unfair, especially for a field like paleontology where the initial discovery is the most important event.

Will virtual screening ever work?

ResearchBlogging.org
Virtual screening (VS), wherein a large number of compounds are screened, either by docking against a protein target of interest or by similarity searching against a known active, is one of the most popular computational techniques in drug discovery. The goal of VS is to complement high-throughput screening (HTS) and the ideal goal is to at least partly substitute HTS in finding new hits.

But this goal is still far from being achieved. VS still has to make a significant contribution in the discovery of a major drug and typical hit rates range from a few tenths of a percent to perhaps a percent or two. VS has been intensively investigated for more than a decade. What do we know about its limitations, and where do we go from here?

Gisbert Schneider of ETH Zurich has some thoughts on VS in a recent review. Success in VS ultimately boils down to understanding the detailed structure and dynamics of protein-ligand complexes, a goal that we are still miles away from. We still struggle to realistically include entropy in any calculation, and we are still not completely clear about the role that buried water molecules play in dictating ligand binding. Plus we cannot yet take allosteric binding properly into account, let alone more complex interactions like protein-protein interactions. Thus, maybe, as pointed out in a past post and article, the correct question to ask would be the "anti-question", namely, why does VS work at all in spite of this supposedly woeful lack of understanding?

First of all it is important to know what VS can do well and what it can't. As the article notes, VS is still best for negative selection, that is for weeding out inactive molecules which are bad binders. One of the goals of VS is also to duplicate the correct protein-bound x-ray conformation of the ligand, and in this endeavor (termed pose prediction) VS seems to be succeeding much better than in the ultimate goal which is to rank ligand binding to a protein in order of free energy of binding. As the article notes, the true binding interaction energy landscape for a protein might be more of a plateau; thus there may be a variety of protein-ligand contacts corresponding to a 'good' solution, rather than a global optimum. Plus, one may end up modeling details that are not very relevant to the gist of the ligand binding event; in such a case productive contacts can be preserved with no great sacrifice of qualitative prediction.

Nonetheless, tiny details can sometimes radically shift the balance. No wonder that VS has been heavily dependent on the target rather than on the computational algorithm. Nature continues to throw up surprises as protein entropy, hydrophobic interactions and subtle behavior of water molecules continue to be uncovered as powerful forces operating for a particular protein-ligand complex.

In the end, modeling the dynamic behavior of macromolecules is an absolute must for lending general utility to VS campaigns. In the absence of adequate modeling of entropy, it may be wise from a practical viewpoint to aim for ligand chemotypes whose binding is dominated more by enthalpic effects. It's interesting to note a past set of studies which I had highlighted which suggested that it's really the enthalpy rather than entropy which is rendered favorable in a drug discovery project as one proceeds from hit to lead.

Finally, the author makes an appeal to fields spread far and wide to come up with ideas that could be applied in VS and related approaches. It is likely that while incremental improvements will continue to be made in the field through better understanding of protein-ligand interactions, only a novel idea would revolutionize the field. Thus insights could possibly come from unlikely quarters, including complexity theory, non linear dynamics, other aspects of physics and even engineering and architecture.

How this might happen is not at all clear, but it definitely calls for more multidisciplinary work and for more scientists from diverse fields to become interested in the problem. After all VS is fundamentally an optimization problem, one of locating the optimal ligand energetic minimum in a multidimensional landscape of protein, ligand, ions and solvent. I can't see why any mathematician, physicist or engineer worth his or her salt won't find it exciting.

Schneider, G. (2010). Virtual screening: an endless staircase? Nature Reviews Drug Discovery, 9 (4), 273-276 DOI: 10.1038/nrd3139

Distinguishing statistical significance from clinical significance

In the past post I was talking about the difference between statistical and clinical significance and how many reported studies have apparently mixed up the two. Now here's a nice case where people seem to be aware of the difference. The article is also interesting in its own right. It deals with AstraZeneca's cholesterol lowering statin drug Crestor being approved by the FDA as a preventive measure for heart attacks and stroke. If this works out Crestor could be a real cash cow for the company since its patent does not expire till 2016 (unlike Lipitor which is going to hit Pfizer hard next year).

The problem seems to be that prescription of the drug would be based on high levels not of cholesterol but of C-Reactive Protein which is an inflammatory marker of high cholesterol. The CRP-inflammation-cholesterol connection is widely believed to hold but there is no consensus in the medical community about the exact causative link (many factors can lead to high CRP levels).

The more important recent issue seems to be a study published in The Lancet which indicates a 9% increased risk of Type 2 diabetes associated with Crestor. As usual the question is whether these risks outweigh the benefits. The Crestor trial was typical of heart disease trials and involved a large population of 18,000 subjects. As the article notes, statistical significance in the reduction of heart attacks in this population does not necessarily translate to clinical significance:
Critics said the claim of cutting heart disease risk in half — repeated in news reports nationwide — may have misled some doctors and consumers because the patients were so healthy that they had little risk to begin with.

The rate of heart attacks, for example, was 0.37 percent, or 68 patients out of 8,901 who took a sugar pill. Among the Crestor patients it was 0.17 percent, or 31 patients. That 55 percent relative difference between the two groups translates to only 0.2 percentage points in absolute terms — or 2 people out of 1,000.

Stated another way, 500 people would need to be treated with Crestor for a year to avoid one usually survivable heart attack. Stroke numbers were similar.

“That’s statistically significant but not clinically significant,” said Dr. Steven W. Seiden, a cardiologist in Rockville Centre, N.Y., who is one of many practicing cardiologists closely following the issue. At $3.50 a pill, the cost of prescribing Crestor to 500 people for a year would be $638,000 to prevent one heart attack.

Is it worth it? AstraZeneca and the F.D.A. have concluded it is.

Others disagree.

“The benefit is vanishingly small,” Dr. Seiden said. “It just turns a lot of healthy people into patients and commits them to a lifetime of medication.”
To some this may seem indeed like a drug of the affluent. Only time will tell.

How useful is cheminformatics in drug discovery?

ResearchBlogging.org
Just like bioinformatics, cheminformatics has come into its own an independent framework and tool for drug design. As a measure of the field's independence and importance, consider that at least five journals primarily dedicated to it have emerged in the last couple of years, and 15000 articles on it have been published since 2003.

But how useful is it in drug discovery? The answer, just like for other approaches and technologies, is that it depends. For calculating and analyzing some properties it is more useful than for others. A group at Abbott summarizes the current knowledge of cheminformatics approaches as applied to various parameters.

To do this the group utilizes a useful classification scheme made (in)famous by ex-Secretary of Defense Donald Rumsfeld (Rumsfeld et al., J. Improb. Res. 2002). This is the classification of knowledge and facts into 'known knowns', 'unknown knowns', 'known unknowns' and 'unknown unknowns'.

The 'known known' category of properties calculated by cheminformatics consists of those that we think we have a very accurate handle on and that are easy to calculate. These include molecular weight, substructure (SS) searches, and ligand efficiency. Molecular weight can be easily calculated, and a variety of studies have indicated that high MW generally impacts drug discovery negatively. Thus the general thrust is on keeping your compounds small. Ligand efficiency is calculated as the free energy of binding per heavy atom. Medicinal chemists are more familiar with IC50 than free energy. Since IC50s are usually easy to measure and the number of heavy atoms are of course known, ligand efficiency would be a 'known known'. However the caveat is that IC50 is not the same as in vivo activity, so biochemical potency in terms of ligand efficiency might be a very different kettle of fish. Lastly, substructure searches can be carried out easily by many computer programs. Such searches are usually used when a compound having a substructure similar to one that is known is sought. The problem with SS searches arises when the decision on which SS to search becomes subjective. SS is also valuably used when certain SSs are to be avoided. In general SS searching belongs with the 'known known' category because of its ease, but subjective interpretations can render this more fuzzy.

In the 'unknown knowns' category lie an interesting set of properties; those which we think we know how to calculate but which sometimes look deceptively simple and are often subject to overconfidence in calculation. The first property in this category is one of the most important ones in drug design; logP, which is regarded to be a measure of the lipophilicity of the compound, a key parameter dictating absorption, bioavailability, and partitioning of drugs between membranes and body fluids. Several programs can calculate logP. However, as the article notes, a recent study of no less than 30 such programs located a mean error in logP calculation of 1 log unit, which means that some programs would do much worse. Thus calculation of logP, just like other computational techniques, crucially depends on the method used to calculate it. The caveat is that absolute cutoffs for logP values in library design of compounds searches might mislead, but many programs seem to do a fairly good job in producing trends. Solubility is another parameter that is notoriously difficult to calculate, especially since its calculation hinges on calculation of pKa values. pKa values in turn again are very method-dependent, and trouble arises especially for charged compounds, which include most drugs. Medicinal chemists are understandably suspicious of theoretical solubility prediction, but as for logP, such calculations may at least be used for quick estimation of qualitative trends. Another parameter in this category is plasma protein binding. Although we know a fair amount about the effect of this parameter on drug ADME, good luck trying to calculate this. Lastly, in vivo ADME is a minefield of complications. I personally would not have placed this in the 'unknown knowns' category, but at least some of the properties in this category can be calculated on the basis of specialized fragment-based models.

What about 'known unknowns'? This includes polar surface area (PSA)'. PSA is a known unknown because we know that is not exactly a real, measurable quantity. However it is still a useful parameter since PSA has been shown to relate to membrane permeability and hence is especially useful for guessing blood-brain barrier penetration for CNS drugs. A set of rules similar to the Lipinski rules says that compounds with a PSA of less than 120 A^2 are more likely to penetrate membranes. As for other properties though, calculation of PSA depends on method and usually involves calculating the 3D structure of a molecule and then calculating the PSA by assigning some kind of a 'surface area' associated with polar atoms. As the article notes, one can get wildly different results depending on which atoms are assigned as 'polar'. Plus compounds containing sulfur or phosphorus can result in big discrepancies. Clearly PSA is a useful parameter, but we have to go some way in calculating it reliably. Finally, similarity searching is all the rage these days and promises a windfall of potential discoveries. In its simplest incarnation, similarity searching aims to discover compounds similar in some structural metric to a given compound, with the assumption that similarity in structure would correspond to similarity in function. But similarity, just like beauty, is notoriously in the eye of the beholder. As the authors quip from a quaint piece of literature-related controversy:
Deciding whether two molecules are similar is much like trying to decide whether something is beautiful. There are no concrete definitions, and most chemists take an “I know it when I see it” attitude (attributed to United States Justice Potter Stewart, concurring opinion in Jacobellis v. Ohio 378 U.S. 184 (1964), regarding possible obscenity in The Lovers)."
As again illustrated, different similarity searching methods give very different hits. One of the most successful metrics for measuring similarities has turned out to be the 'Tanimoto coefficient', but other metrics also abound. Thus it is quite remarkable that similarity is already being used widely in every endeavor from virtual screening to finding new protein targets for old drugs based on drug similarity. One of the most valuable applications of similarity is in finding bioisosteres (chemically similar fragments with improved properties), something which medicinal chemists try to do all the time. A recent review of similarity-based methods in JMC summarizes the utility of such techniques. Nonetheless, similarity based methods are still among the 'known unknowns' because we don't have an objective handle on what constitutes similarity, and we may never have such a handle even if such methods are widely and successfully adopted.

Finally we come to the dreaded 'unknown unknowns'. As the authors ask, can we even list properties which we have no idea about? We can take a shot. This category includes flight of fancy which may never be achieved but which are worth striving for. One such holy grail is the large-scale, high-throughput computation of binding free energies. The binding free energy for a protein ligand complex includes contributions from an enormous number of complicated factors, but the mere attempt to calculate all these factors has valuably increased our understanding of biological systems. Thus we should continue in this endeavor. Another fancy endeavor which is the talk of the town these days is systems biology, where construction of biological networks and applications of graph theory are supposed to shed valuable light on molecular interactions. Such approaches may well be successful, but they always run the risk of becoming too abstract and divorced from reality to be truly understandable. QSAR models can suffer from the same shortcomings.

In the end, only a robust collaboration between informaticians, computer scientists, medicinal chemists and biologists can make sense of the jungle of data uncovered by cheminformatics approaches. It is key for one group of scientists to keep reality checks on others. In the end it's all about reality checks. We all know what happened when these were not applied to the original classification scheme.

Muchmore, S., Edmunds, J., Stewart, K., & Hajduk, P. (2010). Cheminformatic Tools for Medicinal Chemists Journal of Medicinal Chemistry DOI: 10.1021/jm100164z

How pretty is your (im)perfect p?

In my first year of grad school we were once sitting in a seminar in which someone had put up a graph with a p-value in it. My advisor asked all the students what a p-value was. When no one answered, he severely admonished us and said that any kind of scientist engaged in any kind of endeavor should know what a p-value is. It is after all paramount in most statistical studies, and especially in judging the consequences of clinical trials.

So do we really know what it is? Apparently not, not even scientists who routinely use statistics in their work. In spite of this, the use and overuse of the p-value and the constant invoking of "statistical significance" are universal phenomena. When a very low p-value is cited for a study, it is taken for granted that the results of the study must be true. My dad, who is a statistician turned economist, often says that one of the big laments about modern science is that too many scientists are never formally trained in statistics. This may be true, especially with regard to the p-value. An article in ScienceNews drives home this point and I think should be read by every scientist or engineer who remotely uses statistics in his or her work. The article builds on concerns raised recently by several statisticians about the misuse and over-interpretation of these concepts in clinical trials. In fact there seems to be an entire book which is devoted to criticizing the use of statistics; "The Cult of Statistical Significance".

So what's the p-value? A reasonable accurate definition is that it is the probability of getting a result at least extreme as the one you get under the assumption that the null hypothesis is true". The null hypothesis is the hypothesis that your sample is a completely unbiased sample. The p-value is basically related to the probability that you will get a result by chance alone in a completely unbiased sample; thus, the lower it is, the more "confident" you can be that chance did not produce the result. It is usually calculated as a cutoff value. Ronald Fisher, the great statistician who pioneered its use, rather arbitrarily decided this value to be 0.05. Thus generally if one gets a p-value of 0.05, he or she is "confident" that the result which is being observed is one that is "statistically significant".

But this is where misuses and misunderstandings about the p-value just begin to manifest themselves. Firstly, a p-value of 0.05 does not immediately point to a definitive conclusion:
But in fact, there’s no logical basis for using a P value from a single study to draw any conclusion. If the chance of a fluke is less than 5 percent, two possible conclusions remain: There is a real effect, or the result is an improbable fluke. Fisher’s method offers no way to know which is which. On the other hand, if a study finds no statistically significant effect, that doesn’t prove anything, either. Perhaps the effect doesn’t exist, or maybe the statistical test wasn’t powerful enough to detect a small but real effect.
Secondly, as the article says, the most common and erroneous conclusion drawn from a p-value of 0.05 is that there is 95% confidence about the null hypotheses being false, that is, there is 95% confidence that the result being observed is statistically significant and not by chance alone. But as the article notes, this is a rather blasphemous logical fallacy:
Correctly phrased, experimental data yielding a P value of 0.05 means that there is only a 5 percent chance of obtaining the observed (or more extreme) result if no real effect exists (that is, if the no-difference hypothesis is correct). But many explanations mangle the subtleties in that definition. A recent popular book on issues involving science, for example, states a commonly held misperception about the meaning of statistical significance at the 0.05 level: “This means that it is 95 percent certain that the observed difference between groups, or sets of samples, is real and could not have arisen by chance.”
That interpretation commits an egregious logical error (technical term: “transposed conditional”): confusing the odds of getting a result (if a hypothesis is true) with the odds favoring the hypothesis if you observe that result. A well-fed dog may seldom bark, but observing the rare bark does not imply that the dog is hungry. A dog may bark 5 percent of the time even if it is well-fed all of the time
This underscores the above point, that a low p-value may also mean that the result is indeed a fluke result. An especially striking problem with p-values emerged in studies of the controversial link between antidepressants and suicide. The placebo is the golden standard in clinical trials. But it turns out that not enough attention is always paid to the fact that the frequency of incidents even in two different placebo groups might be different. Thus, when calculating p-values, one must consider not their absolute values but the difference in values with reference to the placebo. Thus, a study investigating possible links between Prozac and Paxil reached a conclusion that was opposite to the real one. Consider this:
“Comparisons of the sort, ‘X is statistically significant but Y is not,’ can be misleading,” statisticians Andrew Gelman of Columbia University and Hal Stern of the University of California, Irvine, noted in an article discussing this issue in 2006 in the American Statistician. “Students and practitioners [should] be made more aware that the difference between ‘significant’ and ‘not significant’ is not itself statistically significant.”

A similar real-life example arises in studies suggesting that children and adolescents taking antidepressants face an increased risk of suicidal thoughts or behavior. Most such studies show no statistically significant increase in such risk, but some show a small (possibly due to chance) excess of suicidal behavior in groups receiving the drug rather than a placebo. One set of such studies, for instance, found that with the antidepressant Paxil, trials recorded more than twice the rate of suicidal incidents for participants given the drug compared with those given the placebo. For another antidepressant, Prozac, trials found fewer suicidal incidents with the drug than with the placebo. So it appeared that Paxil might be more dangerous than Prozac.

But actually, the rate of suicidal incidents was higher with Prozac than with Paxil. The apparent safety advantage of Prozac was due not to the behavior of kids on the drug, but to kids on placebo — in the Paxil trials, fewer kids on placebo reported incidents than those on placebo in the Prozac trials. So the original evidence for showing a possible danger signal from Paxil but not from Prozac was based on data from people in two placebo groups, none of whom received either drug. Consequently it can be misleading to use statistical significance results alone when comparing the benefits (or dangers) of two drugs.
As a general and simple example of how p-values can mislead, consider two drugs X and Y, one of whose effects seem more significant than the other when compared to placebo. In this case Drug X has a p-value of 0.04 while the second one Y has a p-value of 0.06. Thus Drug X has more "significant" effects than Drug Y, right? Well, but what happens when you just compare the two drugs to each other rather than to placebo? As the article says:
If both drugs were tested on the same disease, a conundrum arises. For even though Drug X appeared to work at a statistically significant level and Drug Y did not, the difference between the performance of Drug X and Drug Y might very well NOT be statistically significant. Had they been tested against each other, rather than separately against placebos, there may have been no statistical evidence to suggest that one was better than the other (even if their cure rates had been precisely the same as in the separate tests)
Nor does, as some would think, the hallowed "statistical significance" equate to "importance". Whether the result is really important or not depends on the exact field and study under consideration. Thus, a new drug that is statistically significant may only lead to one or two extra benefits per thousand people, which may not be clinically signicant. It is flippant to publish a statistically significant result as a generally significant result for the field.

The article has valuable criticisms of many other common practices, including the popular practice of meta-analyses (combining and analyzing different analyses), where differences between various trials may be obscured. These issues also point to a perpetual problem in statistics which is frequently ignored; that of ensuring that variations in your sample are uniformly distributed. Usually this is an assumption, but it's an assumption that may be flawed, especially in genetic studies of disease inheritance. This is part of an even larger and most general situation in any kind of statistical modeling; that of determining the distribution of your data. In fact a family friend who is an academic statistician once told me that about 90% of his graduate students' time is spent in determining the form of the data distribution. Computer programs have made the job easier, but the problem is not going to go away. The normal distribution may be one of the most remarkable aspects of her identity that nature reveals to us, but assuming such a distribution is also a slap in our face since it indicates that we don't really know the shape of the distribution.

A meta-analyses study with the best-selling anti-diabetic drug Avandia locks in on the problems:
Meta-analyses have produced many controversial conclusions. Common claims that antidepressants work no better than placebos, for example, are based on meta-analyses that do not conform to the criteria that would confer validity. Similar problems afflicted a 2007 meta-analysis, published in the New England Journal of Medicine, that attributed increased heart attack risk to the diabetes drug Avandia. Raw data from the combined trials showed that only 55 people in 10,000 had heart attacks when using Avandia, compared with 59 people per 10,000 in comparison groups. But after a series of statistical manipulations, Avandia appeared to confer an increased risk.

In principle, a proper statistical analysis can suggest an actual risk even though the raw numbers show a benefit. But in this case the criteria justifying such statistical manipulations were not met. In some of the trials, Avandia was given along with other drugs. Sometimes the non-Avandia group got placebo pills, while in other trials that group received another drug. And there were no common definitions.
Clearly it is extremely messy to compare the effects of a drug across different trials with different populations and parameters. So what is the way out of this tortuous statistical jungle? One simple remedy would be to simply insist on a lower p-value for significance, typically 0.0001. But there's another way. The article says that Bayesian methods which were neglected for a while in such studies are now being preferred to address the shortcomings of such interpretations since they include the need to acquire previous knowledge in drawing new conclusions. Bayesian or "conditional probability" has a long history and it can very valuably provide counter-intuitive but correct answers. The problem with prior information is that it itself might not be available and may need to be guessed, but an educated guess about it is better than not using it at all. The article ends with a simple example that demonstrates this:
For a simplified example, consider the use of drug tests to detect cheaters in sports. Suppose the test for steroid use among baseball players is 95 percent accurate — that is, it correctly identifies actual steroid users 95 percent of the time, and misidentifies non-users as users 5 percent of the time.

Suppose an anonymous player tests positive. What is the probability that he really is using steroids? Since the test really is accurate 95 percent of the time, the naïve answer would be that probability of guilt is 95 percent. But a Bayesian knows that such a conclusion cannot be drawn from the test alone. You would need to know some additional facts not included in this evidence. In this case, you need to know how many baseball players use steroids to begin with — that would be what a Bayesian would call the prior probability.

Now suppose, based on previous testing, that experts have established that about 5 percent of professional baseball players use steroids. Now suppose you test 400 players. How many would test positive?

• Out of the 400 players, 20 are users (5 percent) and 380 are not users.

• Of the 20 users, 19 (95 percent) would be identified correctly as users.

• Of the 380 nonusers, 19 (5 percent) would incorrectly be indicated as users.


So if you tested 400 players, 38 would test positive. Of those, 19 would be guilty users and 19 would be innocent nonusers. So if any single player’s test is positive, the chances that he really is a user are 50 percent, since an equal number of users and nonusers test positive.
In the end though, the problem is certainly related to what my dad was talking about. I know almost no scientists around me, whether chemists, biologists or computer scientists (these mainly being the ones I deal with) who were formally trained in statistics. A lot of statistical concepts that I know were hammered into my brain by my dad who knew better. Sure, most scientists including myself were taught how to derive means and medians by plugging numbers into formulas, but very few were exposed to the conceptual landscape, very few were taught how tricky the concept of statistical significance can be, how easily p-values can mislead. The concepts have been widely used and abused even in premier scientific journals and people who have pressed home this point have done all of us a valuable service. In other cases it might simply be a nuisance, but when it comes to clinical trials, simple misinterpretations of p-values and significance could translate to differences between life and death.

Statistics is not just another way of obtaining useful numbers; like mathematics it's a way of looking at the world and understanding its hidden, counterintuitive aspects that our senses don't reveal. In fact isn't that what science is precisely supposed to be about?

Some books on statistics I have found illuminating:

1. The Cartoon Guide to Statistics- Larry Gonick and Woollcott Smith. This actually does a pretty good job of explaining serious statistical concepts.
2. The Lady Tasting Tea- David Salzburg. A wonderful account of the history and significance of the science.
3. How to Lie with Statistics- Darrell Huff. How statistics can disguise facts, and how you can use this fact to your sly advantage.
4. The Cult of Statistical Significance. I haven't read it yet but it sure looks useful.
5. A Mathematician Reads The Newspaper- John Allen Paulos. A best-selling author's highly readable exposition on using math and statistics to make sense of daily news.

Chips worth their salt

This from the WSJ caught my eye today:
"PepsiCo Develops 'Designer Salt' to Chip Away at Sodium Intake"

Later this month, at a pilot manufacturing plant here, PepsiCo Inc. plans to start churning out batches of a secret new ingredient to make its Lay's potato chips healthier.

The ingredient is a new "designer salt" whose crystals are shaped and sized in a way that reduces the amount of sodium consumers ingest when they munch. PepsiCo hopes the powdery salt, which it is still studying and testing with consumers, will cut sodium levels 25% in its Lay's Classic potato chips. The new salt could help reduce sodium levels even further in seasoned Lay's chips like Sour Cream & Onion, PepsiCo said, and it could be used in other products like Cheetos and Quaker bars...working with scientists at about a dozen academic institutions and companies in Europe and the U.S., PepsiCo studied different shapes of salt crystals to try to find one that would dissolve more efficiently on the tongue. Normally, only about 20% of the salt on a chip actually dissolves on the tongue before the chip is chewed and swallowed, and the remaining 80% is swallowed without contributing to the taste, said Dr. Khan, who oversees PepsiCo's long-term research.

PepsiCo wanted a salt that would replicate the traditional "salt curve," delivering an initial spike of saltiness, then a body of flavor and lingering sensation, said Dr. Yep, who joined the company in June 2009 from Swiss flavor company Givaudan SA.

"We have to think of the whole eating experience—not just the physical product, but what's actually happening when the consumer eats the product," Dr. Yep explained.

The result was a slightly powdery ingredient that tastes like regular salt.
"Secret new ingredient"?! Perhaps not. The first thing that popped into my mind after reading this was "polymorph". People in drug discovery face the polymorph beast all the time. Polymorphs are different crystal packing arrangements of a molecule that can have dramatically different dissolution rates. Often a drug which otherwise has impeccable properties crystallizes in a form that renders it as "brick dust" which is unable to dissolve. Polymorphs also serve as a way to get around patents; you can actually patent a different polymorph of an existing drug. And of course, due to their unpredictable nature (computational prediction of crystal packing is still in a rather primitive stage) polymorphs cause extremely serious problems; the HIV protease inhibitor Ritonavir had to actually be withdrawn from the market because of the appearance of an unexpected polymorph.

To me it sounds like PepsiCo has hit on the right polymorph for common salt that has enhanced dissolution rates. I hope they have done adequate stability studies on it, because polymorphs can actually interconvert into each other based on environmental conditions. But it's a neat idea that can potentially be used for other food ingredients.