Field of Science

Showing posts with label virtual screening. Show all posts
Showing posts with label virtual screening. Show all posts

What they found in the virtual screening jungle

ResearchBlogging.org
If successful, virtual screening (VS) promises to become an efficient way to find new pharmaceutical hits, competitive with high-throughput screening (HTS). Briefly, virtual screening screens libraries of millions of compounds to find new and diverse hits, either based on similarity to a known active or by complementarity to a protein binding site. The former protocol is called ligand-based VS (LBVS) and the latter is called structure-based VS (SBVS). In a typical VS campaign, either LBVS or SBVS is used to screen compounds which are then ranked by how well they are likely to be active. The top few percent compounds are then actually tested in assays, thus validating the success or failure of the VS procedure. VS has the potential to cut down on time and expenses inherent in the HTS process.

Unfortunately the success rate of VS has been relatively poor, ranging from a few tenths of a percent to no more than a few percent. If VS is to become a standard part of drug discovery, the factors that influence its failures and successes warrant a thorough review. A recent review in JMC addresses some of these factors and raises some intriguing questions.

From a single publication with the phrase ‘virtual screening’ in 1997, there are now about a hundred such papers every year. The authors pick about 400 successful VS results from three dominant journals in the field- J. Med. Chem., Bioorg. Med. Chem. Lett. and ChemMedChem, along with some from J. Chem. Inf. Mod. They then classify these into ligand-based and structure-based techniques. As mentioned before, VS comes in two flavors. LBVS starts from a known potent compound and then looks for “similar” compounds (with dozens of ways of defining ‘similarity’), with the hope that chemical similarity will translate into biological similarity. Structure-based techniques start with a protein structure, either an x-ray crystal structure, NMR structure or a structure built from homology modeling.

The authors look for correlations between the types of methods and various parameters of success and make some interesting observations, some of which are rather counterintuitive. Here are a couple that are especially interesting.

While SBVS methods dominate, LBVS methods seem to find more potent hits, usually defined as less than 1 μM in activity. Why does this happen? Actually the authors don’t seem to dwell on this point but I have some thoughts on this. Typically when you start a ligand-based campaign, your starting point is a bonafide highly potent ligand. If you have a choice of ligands with a range of activity, you will naturally pick the most potent among them as your query. Now, if your method works, is it surprising that the hits you find based on this query will also be potent? You get what you put in.

Contrast this to a structure-based approach. You usually start with a crystal structure having a co-crystallized ligand in it. Co-crystallized ligands are usually but not always highly potent. The next step would be to use a method like docking to find hits that are complementary to your protein binding-site. But the binding site is conformationally pre-organized and optimized to bind its co-crystallized ligand. Thus, the ligands you screen will not be ranked highly by your docking protocol if they are ill-optimized for the binding site. For instance there could be significant induced fit changes during their binding. Even in the absence of explicit induced fit, fine parameters like precise hydrogen bonding geometries will greatly affect your score; after all, the protein binding site has hydrogen bonding geometries tailored for optimally binding its cognate ligand. If the hydrogen bonding geometries for your ligands are off even by a bit, the score will suffer. No wonder that the hits you find span a range of activities; you are using a binding site template that is not optimized to bind most of your ligands. The other reason which could thwart SBVS campaigns is simply that there is more work necessary in ‘preparing’ a crystal structure for docking. You have to add hydrogens to a structure, make sure all the ionization states are right and optimize the hydrogen bonding network in the protein. If any one of these steps goes wrong you will start with a fundamentally crappy protein structure for screening. Thus this protocol usually requires expert inspection, unlike LBVS where you just have to ‘prepare’ a single template ligand by making sure that the ionization state and bond orders are ok. These differences mean that your starting point for SBVS is more tortuous and much more likely to be messy than it is for LBVS. Again, you get out what you put in.

The second observation that the authors make is also interesting, and it bears on the protein preparation step we just mentioned. They find that VS campaigns where putative hits are docked into homology models seem to find more potent hits compared to those using an x-ray structure. This is surprising since x-ray structures are supposed to be the most rock-solid structures for docking. The authors speculate that this difference could be due to the fact that building good homology model requires a fair level of expertise; thus, successful VS campaigns using homology models are likely to be carried out by experts who know what they are doing, whereas x-ray structures are more likely to be used by novices who simply use the default parameters for docking.

Thirdly, the authors note an interesting correlation between the potency and frequency of hits found and the families of proteins targeted. GPCRs seem to be the most successful targeted family, followed by enzymes and then kinases. This is a pretty interesting observation and to me it points to a crucial factor which the authors don’t seem to really discuss- the nature of the libraries used for screening. These libraries are usually biased by the preferences and efforts of medicinal chemists in making certain kinds of compounds. I already blogged about a paper that looked at the surprising success of VS in finding GPCR ligands, and that paper ascribed this success to ‘library bias’ which was the fact that libraries are sometimes ‘enriched’ for GPCR-active ligands, such as aminergic compounds. Ditto for kinases; kinase inhibitor-like molecules now abound in many libraries. This is partly due to the importance of these targets and partly because of the prevalence of synthetic reactions (like cross-coupling reactions) that make it easy for medicinal chemists to synthesize such ligands and populate libraries with them. I think it would have been very interesting for the authors to analyze the nature of the screened libraries; unfortunately such information is proprietary in industrial publications. But in the absence of such data, one would have to assume that we are dealing with a fundamentally biased set of libraries, which would explain selective target enrichment.

Finally, the authors find that most successful VS efforts have come from academia, while most of the potent hits have come from industry. This seems to be consistent with the role of the former in validating methodologies and that of the former in discovering new drugs.

There are some caveats as usual. Most of the studies don’t include a detailed analysis of false positives and negatives since such analysis is time consuming. But this analysis can be extremely valuable in truly validating a method. Standards for assessing the success of VS are also not consistent and universal and these will have to be decided for true comparisons. But overall, virtual screening seems to hold promise. At the very least there are holes and gaps to fill. And researchers are always fond of these.

Ripphausen, P., Nisius, B., Peltason, L., & Bajorath, J. (2010). Quo Vadis, Virtual Screening? A Comprehensive Survey of Prospective Applications Journal of Medicinal Chemistry DOI: 10.1021/jm101020z

Steering library bias toward A2A adenosine receptor ligand discovery

ResearchBlogging.org
The A2A adenosine receptor is an important GPCR, well-known for binding caffeine. Adenosine receptors are emerging as relevant drug targets for a variety of disorders including Parkinson's disease, and there is interest in discovering new ligands that bind to them. Among adenosine receptor subtypes, the A2A receptor is one of the few GPCRs whose crystal structure is available. Thus the A2A is amenable to structure-based design efforts, and virtual screening is an especially attractive endeavor in this regard.

In the present report, a team of researchers from NIH and UCSF led by Brian Shoichet, John Irwin and Kenneth Jacobson use virtual screening to discover new ligands for the A2A. There are several points to note here. The authors use the ZINC library of drug like molecules to dock about a million and a half compounds into the binding pocket of the A2A crystal structure. They pick the best-scored 500 (0.035% of the total) ligands and investigate their fit in the binding site. Using criteria like electrostatic and VdW complementarity and novelty of chemotype, they finally select 20 of these 500 hits and test them in assays. Out of these 20, 7 inhibited binding by more than 40% at 20 μM concentration, thus constituting a hit rate of 35%. While the compounds formed the same kinds of interactions as some other A2A ligands, they were also relatively diverse in structure. The ligands were also tested in aggregation-based screens to determine that their activity was not a spurious artifact of aggregation-based inhibition.

This is a pretty good hit rate. Generally virtual screening campaigns are lucky to have a hit rate of a few percent. Curiously, the authors also found a similarly high hit rate during a past VS campaign against the well-known β2 adrenergic receptor. What could be responsible for this high hit rate against GPCRs? The reasons are interesting. One reason could be that GPCRs are very well adapted to bind small molecules in compact pockets, enclosing them and forming many kinds of productive interactions. But more intriguingly, as the authors have noted earlier, there is "biogenic bias" in favor of certain target-specific chemotypes in commercial libraries that are screened, both during VS as well as HTS. This in turn reflects the biases of medicinal chemists in picking and synthesizing certain kinds of chemotypes based on the importance of drug targets and past successes in hitting these targets. GPCRs clearly are enormously important, and GPCR-friendly ligand chemotypes thus constitute a large part of screening libraries. These chemotypes are much more prevalent than those for kinases or ion channels for instance.

This observation has both positive and negative implications. The positive implication is that one is likely to keep finding high hit rates for GPCRs using VS. However, the negative implication is that one is also going to be constrained by biogenic bias, and this might preclude finding more diverse and entirely novel subtypes. Thus, while VS campaigns for GPCRs might find a good number of hits, the novelty of these hits might not always be satisfying. One other quite intriguing point emerging in this study is that the kind of hits found (agonist, inverse agonist, antagonist etc.) reflects the ligand which the target structure used for VS is co-crystallized with. Thus the A2A houses an antagonist in the binding site, leading to a preponderance of antagonists in the top docking hits. Indeed, agonists ranked abysmally low in the list.

GPCR ligand discovery is one of the most important goals in drug discovery. This and other similar studies demonstrate that, with all its caveats, VS can be productively used to mine for new GPCR drugs.

Carlsson, J., Yoo, L., Gao, Z., Irwin, J., Shoichet, B., & Jacobson, K. (2010). Structure-Based Discovery of A2A Adenosine Receptor Ligands Journal of Medicinal Chemistry DOI: 10.1021/jm100240h

Will virtual screening ever work?

ResearchBlogging.org
Virtual screening (VS), wherein a large number of compounds are screened, either by docking against a protein target of interest or by similarity searching against a known active, is one of the most popular computational techniques in drug discovery. The goal of VS is to complement high-throughput screening (HTS) and the ideal goal is to at least partly substitute HTS in finding new hits.

But this goal is still far from being achieved. VS still has to make a significant contribution in the discovery of a major drug and typical hit rates range from a few tenths of a percent to perhaps a percent or two. VS has been intensively investigated for more than a decade. What do we know about its limitations, and where do we go from here?

Gisbert Schneider of ETH Zurich has some thoughts on VS in a recent review. Success in VS ultimately boils down to understanding the detailed structure and dynamics of protein-ligand complexes, a goal that we are still miles away from. We still struggle to realistically include entropy in any calculation, and we are still not completely clear about the role that buried water molecules play in dictating ligand binding. Plus we cannot yet take allosteric binding properly into account, let alone more complex interactions like protein-protein interactions. Thus, maybe, as pointed out in a past post and article, the correct question to ask would be the "anti-question", namely, why does VS work at all in spite of this supposedly woeful lack of understanding?

First of all it is important to know what VS can do well and what it can't. As the article notes, VS is still best for negative selection, that is for weeding out inactive molecules which are bad binders. One of the goals of VS is also to duplicate the correct protein-bound x-ray conformation of the ligand, and in this endeavor (termed pose prediction) VS seems to be succeeding much better than in the ultimate goal which is to rank ligand binding to a protein in order of free energy of binding. As the article notes, the true binding interaction energy landscape for a protein might be more of a plateau; thus there may be a variety of protein-ligand contacts corresponding to a 'good' solution, rather than a global optimum. Plus, one may end up modeling details that are not very relevant to the gist of the ligand binding event; in such a case productive contacts can be preserved with no great sacrifice of qualitative prediction.

Nonetheless, tiny details can sometimes radically shift the balance. No wonder that VS has been heavily dependent on the target rather than on the computational algorithm. Nature continues to throw up surprises as protein entropy, hydrophobic interactions and subtle behavior of water molecules continue to be uncovered as powerful forces operating for a particular protein-ligand complex.

In the end, modeling the dynamic behavior of macromolecules is an absolute must for lending general utility to VS campaigns. In the absence of adequate modeling of entropy, it may be wise from a practical viewpoint to aim for ligand chemotypes whose binding is dominated more by enthalpic effects. It's interesting to note a past set of studies which I had highlighted which suggested that it's really the enthalpy rather than entropy which is rendered favorable in a drug discovery project as one proceeds from hit to lead.

Finally, the author makes an appeal to fields spread far and wide to come up with ideas that could be applied in VS and related approaches. It is likely that while incremental improvements will continue to be made in the field through better understanding of protein-ligand interactions, only a novel idea would revolutionize the field. Thus insights could possibly come from unlikely quarters, including complexity theory, non linear dynamics, other aspects of physics and even engineering and architecture.

How this might happen is not at all clear, but it definitely calls for more multidisciplinary work and for more scientists from diverse fields to become interested in the problem. After all VS is fundamentally an optimization problem, one of locating the optimal ligand energetic minimum in a multidimensional landscape of protein, ligand, ions and solvent. I can't see why any mathematician, physicist or engineer worth his or her salt won't find it exciting.

Schneider, G. (2010). Virtual screening: an endless staircase? Nature Reviews Drug Discovery, 9 (4), 273-276 DOI: 10.1038/nrd3139

The anti-question, or when bias can be a good thing

A recent publication indicates that more bias in the form of natural product scaffolds not yet synthesized could improve hit rates in screening

ResearchBlogging.org

Most drug discovery projects are inaugurated with some kind of screening campaign where millions of molecules are screened against a biological target. Even though the hit rate from High-Throughout Screening (HTS) can be quite low, HTS still provides one of the best starting points to discover interesting new structures that display biological activity. In spite of this, there is frequent disappointment at the low rates from HTS which could be as low as 0.05%.

But instead of focusing on the low hit rate from HTS, what if we express surprise that this hit rate is actually high? This thought takes me into a slight digression. In his remarkable book The Black Swan, the author Nassim Nicholas Taleb talks about an "anti-library", the set of all books you have not read. The anti-library is in some ways more important than your library because it really tells you what you are ignorant about.

Similarly we can define an "anti-question". The anti-question is a question opposite to one which we might usually ask. So instead of asking; "Why is this drug specific for this protein?", we could ask "Why is this drug not hitting other proteins?". The value of the anti-question is that it forces us to analyze and evaluate things that we otherwise may not and enables us to think outside the box. As the wise doctor constantly exhorts detective Sponer in "I Robot" to get to the all-important right question, so it could be important to get to the right anti-question.

In the context of HTS, the anti-question actually turns out to be logical. Instead of asking, "Why is the hit rate from HTS so low"?, one should ask "Given the number of small molecules in small-molecule space (~10*60) compared to the extremely low number typically screened in HTS campaigns (10*6), why should we get any hits from HTS at all?". Even narrowing down the unimaginably large small-molecule universe to more drug-like or lead-like entities, we still run into a numbers paradox since even this number is orders of magnitude greater than what is usually screened.

In their most recent paper, Brian Shoichet and his team ask this important anti-question, and it leads them down an interesting road. Most campaigns that screen libraries focus on readily available commercial compounds and fragments that can be synthesized by organic chemists. This bias in turn reflects what has been more or less synthetically accessible through more than a hundred years of synthesis. Compared to this, the Kyoto Encyclopedia of Genes and Genomes (KEGG) contains metabolites whose structures are untainted by the minds of organic chemists. These are scaffolds among secondary metabolites and natural products that have simply been found.

There is another set of structures; the Generated Database (GDB), a theoretical set which contains all possible molecules containing less than 11 heavy atoms consisting of first-row elements (C, O, N, F). This number is not as large as may be imagined and amounts to about 26 million. In the study the authors essentially compare the set of purchasable or commercial KEBB compounds found in their own annotated library called ZINC with the GDB. They use a similarity measure called a Tanimoto coefficient derived from 2D fingerprint comparison to accomplish this. 2D fingerprints use different kinds of protocols for breaking up a molecule into bit strings and then compare bit strings by distances and atom types.

The comparison indicates something interesting; the compounds in the purchasable set are much more similar to the KEBB compounds than are the compounds from the rest of the GDB. In other words, purchasable compounds contain scaffolds that are biased towards those in the KEBB. This is a good thing, since metabolites are usually primed by nature to show at least some biological activity. Another noteworthy finding was that the bias also increased with molecular size, as compounds became more drug-like or lead-like in terms of size.

However, the more surprising and useful observation was that there are hundreds of scaffolds in the KEBB that are notpresent in the commercial library. The authors also do this comparison for other popular commercial libraries designed specifically for screening and find a similar result. The bottom line; while synthesized commercial libraries of molecules show a bias toward natural products and metabolites, there are also several natural product scaffolds that are not found in these libraries.

So what is the prescription? Introduce further bias! The compounds in the KEGG are more or less optimized for biological activity. If their scaffolds are not yet present in the commercial libraries, organic chemists should go ahead and focus on synthesizing these scaffolds and adding them to screening libraries. More such scaffolds could increase the hit rate in HTS by enriching libraries in biologically relevant scaffolds. Of course the usual caveats of false positives and promiscuous compounds should be kept in mind, and it's also not clear that proteins like kinases which are optimized to bind certain core scaffold structures would greatly benefit from these diverse scaffolds. But in terms of unmined drug space, introducing such further bias would be beneficial.

This study again goes to show the possibilities for finding new stars in the constellations and galaxies of the drug universe. Hopefully the universe will keep on expanding.

Hert, J., Irwin, J., Laggner, C., Keiser, M., & Shoichet, B. (2009). Quantifying biogenic bias in screening libraries Nature Chemical Biology DOI: 10.1038/nchembio.180

Drug Discovery, Models and Computers: A (necessarily incomplete) Personal Take

Drugs and rational drug discovery

Natural substances have been used to treat mankind’s diseases and ills since the dawn of humanity. The Middle Ages saw the use of exotic substances like sulfur and mercury to attempt to cure afflictions; most of these efforts resulted in detrimental side effects or death because of lack of knowledge of drug action. Quinine was isolated from the bark of the Cinchona tree and used for centuries to treat malaria. Salicylic acid was isolated from the Willow tree and was used for hundreds of years to treat fevers, knowledge that led to the discovery of Aspirin. The history of medicine has seen the use of substances ranging from arsenic to morphine, some of which are now known to be highly toxic or addictive.

The use of these substances reflected the state of medical knowledge of the times, when accidentally generated empirical data was the most valuable asset in the treatment of disease. Ancient physicians from Galen to Sushruta made major advances in our understanding of the human body and of medical therapies, but almost all of their knowledge was derived through patient and meticulously documented trial and error. A lack of knowledge of the scientific basis of disease meant that there were few systematic rational means of discovering new medicines, and serendipity and the traditional folk wisdom passed on through the centuries played the most important role in warding off disease.

This state of affairs continued till the 19th and 20th centuries when twin revolutions in biology and chemistry made it possible to discover drugs in a more logical manner. Organic chemistry formally began in 1848 when Friedrich Wöhler found that he could synthesize urea from simple inorganic substances like ammonium cyanate, thus dispelling the belief that organic substances could only be synthesized by living organisms (1). The further development of organic chemistry was orchestrated by the formulation of the structural theory in the late 19th century by Kekulé, Cooper, Kolbe, Perkin and others (1). This framework made it possible to start to elucidate the precise arrangement of atoms in biologically active compounds. Knowledge of this arrangement in turn led to routes for synthesis of these molecules. These investigations also provided impetus to the synthesis of non-natural molecules of practical interest, sparking off the field of synthetic organic chemistry. However, while the power of synthetic organic chemistry later provided several novel drugs, the legacy of natural products is still prominent, and about half of the drugs currently on the market are either natural products or derived from natural products (2).

Success in the application of chemistry to medicine was exemplified in the early 20th century by tentative investigations of what we currently call structure-activity relationships (SAR). Salvarsan, an arsenic compound used for treating syphilis, was perhaps the first example of a biologically active substance that had been improved by systematic investigation and modification. As the same time, chemists like Emil Fischer were instrumental in synthesizing further naturally occurring substances like carbohydrates and proteins, thus extending the scope of organic synthesis into biochemistry.

The revolution in structure determination initiated by physicists led to vastly improved synthesis and studies of bioactive substances. At this point, rational drug discovery began to take shape. Chemists working in tandem with biologists made hundreds of substances which were tested for their efficacy against various diseases. Knowledge from biological testing was in turn translated into modifications of the starting compounds. The first successful example of such rational efforts was the synthesis of sulfa drugs used to treat infections in the 1930s (3). These compounds were the first effective antibiotics and were followed by the famous discovery, but this time serendipitous, of penicillin by Alexander Fleming in 1928 (4).

Rational drug discovery received a substantial impetus because of the post-World War 2 breakthroughs of structure determination by x-ray crystallography that revealed the structures of small molecules, proteins and DNA. The discovery of the structure of DNA in 1953 by Watson and Crick heralded the advent of molecular biology (5). This landmark event led in succession to the elucidation of the genetic code and the transfer of genetic information from DNA to RNA that results in protein synthesis. The first structure determination of a protein- hemoglobin by Perutz (6)- was followed by the structure determination of several other proteins, some of which were pharmacologically important. Such advances and preceding ones by Pauling and others (7) led to the elucidation of common motifs in proteins such as alpha helices and beta sheets. The simultaneous growth of techniques in biological assaying and enzyme kinetics made it possible to monitor the binding of drugs to biomolecules. At the same time, better application of statistics and the standardization of double blind, controlled clinical trials caused a fundamental change in the testing and approval of new medicines. A particularly noteworthy example of one of the first drugs discovered through rational investigations is cimetidine (8), a drug for acid reflux that was for several years the best-selling drug in the world.

Structure-based drug design and CADD

As x-ray structures of protein-ligand complexes began to emerge in the 70s and 80s, rational drug discovery received enormous benefits. The development was also accompanied by High-Throughput Screening, an ability to screen thousands of ligands against a protein target to identify likely binders. These studies led to what today is known as “structure-based drug design” (SBDD) (9). In SBDD, the structure of a protein bound to a ligand is used as a starting point for further modification and improvement of properties of the drug. While care has to taken in order to fit the structure well to the electron density in the data (10), well-resolved data can greatly help in identifying points of contact between the drug and the protein active site as well as the presence of special chemical moieties such as metals and cofactors. Water molecules identified in the active site can play crucial roles in bridging interactions between the protein and ligand (11). Early examples of classes of drugs discovered using structure-based design include Captopril (12) (angiotensin-converting enzyme inhibitor- hypertension) and Trusopt13 (carbonic anhydrase inhibitor- glaucoma) and recent examples include Aliskiren (14) (renin inhibitor- hypertension) and HIV protease inhibitors (13).

As SBDD progressed, another approach called ligand-based design (LBD) has also recently emerged. Obtaining x-ray structures of drugs bound to proteins is still a tricky endeavor, and one is often forced to proceed on the basis of the structure of an active compound alone. Techniques developed to tackle this problem involve QSAR (Quantitative Structure-Activity Relationships) (15) and pharmacophore construction in which the features essential for a particular ligand to bind to a certain protein are conjectured from affinity data for several similar and dissimilar molecules. Molecules based on the minimal set of interacting features are then synthesized and tested. However, since molecules can frequently adopt diverse conformations when binding to a protein, care has to be exercised in developing such hypotheses. In addition, it is relatively easy to be led astray by a high correlation between affinity data in the training set. It is paramount in such cases to remember the general discrepancy between correlation and causation, and overfitting of models can lead to both spurious correlations and absence of causation (16). While LBD is more recent than SBDD, it has turned out to be valuable in certain cases. Noteworthy is a recent example where an inhibitor of NAADP was discovered by shape-based virtual screening (17) (vida infra)

As rational drug discovery progressed, software and hardware capacities of computers also grew exponentially, and CADD (Computer-Aided Drug Design) began to be increasingly applied to drug discovery. An effort was made to integrate CADD in the traditional chemistry and biology workflow and its principal development took place in the pharmaceutical industries, although academic groups were also instrumental in developing some capabilities (18). The declining costs of memory and storage, increasing processing power and facile computer graphics software put CADD within the grasp of relatively untrained computational chemists or experimental scientists. While the general verdict on the contribution of CADD to drug discovery is still forthcoming, many drugs currently on the market now include CADD as an important component of their discovery and development (19). Many calculations that once were impractical because of constraints of time and computing power can now be routinely performed, some on a common desktop. Currently the use of CADD in drug design aims to address three principal problems, all of which are valuable to drug discovery.

Virtual Screening

Virtual screening (VS) is defined by the ability to test thousands or millions of potential ligands against a protein, distinguish the actives from inactives and rank the ‘true’ binders in a certain top fraction. If validated, VS would serve as a valuable complement, if not substitute, for HTS and would save significant amounts of resources and time in HTS. Just like HTS, VS has to circumvent the problem of false positives and false negatives, the latter of which in some ways are more valuable since by definition they would not be identified. VS can be either structure-based or ligand-based. Both approaches have enjoyed partial success although recent studies have validated 3D ligand-based techniques in which ligand structures are compared to known active ligands by means of certain metrics as having a greater hit rate than structure-based techniques (20). Virtual libraries of molecules such as DUD (21) (Directory of Useful Decoys) and ZINC (22) have been built to test the performance of several VS programs and compare them with each other. These libraries typically consist of a few actives and several thousand decoys, with the goal being to rank the true actives above the true decoys using some metric.

Paramount in such retrospective assessment is an accurate method for evaluating the success and failure of these methods (23,24). Until now ‘enrichment factors’ have mostly been used for this purpose (24). The EF refers to the number of ‘true’ actives that rank in a certain top fraction (typically 1% or 10%) as a function of the screened database. However the EF suffers from certain drawbacks, such as being dependent on the number of decoys in the dataset. To circumvent this problem, recent studies have suggested the use of the ROC (Receiver Operator Characteristic) curve, a graph that plots false positives vs. true positives (24,25) (Figure 1). The curve indicates what the false positive rate is for a given true positive rate and the measured variable is the Area Under the Curve (AUC). A completely random performance gives a straight line (AUC 0.5), while better performance results in a hyperbolic curve (AUC > 0.5).

Image Hosted by ImageShack.us


Figure 1: ROC curve for three different VS scenarios. Completely random performance will give the straight white line (AUC 0.5), an ideal performance (no false positives and all true positives) will give the red line (AUC 1.0) and a good VS algorithm will produce the yellow curve (0.5 < AUC < 1.0)

Until now VS has provided limited evidence of success. Yet its capabilities are being improved and it has become a part of the computational chemist’s standard repertoire. In some cases VS can provide more hits compared to HTS (26) and in others, VS at the very least provides a method to narrow down the number of compounds actually assayed (27). As advances in general SBDD and LBD continue, the power of VS to identify true actives will undoubtedly increase.

Pose-prediction

The second goal sought by computational chemists is to predict the binding orientation of a ligand in the binding pocket of a protein, a task that falls within the domain of SBDD. This endeavor if successful will provide an enormous benefit in cases where crystal structures of protein-ligand complexes are not easily obtained. Since such cases are still very common, pose-prediction continues to be both a challenge as well as a valuable objective. There are two principal problems in pose prediction. The first one relates to the scoring of the poses obtained in order to identify the top-scoring pose as the ‘real’ pose; current docking programs are notorious for their scoring unreliability, certainly in an absolute sense and sometimes even in a relative sense. The problem of pose prediction ultimately is defined by the ability of an algorithm to find the global minimum orientation and conformation of a ligand on the potential energy surface (PES) generated by the protein active site (28). As such it is susceptible to the common inadequacies inherent in comprehensively sampling a complex PES. Frequently however, as in the case of CDK7, past empirical data including knowledge of poses of known actives (roscovitine in this case) provides confidence about the pose of the unknown ligand.

Another serious problem in pose prediction is the inability of many current algorithms to adequately sample protein motion. X-ray structures provide only a static snapshot of ligand binding that may obscure considerable conformational changes in protein motifs. Molecular dynamics simulations followed by docking (‘ensemble docking’) have remedied this limitation to some extent (29), induced-fit docking algorithms have now been included in programs such as GLIDE30, and complementary information from dynamical NMR studies may help judicious selection between several protein poses. Yet simulating large-scale protein motions are still outside the domain of most MD simulations, although significant progress has been made in recent years (31,32).

An example of how pose prediction can shed light on anomalous binding modes and possibly save the allocation of time and financial resources was experienced by the present author during his study of a paper detailing the development of inhibitors of the p38 MAP kinase (33). In one instance the authors followed the SAR data in the absence of a crystal structure and observed contradictory changes in activity influenced by structural modifications. Crystallography on the protein ligand complex finally revealed an anomalous conformation of the ligand in which the oxygen of an amide at the 2 position of a thiophene was cis to the thiophene sulfur, when chemical intuition would have expected it to be trans. The crystal structure showed that an unfavorable interaction of a negatively charged glutamate with the sulfur in the more common trans conformation forced the sulfur to adopt the slightly unfavorable cis position with respect to the amide oxygen. Surprisingly this preference was seen in all top 5 GLIDE poses of the docked compound. This example indicates that at least in some cases pose prediction could serve as a valuable timesaving complement and possible alternative to crystallography.

Binding affinity prediction

The third goal is possibly the most challenging endeavor for computational chemistry. Rank-ordering ligands in terms of their binding affinity involves accurate scoring, which as noted above is a recalcitrant problem. The problem is a fundamental one since it really involves calculating absolute free energies of protein ligand binding. The most accurate and sophisticated approaches for calculating these energies are the Free-Energy Perturbation (FEP) (34) or Thermodynamic Integration (TI) methods based on MD simulations and statistical thermodynamics. The methods involve ‘mutating’ one ligand to another in hundreds of thousands of infinitesimal steps and evaluating the binding enthalpy and entropy at every step. As of now, these techniques are some of the most computationally expensive techniques in the field. This problem typically limits their use only to evaluating free energy changes between ligand that differ little in structure. Therefore successful examples where they have found their greatest use involve cases where small substituents on aromatic rings are modified to evaluate changes in binding affinity (35). However as computing power grows, these techniques will continue to find more applications in drug discovery.

Apart from these three goals, a major goal of computational science in drug discovery is to aid the later stages of drug development when pharmacokinetics (PK) and ADMET (Absorption Distribution Metabolism Excretion Toxicity) issues are key. Optimizing the binding affinity of a particular compound to a protein only results in an efficient ligand and not necessarily an efficient drug. Computational chemistry can make valuable contributions to these later developmental stages by trying to predict the relevant properties of ligands in the early stages, thus limiting the typically high attrition of drugs in the advanced phases. While much remains to be accomplished in this context, some progress has been made (36). For example, the well-known Lipinski Rule of Five (37) provides a set of physicochemical properties necessary for drugs to have good bioavailability and computational approaches are starting to help evaluate these properties during early stages. The QikProp program developed by Jorgensen et al. calculates properties like Caco-2 cell permeability, possible metabolites, % absorption in the GI tract and logP values (38). Such programs are still largely empirical, depending on a large dataset of properties of known drugs for comparison and fitting.

Models, computers and drug discovery

In applying models to designing drugs and simulating their interactions with proteins, the most valuable lesson to remember is that these are models that are generated by computers. Models seldom mirror reality; in fact they often may succeed in spite of reality. Models are not usually designed to simulate reality but they are designed to produce results that agree with experiment. There are many approaches that produce such results. These approaches may not always encompass factors operating in real environments. In QSAR for instance, it has been shown that adding enough number of parameters to your model can lead to a good fit to the data with a high correlation coefficient. However the model may be overfitted; that is, it may seem to fit the known data very well but may fail to predict the unknown data, which is what it was designed to do (16,39). In such cases, using more advanced statistical methods and using ‘bootstrapping’ (leaving out a part of the data and looking at the resulting fit to investigate whether that part of data is predicted) can lead to improvement in results (39).

Models can also be used in spite of outliers. A high correlation coefficient of 0.85 that leads to acceptance of a model may nonetheless lead to one or two outliers. It then becomes important to be aware of the physical anomaly which the outliers represent. The reason for this is clear. If the variable producing the outlier does not constitute a part of the model building, then applying the well-trained model to a system where that particular variable suddenly becomes dominant will result in a failure of the model. Such outliers, termed ‘black swans’, can prove extremely deleterious if their value is unusually high (40). This phenomenon is known to operate in the field of financial engineering (40). In modeling for instance, if the training set for a docking model consists of largely lipophilic protein active sites, then the model may fail to deliver cogent results if applied to a set of ligands binding to a protein that has an anomalously polar or charged active site. If the value of this protein is unusually high for a particular pharmaceutical project, an inability to predict its behavior under unforeseen circumstances may lead to valuable losses. Clearly in this case the physical variable, namely the polarity of the active site, was not taken into account in spite of the fact that the model delivered a high initial correlation merely because of the addition of a large number of parameters or descriptors, none of which was related in a significant way to the polarity of the binding pocket. The difference between correlation and causation is especially relevant in this respect. This hypothetical example illustrates one of the limitations of models iterated above; that they may not bear relationship to actual physical phenomena and may yet fit the data well enough because of various reasons to elicit confidence in their predictive ability.

In summary, models of the kind that are used in computational chemistry have to be carefully evaluated, especially in the context of practical applications like drug discovery where time and financial resources are valuable. Training the model on high-quality datasets, reiterating the difference between correlation and causation and better application of statistics and bootstrapping can help to avert model failure.

In the end however, it is experiment that is of paramount importance for building the model. Inaccurate experimental data with uncertain error margins will undoubtedly hinder the success of every subsequent step in model building. To this end, generating, presenting and evaluating accurate experimental data is a responsibility that needs to be fulfilled by both computational chemists and experimentalists, and it is only a fruitful and synergistic alliance between the two groups that can help overcome the complex challenges in drug discovery.


References

(1) Berson, J. A. Chemical creativity : ideas from the work of Woodward, Hückel, Meerwein and others; 1st ed.; Wiley-VCH: Weinheim ; Chichester, 1999.
(2) Paterson, I.; Anderson, E. A. Science 2005, 310, 451-3.
(3) Hager, T. The demon under the microscope : from battlefield hospitals to Nazi labs, one doctor's heroic search for the world's first miracle drug; 1st ed.; Harmony Books: New York, 2006.
(4) Macfarlane, G. Alexander Fleming, the man and the myth; Oxford University Press: Oxford [Oxfordshire] ; New York, 1985.
(5) Judson, H. F. The eighth day of creation : makers of the revolution in biology; Expanded ed.; CSHL Press: Plainview, N.Y., 1996.
(6) Ferry, G. Max Perutz and the secret of life; Cold Spring Harbor Laboratory Press: New York, 2008.
(7) Hager, T. Linus Pauling and the chemistry of life; Oxford University Press: New York, 1998.
(8) Black, J. Annu Rev Pharmacol Toxicol 1996, 36, 1-33.
(9) Jhoti, H.; Leach, A. R. Structure-based drug discovery; Springer: Dordrecht, 2007.
(10) Davis, A. M.; Teague, S. J.; Kleywegt, G. J. Angew. Chem. Int. Ed. Engl. 2003, 42, 2718-36.
(11) Ball, P. Chem. Rev. 2008, 108, 74-108.
(12) Smith, C. G.; Vane, J. R. FASEB J. 2003, 17, 788-9.
(13) Kubinyi, H. J. Recept. Signal Transduct. Res. 1999, 19, 15-39.
(14) Wood, J. M.; Maibaum, J.; Rahuel, J.; Grutter, M. G.; Cohen, N. C.; Rasetti, V.; Ruger, H.; Goschke, R.; Stutz, S.; Fuhrer, W.; Schilling, W.; Rigollier, P.; Yamaguchi, Y.; Cumin, F.; Baum, H. P.; Schnell, C. R.; Herold, P.; Mah, R.; Jensen, C.; O'Brien, E.; Stanton, A.; Bedigian, M. P. Biochem Biophys Res Commun 2003, 308, 698-705.
(15) Hansch, C.; Leo, A.; Hoekman, D. H. Exploring QSAR; American Chemical Society: Washington, DC, 1995.
(16) Doweyko, A. M. J. Comput. Aided Mol. Des. 2008, 22, 81-9.
(17) Naylor, E.; Arredouani, A.; Vasudevan, S. R.; Lewis, A. M.; Parkesh, R.; Mizote, A.; Rosen, D.; Thomas, J. M.; Izumi, M.; Ganesan, A.; Galione, A.; Churchill, G. C. Nat. Chem. Biol. 2009, 5, 220-6.
(18) Snyder, J. P. Med. Res. Rev. 1991, 11, 641-62.
(19) Jorgensen, W. L. Science 2004, 303, 1813-8.
(20) McGaughey, G. B.; Sheridan, R. P.; Bayly, C. I.; Culberson, J. C.; Kreatsoulas, C.; Lindsley, S.; Maiorov, V.; Truchon, J. F.; Cornell, W. D. J. Chem. Inf. Model. 2007, 47, 1504-19.
(21) Huang, N.; Shoichet, B. K.; Irwin, J. J. J. Med. Chem. 2006, 49, 6789-801.
(22) Irwin, J. J.; Shoichet, B. K. J. Chem. Inf. Model. 2005, 45, 177-82.
(23) Jain, A. N.; Nicholls, A. J. Comput. Aided Mol. Des. 2008, 22, 133-9.
(24) Hawkins, P. C.; Warren, G. L.; Skillman, A. G.; Nicholls, A. J. Comput. Aided Mol. Des. 2008, 22, 179-90.
(25) Triballeau, N.; Acher, F.; Brabet, I.; Pin, J. P.; Bertrand, H. O. J. Med. Chem. 2005, 48, 2534-47.
(26) Babaoglu, K.; Simeonov, A.; Irwin, J. J.; Nelson, M. E.; Feng, B.; Thomas, C. J.; Cancian, L.; Costi, M. P.; Maltby, D. A.; Jadhav, A.; Inglese, J.; Austin, C. P.; Shoichet, B. K. J. Med. Chem. 2008, 51, 2502-11.
(27) Peach, M. L.; Tan, N.; Choyke, S. J.; Giubellino, A.; Athauda, G.; Burke, T. R.; Nicklaus, M. C.; Bottaro, D. P. J. Med. Chem. 2009.
(28) Jain, A. N. J. Comput. Aided Mol. Des. 2008, 22, 201-12.
(29) Rao, S.; Sanschagrin, P. C.; Greenwood, J. R.; Repasky, M. P.; Sherman, W.; Farid, R. J. Comput. Aided Mol. Des. 2008, 22, 621-7.
(30) Sherman, W.; Day, T.; Jacobson, M. P.; Friesner, R. A.; Farid, R. J. Med. Chem. 2006, 49, 534-53.
(31) Shan, Y.; Seeliger, M. A.; Eastwood, M. P.; Frank, F.; Xu, H.; Jensen, M. O.; Dror, R. O.; Kuriyan, J.; Shaw, D. E. PNAS 2009, 106, 139-44.
(32) Jensen, M. O.; Dror, R. O.; Xu, H.; Borhani, D. W.; Arkin, I. T.; Eastwood, M. P.; Shaw, D. E. PNAS 2008, 105, 14430-5.
(33) Goldberg, D. R.; Hao, M. H.; Qian, K. C.; Swinamer, A. D.; Gao, D. A.; Xiong, Z.; Sarko, C.; Berry, A.; Lord, J.; Magolda, R. L.; Fadra, T.; Kroe, R. R.; Kukulka, A.; Madwed, J. B.; Martin, L.; Pargellis, C.; Skow, D.; Song, J. J.; Tan, Z.; Torcellini, C. A.; Zimmitti, C. S.; Yee, N. K.; Moss, N. J. Med. Chem. 2007, 50, 4016-26.
(34) Jorgensen, W. L.; Thomas, L. L. J. Chem. Theor. Comp. 2008, 4, 869-876.
(35) Zeevaart, J. G.; Wang, L. G.; Thakur, V. V.; Leung, C. S.; Tirado-Rives, J.; Bailey, C. M.; Domaoal, R. A.; Anderson, K. S.; Jorgensen, W. L. J. Am. Chem. Soc. 2008, 130, 9492-9499.
(36) Martin, Y. C. J. Med. Chem. 2005, 48, 3164-70.
(37) Lipinski, C. A.; Lombardo, F.; Dominy, B. W.; Feeney, P. J. Adv. Drug. Del. Rev. 1997, 23, 3-25.
(38) Ioakimidis, L.; Thoukydidis, L.; Mirza, A.; Naeem, S.; Reynisson, J. Qsar & Comb. Sci. 2008, 27, 445-456.
(39) Hawkins, D. M. J. Chem. Inf. Comput. Sci. 2004, 44, 1-12.
(40) Taleb, N. The black swan : the impact of the highly improbable; 1st ed.; Random House: New York, 2007.

New ligands for everyone's favorite protein

ResearchBlogging.org

A landmark event in structural biology and pharmacology occurred in 2007 when the structure of the ß2-adrenergic receptor was solved using xray crystallography by Brian Kobilka's and Raymond Stevens's groups at Stanford and Scripps respectively. The structure was co-crystallized with the inverse agonist carazolol. Until then the only GPCR structure available was that of rhodopsin and all homology models of GPCR were based on this structure. The availability of this new high resolution structure opened new avenues for structure-based GPCR ligand discovery.

The ß2 binding pocket is especially suited for drug design since it is tight, narrow and lined with mostly hydrophobic residues with polar residues well-separated. Two crucial residues, an Asp and a Ser bind to the ubiquitous charged amino nitrogen present in most catecholamines and the aromatic section of the molecule docks deep into the hydrophobic pocket. These particular features also make computational docking more facile; a mix of polar and non-polar features with bridging waters can make docking and scoring more challenging.

Since the ß2 structure has been published, attempts are being made to use it as a template to build homology models of other GPCRs. A couple of months back I described an interesting proof-of-principle paper by Stefano Costanzi that sought to investigate how well a homology model based on the ß2 would perform. In that study carazolol itself was used as a ligand for docking into the homology model. Comparison with the original crystal structure revealed that while the ligand docked more or less satisfactorily, an important deviation in its orientation could be explained by a counterintuitive orientation of a Phe residue in the binding site. The study indicated that the devil is in the details when one is considering homology models.

However, finding ligands for the ß2 itself is also an important and interesting endeavor. Virtual screening could help in such studies. To this end Brian Shoichet, Brian Kobilka and their group have used the DOCK program to virtually screen one million lead-like ligands from their ZINC database against the ß2. Out of the 1 million ranked poses, they chose and clustered the top 500 compounds (0.05% of the database) into 25 unique chemotypes, a choice also guided by visual inspection of the protein-ligand interactions and commercial availability. They then tested these 25 compounds against the ß2 and found 6 compounds with IC50s better than 4 µM. One of these compounds with an IC50 of 9 nM is perhaps the most potent inverse agonist of the ß2 known. The binding poses revealed substantial overlap of similar functional groups with the carazolol structure. Two compounds turned out to have novel chemotypes and bore very little similarity with known ß2 ligands. A negative test was also run where a known predicted binder was chemical modified so that it would not bind.

Interestingly all the compounds found were inverse agonists. The ZINC library is somewhat biased against aminergic ligands as is most of chemical space. The catecholamine scaffold is one of the favourite scaffolds in medicinal chemistry. However, subtle difference in protein structure can sometimes turn an inverse agonist into an agonist. In this case, small changes in the orientation of the crucial Ser residue near the mouth of the binding pocket. In a past study for instance, slightly changing the rotameric features of the Ser residue thus resulting in a different orientation of the hydroxyl was sufficient to retrieve agonists.

The study thus shows the value of virtual screening in the discovery of new ß2 ligands and indicates the effect of library bias and protein structure on such ligand discovery. Many factors can contribute to the success or failure of such a search; nature is a multi-armed demon.

Reference:
Kolb, P., Rosenbaum, D., Irwin, J., Fung, J., Kobilka, B., & Shoichet, B. (2009). Structure-based discovery of ß2-adrenergic receptor ligands Proceedings of the National Academy of Sciences DOI: 10.1073/pnas.0812657106

Post-docking as a post-doc, and some fragment docking

I am now ready to post-doc. I am also now ready to post-dock, that is, engage in activities beyond docking. Sorry, I could not resist cracking that terrible joke. It's been a long journey and I have enjoyed every most moments of it. Thanks to everyone in the chemistry blogworld who regaled, informed, provoked and entertained on this blog. I am now ready to move on to the freakingly chilly Northeast. Location not disclosed for now, but maybe later.

ResearchBlogging.org

Speaking of docking, here is a nice paper from the Shoichet group in which they use fragment docking to divine hits from a large library for a beta-lactamase. Fragment docking can often be tricky compared to "normal" docking since fragments being small usually demonstrate promiscuity, low-affinity and non-selectivity in binding. Fragment docking thus is not yet a completely validated technique.

In their study, the present authors screen their ZINC library for fragments binding to the ß lactamase CTX-M by docking using the program DOCK. They also screen a lead-like library for larger molecules. The top hits from the fragment docking results were assayed and showed micromolar inhibition against the lactamase. These included tetrazole scaffolds not seen before. Importantly, five of these hits could be crystallized and the high-res crystal structures validated the docking modes.

What was interesting was that the same tetrazole scaffolds in the larger lead-like library were ranked very low (>900) and would not have ever been selected had their tetrazole fragments not showed up at the top in the fragment docking results. These compounds, when assayed showed sub-milimolar to micromolar activity against the lactamase. Thus, the protocol essentially demonstrated that fragment docking can reveal hits that can be missed by docking larger lead-like molecules. One of the reasons DOCK succeeds in this capacity is because of its use of a physics-based scoring function that has no bias against fragments. It also helps that the active site of CTM-X is relatively rigid with little protein motion.

The fragments were also assayed against another lactamase for Amp C. Usually, hits for CTM-X and Amp C are mutually exclusive. What was seen was that the higher the potency of the fragments for CTX-M, the higher the specificity for CTX-M, not surprising considering that increased potency translates to a much better complementary fit of the fragments for CTX-M.

Fragment docking can be messy since fragments can bind non-selectively and haphazardly to many different parts of many different proteins. But this study indicates that fragment docking is not an uninteresting strategy to possibly find hits from other lead-like libraries that may be otherwise concealed.

The potencies of the compounds found may look pretty weak, but because there are extremely few molecules inhibiting these medicinally important lactamases, such advances are welcome. Lactamases are of course an important target for overcoming resistance in antibiotic treatment.

Reference:
Chen, Y., & Shoichet, B. (2009). Molecular docking and ligand specificity in fragment-based inhibitor discovery Nature Chemical Biology DOI: 10.1038/nchembio.155

Met/VEGFR/FGFR inhibitors by VS

ResearchBlogging.org

This one comes from a NCI group. Starting with a database of 3.5 million compounds, 175 prospective candidates were finally selected for inhibition of Met kinase by virtual screening based on filtering by size, log P, polarity, h-bond donors and acceptors etc. The 175 compounds were docked by GOLD and 70 were selected on the basis of force-field calculated protein-ligand interaction energies (always a risky endeavor, but for a similar series, errors may cancel). For paring down the compound set, an empirical pharmacophore based on the usual kinase interactions (eg. h-bonding to hinge residues and h-bonding to a crucial Tyr to keep the kinase locked in the inactive conformation) was used.

The 70 compounds bought from vendors were assayed in cell-free Met assays as well as HGF (hepatocyte growth factor) induced Met activation. 3 compounds were chosen with modest Met inhibition IC50 values of 0.6 and 40 µM. The IC50s for the intact cells were 30, 50 and 30 µM. In spite of this less than stellar inhibition performance, the three were tested for the inhibition of other kinases and found to inhibit VEGFR and FGFR to similar extents. Since the single-kinase inhibition paradigm seems to be called into question these days, one might as well diversify. The different inhibition values were rationalized on the basis of hydrogen bonding and other interactions (or the lack thereof) by docking. Wisely, docking and docking scores were not used for judging binding affinity, just the orientations of the compounds in the active sites.

The flavone-like looks of the compounds make me quite suspicious about PK and Tox. Only time will tell. But Met is emerging as a valuable target.

Reference:
Megan L. Peach, Nelly Tan, Sarah J. Choyke, Alessio Giubellino, Gagani Athauda, Terrence R. Burke, Jr., Marc C. Nicklaus and Donald P. Bottaro (2009)
Directed Discovery of Agents Targeting the Met Tyrosine Kinase Domain by Virtual Screening
J. Med. Chem ASAP

Water-Inclusive Docking with Remarkable Approximations

ResearchBlogging.org
The role of water in mediating protein-ligand interactions has now been well-recognized by both experimentalists and modelers. However it's been relatively recently that modelers have actually started taking the unique roles that water plays into account. While the role of water in bridging ligand and protein atoms is obvious, a more subtle but crucial role of water is to fill up hydrophobic pockets in proteins. Such waters can be very unhappy in such pockets because of both unfavourable entropy (not much movement) and enthalpy (inability to form a full complement of 4 hydrogen bonds). If one can design a ligand that will displace such waters, significant gains in affinity would be obtained. One docking approach that does take such properties of waters into consideration is Schrodinger's Glide, with a recent paper attesting to the importance of such a method for Factor Xa inhibitors.

Clearly the exclusion of water molecules during docking and virtual screening (VS) will hamper enrichment factors, namely how well you can rank actives above inactives. Now a series of experiments from Brian Shoichet's group illustrates the benefits of including waters in active sites when doing virtual screening. These experiments seem to work in spite of two approximations that should have posed significant problems, but surprisingly did not.

To initiate the experiments, the authors chose a set of 24 targets and their corresponding ligands from their well-known DUD ligand set. This is a VS data set in which ligands are distinguished by topology but not by physical properties such as size and lipophilicity. This feature makes sure that ligands aren't trivially distinguished by VS methods on the basis of such properties alone. Importantly, the complexes were chosen so that the waters in them are bridging waters with at least two hydrogen bonds to the protein, and not waters which simply occupy hydrophobic pockets. Note that this would exclude a lot of important cases where affinity comes from displacement of such waters.

Now for the approximations. Firstly, the authors treated each water molecule separately in multiple configurations. They then scored the docked ligands against each such configuration as well as the rest of the protein. The waters were treated as either "on" or "off", that is, either displaced or not displaced. Whether to keep a water or not depended on whether the score improved or not when it was displaced by a ligand. The best scored ligands were then selected and figured high on the enrichment curve. This is a significant approximation because the assumption here is that every water contributes to ligand binding affinity independently of the other waters. While this would be true in certain cases, there is no reason to assume that it would generally hold.

The second approximation was even more important and startling. All the waters were regarded as energetically equivalent. From our knowledge of protein-ligand interactions, we know that the reason why evaluating waters in protein active sites is such a tricky business is precisely because each water has a different energetic profile. In fact the Factor Xa study cited above takes this profile into consideration. Without such an analysis it would be difficult to tell the medicinal chemist which part of the molecule to modify to get the best binding affinity from water displacement.

The most important benefit of this approximate approach was a linear increase in computational time instead of an exponential one. This was clearly because of the separate-water configuration approximation. The calculation of individual water free energies would also have added to this time.

In spite of these crucial approximations, the results indicate that the ability to distinguish actives from inactives was considerably improved for 12 out of 24 targets. This is not saying much, but even 50% sounds like a lot in the face of such approximations. Clearly an examination of the protein active site will also help to evaluate which cases will benefit, but it will also naturally depend on the structure of the ligand.

For now, this is an encouraging result and indicates that this approach could be implemented in virtual screening. There are probably very few cases where docking accuracy decreases when waters are included. With the sparse increases in computational time, this would be a quick and dirty but viable approach for virtual screening.

Reference:
Niu Huang, Brian K. Shoichet (2008). Exploiting Ordered Waters in Molecular Docking Journal of Medicinal Chemistry, 51 (16), 4862-4865 DOI: 10.1021/jm8006239

Datasets in Virtual Screening: starting off on the right foot

ResearchBlogging.org
As Niels Bohr said, prediction is very difficult, especially about the future. In case of computational modeling, the real grist of value is in prediction. But for any methodology to predict it must first be able to evaluate. Sound evaluation of known data (retrospective testing) is the only means to proceed to accurate prediction of new data (prospective testing).

Over the last few years, several papers have come out involving the comparisons of different structure-based and ligand-based methods for virtual screening (VS), binding mode prediction and binding affinity prediction. Every one of these goals if accurately achieved could lead to the saving of immense amounts of time and money for the industry. Every paper concludes that some method is better than other. For virtual screening for example, it has been concluded by many that ligand-based 3D methods are better than docking methods, and 2D ligand-based methods are at least as good if not better.

However, such studies have to be conducted very carefully to make sure that you are not biasing your experiment for or against any method, or comparing apples and oranges. In addition, you have to use the correct metrics for evaluation of your results. Failure to do either of these and other things can lead to erroneous or/and inflated or artificially enhanced results leading to fallible prediction.

Here I will talk about two aspects of virtual screening; choosing the correct dataset, and choice of evaluation metric. The basic problems in VS are false positives and false negatives and one wants to minimize the occurrence of these. Sound statistical analysis can do wonders for generating and evaluating good virtual screening data. This has been documented in several recent papers, notably one by Ant Nicholls from OpenEye. If you have a VS method, it's not of much use randomly picking a random screen of 100,000 compounds. You need to choose the nature and number of actives and inactives in the screen judiciously to avoid bias. Here are a few things to be remembered that I got from the literature:

1. Standard statistical analysis tells you that the error in your results depends upon the number of representatives in your sample. Thus, you need to have an adequate number of actives and inactives in your screening dataset. What is much more important is the correct ratio of inactives to actives. The errors inherent in choosing various such ratios have been quantified; for example, with an inactive:active ratio of 4:1, the error incurred is 11% more than that incurred by a theoretical ratio of infinite:1. For a ratio of 100:1 it's only 0.5% more than with infinite. Clearly we must use a good ratio of inactives to actives to reduce statistical error. Incidentally you can also increase the number of actives to reduce this error. But this is not compatible with real-life HTS where actives are (usually) very less, sometimes not more than 0.1% of the screen.

2. Number is one thing. The nature of your actives and decoys is equally important; simply overwhelming your screen with decoys won't do the trick. For example, consider a kinase inhibitor virtual screen in which the decoys are things like hydrocarbons and inorganic ions. In his paper, Nicholls calls distinguishing these decoys the "dog test", that is, even your dog should be able to distinguish them from actives (not that I am belittling dogs here). We don't want a screen that makes it too easy for the method to reject actives. Thus, simply throwing a screen of random compounds at your method might make it too easy for your method to screen actives and mislead.

We also don't want a method that rejects chemically similar molecules on the basis of some property like logP or molecular weight. For example consider a method or scoring function that is sensitive to logP, and suppose it is supplied with two hypothetical molecules which have an identical core and a nitrogen in the side chain. If one side chain has a NH2 and another one is N-alkylated where the alkyl is butyl, then there will be a substantial difference in logP between the two, and your method will fail to recognise them as "similar", especially from a 2D perspective. Thus, a challenging dataset for similarity based methods is one in which the decoys and actives are property-matched. Just such a dataset has been put together by Irwin, Huang and Shoichet- this is the DUD (Directory of Useful Decoys) dataset of property-matched compounds. In it, 40 protein targets and their corresponding actives have been selected. 36 property-matched decoys for every active have been chosen. This dataset is much more challenging for many methods that do well on other random datasets. For more details, take a look at the original DUD paper. In general, there can be different kinds of decoys; random, drug-like, drug-like and property-matched etc. and one needs to know exactly how to choose the correct dataset. With datasets like DUD, there is an attempt to provide possible benchmarks for the modeling community.

3. Then there is the extremely important matter of evaluation. After doing a virtual screen with a well-chosen dataset and well-chosen targets, how do you actually evaluate the results and put your method in perspective? There are several metrics but until now, the most popular way of doing this is by calculating enrichment and this is the way it has been done in several publications. The idea is simple; you want your top ranked compounds to contain the most number of actives. Enrichment is simply the fraction of actives found in a certain fraction of screened compounds. Ideally you want your enrichment curve to shoot up at the beginning, that is you want most (ideally all) of the actives to show up in the first 1% or so of your ranked molecules. Then you compare that enrichment curve to a curve (actually a straight line) that would stem from an ideal result.
The problem with enrichment is that it is a function of the method and the dataset, hence of the entire experiment. For example, the ideal straight line depends on the number of actives in the dataset. If you want to do a controlled experiment, then you want to make sure that the only differences in the results come from your method, and enrichment introduces another variable that complicates interpretation. Other failings of enrichment are documented in this paper.

Instead, what's recommended for evaluating methods are R.O.C curves.

Essentially, R.O.C curves can be used in any situation where one needs to distinguish signal from noise and boy, is there a lot of noise around. R.O.C curves have an interesting history; they were developed by radar scientists during World War 2 to distinguish the signal of enemy warplanes from the noise of false hits and other artifacts. In recent times they have been used in diverse fields; psychology, medicine, epidemiology, engineering quality control, anywhere where we want to pick the bad apples from the good ones. Thus, R.O.C curves simply plot the false positive (FP) rate against the true positive (TP) rate. A purely random result gives a straight line at 45 degrees implying that for every FL you get a TP- dismal performance. A good R.O.C curve is a hyperbola that shoots above the straight line, and a very useful measure of your method's performance is the Area Under the Curve (AUC). The AUC needs to be prudently interpreted; for instance an AUC of 0.8 means that you can discriminate a TP by assigning a higher score to it than to a FP in 8 out of 10 cases. Here's a paper discussing the advantages of R.O.C curves for VS and detailing an actual example.

One thing seems to be striking. The papers linked here and at other places document that R.O.C curves may currently be the single-best metric for measuring performance of virtual screening methods. This is probably not too surprising given that they have proved so successful in other fields.

Why should modeling be different? Just like in other fields, rigorous and standard statistical metrics need to be established for the field. Only then will the comparisons between different methods and programs commonly seen these days be valid. For this, as in other fields, experiments need to be judiciously planned (including choosing the correct datasets here) and their results need to be carefully evaluated with unbiased techniques.

It is worth noting that these are mostly prescriptions for retrospective evaluations. When confronted with an unknown and novel screen, which method or combination of methods does one use? The answer to this question is still out there. In fact some of the real-life challenges run contrary to the known scenarios. For example consider a molecular screen from some novel plant or marine sponge. Are the molecules in this screen going to be drug-like? Certainly not. Is this going to have the right ratio of actives to decoys? Who knows (the whole point is to find the actives). Is it going to be random? Yes. If so, how random? In all actual screenings, there are a lot of unknowns out there. But it's still very useful to know about "known unknowns" and "unknown unknowns", and retrospective screening and the design of experiments can help us unearth some of these. If nothing else, it indicates attention to sound scientific and statistical principles.

In later posts, we will take a closer look at statistical evaluation and dangers in pose-prediction including being always wary of crystal structures, as well as something I found fascinating- bias in virtual screen design and evaluation that throws light on chemist psychology itself. This is a learning experience for me as much or more than it is for anyone else.


References:

1. Hawkins, P.C., Warren, G.L., Skillman, A.G., Nicholls, A. (2008). How to do an evaluation: pitfalls and traps. Journal of Computer-Aided Molecular Design, 22(3-4), 179-190. DOI: 10.1007/s10822-007-9166-3

2. Triballeau, N., Acher, F., Brabet, I., Pin, J., Bertrand, H. (2005). . Journal of Medicinal Chemistry, 48(7), 2534-2547. DOI: 10.1021/jm049092j

3. Huang, N., Shoichet, B., Irwin, J. (2006). . Journal of Medicinal Chemistry, 49(23), 6789-6801. DOI: 10.1021/jm0608356