Field of Science

Showing posts with label HTS. Show all posts
Showing posts with label HTS. Show all posts

NSA, data uncertainty, and the problem of separating bad ideas from good ones

‘Citizenfour” is a movie about NSA whistleblower Edward Snowden made by journalist Laura Poitras, one of the two journalists Snowden contacted. It’s a gripping, professional, thought-provoking movie that everyone should consider watching. But this is especially so because at a deeper level, I think it goes to the heart not just of government surveillance but also of the whole problem of picking useful nuggets of data in the face of an onslaught of potential dross. In fact even the very process of classifying data as “nuggets” or “dross” is fraught with problems.

I was reminded of this problem as my mind went back to a piece on Edge.org by noted historian of technology George Dyson in which he takes government surveillance to task, not just on legal or moral grounds but on basic technical ones. Dyson’s concern is simple; when you are trying to identify that nugget of a dangerous idea from the morass of ideas out there, you are as likely to snare creative, good ideas in your net as bad ones. This may lead to a situation rife with false positives where you routinely flag – and, if everything goes right with your program, try to suppress – the good ideas. The problem arises partly because you don’t need to, and in fact cannot, flag every idea as “dangerous” or “safe” with one hundred percent accuracy; all you need to do is to get a rough idea.

“The ultimate goal of signals intelligence and analysis is to learn not only what is being said, and what is being done, but what is being thought. With the proliferation of search engines that directly track the links between individual human minds and the words, images, and ideas that both characterize and increasingly constitute their thoughts, this goal appears within reach at last. “But, how can the machine know what I think?” you ask. It does not need to know what you think—no more than one person ever really knows what another person thinks. A reasonable guess at what you are thinking is good enough.”

And when you are trying to get a rough idea, especially pertaining to someone’s complex thought processes, there’s obviously a much higher chance of making a mistake and failing the discrimination test.
The problem of separating the wheat from the chaff is encountered by every data analyst: for example, drug hunters who are trying to identify ‘promiscuous’ molecules – molecules which will indiscriminately bind to multiple proteins in the body and potentially cause toxic side effects – have to sift through lists of millions of molecules to find the right ones. They do this using heuristic rules of thumb which tell them what kinds of molecular structures have been promiscuous in the past. But the past cannot foretell the future, partly because the very process of defining these molecules is sloppy and inevitably captures a lot of ‘good’, non-promiscuous, perfectly druglike compounds. This problem with false positives applies to any kind of high-throughput process founded on empirical rules of thumb; there are always bound to be several exceptions. The same problem applies when you are trying to sift through millions of snippets of DNA and assigning causation or even correlations between specific genes and diseases.
What’s really intriguing about Dyson’s objection though is that it appeals to a very fundamental limitation in accomplishing this discrimination, one that cannot be overcome even by engaging the services of every supercomputer in the world.
Alan Turing jump-started the field of modern computer science when he proved that even an infinitely powerful algorithm cannot determine whether an arbitrary string of code represents a provable statement (the so-called ‘Decision Problem’ articulated by David Hilbert). Turing provided to the world the data counterpart of Kurt Gödel’s Incompleteness Theorem and Heisenberg’s Uncertainty Principle; there is code whose truth or lack thereof can only be judged by actually running it and not by any preexisting test. Similarly Dyson contends that the only way to truly distinguish good ideas from bad is to let them play out in reality. Now nobody is actually advocating that every potentially bad idea should be allowed to play out, but the argument does underscore the fundamental problem with trying to pre-filter good ideas from bad ones. As he puts it:
“The Decision Problem, articulated by Göttingen’s David Hilbert, concerned the abstract mathematical question of whether there could ever be any systematic mechanical procedure to determine, in a finite number of steps, whether any given string of symbols represented a provable statement or not.

The answer was no. In modern computational terms (which just happened to be how, in an unexpected stroke of genius, Turing framed his argument) no matter how much digital horsepower you have at your disposal, there is no systematic way to determine, in advance, what every given string of code is going to do except to let the codes run, and find out. For any system complicated enough to include even simple arithmetic, no firewall that admits anything new can ever keep everything dangerous out…

There is one problem—and it is the Decision Problem once again. It will never be entirely possible to systematically distinguish truly dangerous ideas from good ones that appear suspicious, without trying them out. Any formal system that is granted (or assumes) the absolute power to protect itself against dangerous ideas will of necessity also be defensive against original and creative thoughts. And, for both human beings individually and for human society collectively, that will be our loss. This is the fatal flaw in the ideal of a security state.”

In one sense this problem is not new since governments and private corporations alike have been trying to separate and suppress what they deem to be dangerous ideas for centuries; it’s a tradition that goes back to book burning in medieval times. But unlike a book which you can at least read and evaluate, the evaluation of ideas based on snippets, indirect connections, Google links and metadata is tenuous at best and wildly unlikely to accurately succeed. That is the fundamental barrier that agencies who are trying to determine thoughts and actions based on Google searches and Facebook profiles are facing, and it is likely that no amount of sophisticated computing power and data will enable them to solve the general problem.
Ultimately whether it’s government agencies, drug hunters, genomics experts or corporations, the temptation to fall prey to what writer Evgeny Morozov calls “technological solutionism” – the belief that key human problems will succumb to the latest technological advances – can be overpowering. But when you are dealing with people’s lives you need to be a little more wary of technological solutionism than when you are dealing with the latest household garbage disposal appliance or a new app to help you find nearby ice cream places. There is not just a legal and ethical imperative but a purely scientific one to treat data with respect and to disabuse yourself of the notion that you can completely understand it if only you threw more manpower, computing power and resources at it. A similar problem awaits us in the application of computation to problems in biotechnology and medicine.

At the end of his piece Dyson recounts a conversation he had with Herbert York, a powerful defense establishment figure who designed nuclear weapons, advised presidents and oversaw billions of dollars in defense and scientific funding. York cautions us to be wary of not just Eisenhower’s famed military-industrial complex but of the scientific-technological complex that has aligned itself with the defense establishment for the last fifty years. With the advent of massive amounts of data this alignment is honing itself into an entity that can have more power on our lives than ever before. At the same time we have never been in greater need of the scientific and technological tools that will allow us to make sense of the sea of data that engulfs. And that, as York says, is precisely the reason why we need to beware of it.
“York understood the workings of what Eisenhower termed the military-industrial complex better than anyone I ever met. “The Eisenhower farewell address is quite famous,” he explained to me over lunch. “Everyone remembers half of it, the half that says beware of the military-industrial complex. But they only remember a quarter of it. What he actually said was that we need a military-industrial complex, but precisely because we need it, beware of it. Now I’ve given you half of it. The other half: we need a scientific-technological elite. But precisely because we need a scientific-technological elite, beware of it. That’s the whole thing, all four parts: military-industrial complex; scientific-technological elite; we need it, but beware; we need it but beware. It’s a matrix of four.”

It’s a lesson that should particularly be taken to heart in industries like biotechnology and pharmaceuticals, where standard, black-box computational protocols are becoming everyday utilities of the trade. Whether it’s the tendency to push a button to launch a nuclear war or to sort drugs from non-drugs in a list of millions of candidates, temptation and disaster both await us at the other end.

Anticancer drugs form colloidal aggregates and lose activity

Over the last few years, one of the most interesting findings in drug screening and testing at a preclinical level has been the observation that many drugs form colloidal aggregates under standard testing conditions and nonspecifically inhibit target proteins which they otherwise would not affect. This are large aggregates, a hundred nanometers or more in diameter, and they cause proteins to stick and partially unfold, creating the illusion of inhibition. This leads to false positives, especially in high-throughput screening protocols. And these false positives can be absolutely rampant.

What's striking is the sheer ubiquity of this phenomenon which has been observed with all kinds of drugs under all kinds of conditions; while the initial observation was limited to isolated protein-based assays, the phenomenon has also been seen in simulated gastric fluids and in the presence of many different kinds of proteins like serum albumin which are found inside the body. The colloid spirit seems to emphatically favor a shotgun approach.

Now a team led by the brother-sister duo Brian and Molly Shoichet (UCSF and Toronto) has found something that should give drug testers further pause for thought; they see some bestselling anticancer drugs forming colloids (shown above) in cell-based assays to an extent that actually diminishes their activity, leading not to false positives but to false negatives. They test seven known anticancer drugs in cell assays both under known colloid forming conditions along with conditions that break the colloids up. This is not as easy as it sounds since it involves adding a detergent which would usually be too toxic to cells; fortunately in this case they find the right one. Another interesting finding is the re-evaluation of a popular dye used to study "leaky" cancer blood vessels; unlike the previously proposed mechanism, the current study seems to suggest that the dye too forms large aggregates and nonspecifically inhibits the protein serum albumin.

The testing essentially reveals that the drugs when they form colloids basically show activity that's so low as to be negligible and equivalent to the controls. That's a self-(un)proclaimed false negative. Now anybody who deals with error analysis knows that false negatives are fundamentally worse than false positives since by definition they cannot even be detected. The present study raises the pertinent question; how many promising drugs might we be missing because they form aggregates and lower the observed response in cells? And since the colloid forming phenomenon has been shown to be so ubiquitous, could it possibly be influencing the mechanism of action of all kinds of drugs inside the body? And in what ways? It's a fascinating question, and one of those that continues to make basic research in drug discovery still so interesting.
Image source and credit: ACS

Screening probes and probing screens

ResearchBlogging.orgHigh Throughput Screening (HTS), with all its strengths and limitations, is still the single-best way to discover novel interesting molecules in drug discovery. Thomas Kodadek of Scripps Florida has an interesting article on screening in the latest issue of Nat. Chem. Biol which is a special issue on chemical probes.

Kodadek talks about the very different properties required for drugs and probes and the limitations and unmet needs in current HTS strategies. He focuses on mainly two kinds of screening; functional assays and binding assays. The former can consist of phenotypic screening wherein one is only interested in a particular cellular response. This is more useful for drugs. However for probes, target selectivity is important and one must have knowledge of the target. HTS hits can hit all kinds of protein targets, thus making it hard to find out if your compounds are being selective. Mutagenesis and siRNA studies can shed light on target selectivity but this is not easy to do.

One of the possible solutions Kodadek suggests to circumvent the problem of gauging selectivity is to use binding assays instead of functional assays. He notes a pretty clever idea used in binding assays; that of throwing in cell extracts with miscellaneous proteins that could mop up greasy, non-selective compounds. This strategy cannot be easily used in functional assays. Binding assays are also typically less expensive than functional assays.

There is also a discussion of some of the very practical problems associated with screening. Screening typically has low hit rates and more importantly, hits from screening are not leads. You usually need a dedicated team of synthetic chemists to make systematic SAR modifications to these hits to optimize them further. As the author says, few synthetic chemists wish to serve as SAR facilities for their biologist colleagues. Plus it is not easy to lure industrial chemists to serve this function in academia (although the present economic climate may have made this easier). Thus, biologists with no synthetic background need to be able to make at least some modifications to their hits. For this purpose Kodadek suggests the use of modular molecules with easily available building blocks which can be cheaply and easily connected together by relatively inexperienced chemists; foremost in his recommendations are peptoids, N-substituted oligoglycines which are biologically active and easy to synthesize. Thus, if libraries for screening are enriched in such kinds of molecules, it could make it easy for biologists without access to sophisticated synthetic chemists and facilities to cobble together leads. Of course this would lead to a loss of diversity in the libraries, but that's the tradeoff necessary for going down the long road from hit to lead.

Lastly, Kodadek briefly talks about prospects for screening in academia. Academic drug discovery is gradually becoming more attractive with the recent long lull in industry. However academic scientists are typically not very familir with the post-synthesis optimization of drugs including optimization of metabolic properties, bioavailability and PK. Academic scientists who can pursue such studies or partner with DMPK contracting companies may be paid back their dues.

One topic which Kodadek does not mention is virtual screening (VS). VS can complement HTS and at least some studies indicate that the rate of success in VS can match, if not exceed, that in HTS. In addition, new ligand-based methods which use properties such as molecular shape to screen for compounds similar to given hits can also valuably complement HTS follow up studies.

Screening is still the best bet for discovering new drugs, but hit rates are typically still very low (1% would be a godsend). Only a concerted effort at designing libraries, ensuring selectivity and synthetic accessibility will make it easier.

Kodadek, T. (2010). Rethinking screening Nature Chemical Biology, 6 (3), 162-165 DOI: 10.1038/nchembio.303

The same and not the same: more aggregates in HTS

ResearchBlogging.org

High-throughput screening (HTS) is now a mainstay of drug discovery and usually the starting point for most drug discovery projects. Industry usually has a lot of resources invested in HTS and therefore needs to be aware of false positives and false negatives that could hamper useful results and lead one down an erroneous path.

Among the many factors responsible for false positives in HTS, one of the most startling and important factors recently unearthed is the non-specific and potent inhibition of enzymes by aggregates of molecules occurring under typical assay conditions. These aggregates are large enough to be observed under a microscope and to be detected by dynamic light scattering. The aggregates adsorb enzyme molecules on their surface, and one of the best tests for detecting their presence is to re-run the enzyme assay under high detergent concentration. High detergent concentrations usually break up the aggregates and lead to a loss of potent inhibition. The phenomenon of aggregation-based inhibition was accidentally discovered by Brian Shoichet's group at UCSF and has been comprehensively explored by him and his students in a series of papers throughout the last decade, although much is still to be known about the exact physical nature of these aggregates. The reason why this has become a big deal is because it has been observed in an unusual number of cases, which leads to the suspicion that much effort might have been already expended in drug discovery campaigns in pursuing such false leads.

In a recent paper, Shoichet and Craik's groups at UCSF accidentally discovered aggregate-based inhibition in discovering inhibitors for the enzyme cruzain which is a part of the metabolic machinery of the parasite responsible for Chagas disease. The authors had started with an initial hit from a virtual screening campaign and were engaged in the usual process of modifying the hit based on SAR. The initial tinkering led to a series of oxadiazole inhibitors which exhibited potent inhibition of cruzain.

However, many of these molecules failed to show activity in cell-based assays. Such a discrepancy between enzyme and cell-based assays can be traced back to many reasons including poor permeability. But in this particular case, kinetic measurements hinted at aggregates of the oxadiazoles that were inhibiting the enzyme. At this stage it was also discovered that unlike the initial hits series, the oxadiazole series had been accidentally assayed under low detergent conditions. The molecules also inhibited another intensely studied enzyme in the Shoichet group- AmpC beta-lactamase. The quintessential test for aggregate-based inhibition, namely increasing the concentration of detergent (Triton in this case), also proved positive confirming the phenomenon. Interestingly the initial set of hit molecules also seemed to exhibit this phenomenon but only in case of AmpC lactamase and not in case of cruzain. In case of cruzain, experiments with differing detergent concentrations proved that the initial set of molecules were equally potent under both conditions, while the oxadiazoles lost activity under high detergent conditions, indicating divergent modes of inhibition between the two sets of molecules.

Finally, note that the aggregation-based inhibition would likely have not been discovered if the oxadiazole series had been assayed under the same low detergent condition as the initial hit series. What seemed like similar molecules turned out to behave very differently under different assay conditions. Sometimes mistakes can reward you with unexpected treasures, and similarity needs to be pried out from the eye of the experimenter. Never underestimate the importance of going wrong (of course revealed only in retrospect).

As the authors narrate, the moral of such studies should not be lost on medicinal chemists, who usually interpret high and low potency of related molecules based on local structural features like hydrogen bonding, electrostatics and hydrophobicity. Aggregation-based enzyme inhibition proves that chemists have to look beyond single molecule structural features toward supramolecular features of several molecules that are interacting with each other. Chemists regularly engaged in HTS campaigns might well keep this valuable piece of advice in mind. Scientific enumeration, it seems, has to always be done at several different levels.

Note: Apologies to Prof. Roald Hoffmann for appropriating the title

Ferreira, R., Bryant, C., Ang, K., McKerrow, J., Shoichet, B., & Renslo, A. (2009). Divergent Modes of Enzyme Inhibition in a Homologous Structure−Activity Series Journal of Medicinal Chemistry DOI: 10.1021/jm9009229

The anti-question, or when bias can be a good thing

A recent publication indicates that more bias in the form of natural product scaffolds not yet synthesized could improve hit rates in screening

ResearchBlogging.org

Most drug discovery projects are inaugurated with some kind of screening campaign where millions of molecules are screened against a biological target. Even though the hit rate from High-Throughout Screening (HTS) can be quite low, HTS still provides one of the best starting points to discover interesting new structures that display biological activity. In spite of this, there is frequent disappointment at the low rates from HTS which could be as low as 0.05%.

But instead of focusing on the low hit rate from HTS, what if we express surprise that this hit rate is actually high? This thought takes me into a slight digression. In his remarkable book The Black Swan, the author Nassim Nicholas Taleb talks about an "anti-library", the set of all books you have not read. The anti-library is in some ways more important than your library because it really tells you what you are ignorant about.

Similarly we can define an "anti-question". The anti-question is a question opposite to one which we might usually ask. So instead of asking; "Why is this drug specific for this protein?", we could ask "Why is this drug not hitting other proteins?". The value of the anti-question is that it forces us to analyze and evaluate things that we otherwise may not and enables us to think outside the box. As the wise doctor constantly exhorts detective Sponer in "I Robot" to get to the all-important right question, so it could be important to get to the right anti-question.

In the context of HTS, the anti-question actually turns out to be logical. Instead of asking, "Why is the hit rate from HTS so low"?, one should ask "Given the number of small molecules in small-molecule space (~10*60) compared to the extremely low number typically screened in HTS campaigns (10*6), why should we get any hits from HTS at all?". Even narrowing down the unimaginably large small-molecule universe to more drug-like or lead-like entities, we still run into a numbers paradox since even this number is orders of magnitude greater than what is usually screened.

In their most recent paper, Brian Shoichet and his team ask this important anti-question, and it leads them down an interesting road. Most campaigns that screen libraries focus on readily available commercial compounds and fragments that can be synthesized by organic chemists. This bias in turn reflects what has been more or less synthetically accessible through more than a hundred years of synthesis. Compared to this, the Kyoto Encyclopedia of Genes and Genomes (KEGG) contains metabolites whose structures are untainted by the minds of organic chemists. These are scaffolds among secondary metabolites and natural products that have simply been found.

There is another set of structures; the Generated Database (GDB), a theoretical set which contains all possible molecules containing less than 11 heavy atoms consisting of first-row elements (C, O, N, F). This number is not as large as may be imagined and amounts to about 26 million. In the study the authors essentially compare the set of purchasable or commercial KEBB compounds found in their own annotated library called ZINC with the GDB. They use a similarity measure called a Tanimoto coefficient derived from 2D fingerprint comparison to accomplish this. 2D fingerprints use different kinds of protocols for breaking up a molecule into bit strings and then compare bit strings by distances and atom types.

The comparison indicates something interesting; the compounds in the purchasable set are much more similar to the KEBB compounds than are the compounds from the rest of the GDB. In other words, purchasable compounds contain scaffolds that are biased towards those in the KEBB. This is a good thing, since metabolites are usually primed by nature to show at least some biological activity. Another noteworthy finding was that the bias also increased with molecular size, as compounds became more drug-like or lead-like in terms of size.

However, the more surprising and useful observation was that there are hundreds of scaffolds in the KEBB that are notpresent in the commercial library. The authors also do this comparison for other popular commercial libraries designed specifically for screening and find a similar result. The bottom line; while synthesized commercial libraries of molecules show a bias toward natural products and metabolites, there are also several natural product scaffolds that are not found in these libraries.

So what is the prescription? Introduce further bias! The compounds in the KEGG are more or less optimized for biological activity. If their scaffolds are not yet present in the commercial libraries, organic chemists should go ahead and focus on synthesizing these scaffolds and adding them to screening libraries. More such scaffolds could increase the hit rate in HTS by enriching libraries in biologically relevant scaffolds. Of course the usual caveats of false positives and promiscuous compounds should be kept in mind, and it's also not clear that proteins like kinases which are optimized to bind certain core scaffold structures would greatly benefit from these diverse scaffolds. But in terms of unmined drug space, introducing such further bias would be beneficial.

This study again goes to show the possibilities for finding new stars in the constellations and galaxies of the drug universe. Hopefully the universe will keep on expanding.

Hert, J., Irwin, J., Laggner, C., Keiser, M., & Shoichet, B. (2009). Quantifying biogenic bias in screening libraries Nature Chemical Biology DOI: 10.1038/nchembio.180

A rash of molecular personalities

ResearchBlogging.org
Just like human beings, molecules have personalities. And just like human beings, they display those personalities best when they react to a stimulus. For a medicinal chemist, one such stimulus is HTS where one can identify different flavors of molecules through their interaction with protein targets. But this is not always done, and quantitative analysis of molecules in HT screens is lacking. Clearly such analyses will help to identify compositions of such screens and give insight into future screens.

In his latest offering, Brian Shoichet does just that. He and his group set out to identify essentially every molecular character from a colorful screen of about 70000 molecular personalities applied to ampicillin resistant beta-lactamase. Their results are surprising.

Out of 70000, about 1274 showed activity. Shoichet has already extensively documented the alarming frequency of aggregate-forming molecules in common HTS screens. It's a very substantial contribution from his laboratory. In this case, 1204 (95%) of the 1274 turned out to be inhibiting the enzyme through non-specific aggregation. This can be found out by adding detergent, which breaks up the aggregates and gets rid of the spurious activity.

So now there were 70 detergent-insensitive compounds. How many of these were true, reversible binders? 25 of these were beta-lactams, and since they are covalent modifiers of the enzyme and known chemical scaffolds, they were not considered further. So out of the remaining ones, 25 were re-synthesized and were found to be false positive through lack of reproducible activity. There were now 20 non beta-lactams. Out of these 9 were again found to be aggregators- the earlier screen had skipped them because of low detergent concentration.

That left 12 molecules. After some more scrutiny, these were all found to be covalent, irreversible modifiers of the enzyme. A neat and simple trick can be used to identify covalent modification; mass spectra of the modified enzyme are clearly different from the apo enzyme.

So how many non-covalent, reversible inhibitors of beta-lactmase were found? Zero.

To shed some more light on this strange phenomenon, the authors turned to docking with DOCK. To make sure the program can identify reversible binders, some known binders were seeded among the unknown binders. After docking and observing that the first 500 hits contained the known binders, 16 out of these 500 compounds were selected based on structural diversity and then assayed. Interestingly, two among these compounds were found to inhibit the enzyme at IC50 values of >100 µM. No wonder the initial screen had missed these phthalimide culprits- the highest concentration in the screen was 30 µM.

In other studies, they also did some SAR on the hits and verified the docking poses by obtaining crystal structures. There are other interesting details in the paper.

But even if the study did not unearth reversible, potent, novel binders, it is of course still very instructive. It tells us about the variety of beasts existing in HTS. It also again sheds light on docking as a valuable complement to HTS. In this case, 70000 compounds may been too less for assaying, and 30 µM must have been two low a threshold for finding hits. In any case, higher thresholds for testing are limited by practical difficulties, including material availability and solubility. But what is valuable is that given due effort, we can identify compounds that give false positive results in screens through novel mechanisms- in this case by aggregation (detected by detergent addition) and by covalent modification (detected by mass spec)

There are clearly some notorious and dirty candidates in HTS screens- more than everyone would be comfortable with- and this study provides a good model for being on one's guard and seeking to identify them as thoroughly as possible. When we lay down the red carpet, we want only the cream of the crop, not asses disguised as lions.

Babaoglu, K., Simeonov, A., Irwin, J.J., Nelson, M.E., Feng, B., Thomas, C.J., Cancian, L., Costi, M.P., Maltby, D.A., Jadhav, A., Inglese, J., Austin, C.P., Shoichet, B.K. (2008). Comprehensive Mechanistic Analysis of Hits from High-Throughput and Docking Screens against ÃŽ²-Lactamase. Journal of Medicinal Chemistry DOI: 10.1021/jm701500e

Interview with Brian Shoichet: aggregation-induced inhibition

Image Hosted by ImageShack.us

Ok, now that we have gotten past the Nobel mania (or maybe not; go Somorjai), we can hopefully come back to real life. I was reading an interview with Brian Shoichet, who is one of the most promising stars in the areas of screening, docking, and structure-based design. He has gotten his fingers in many pies, both computational and experimental.

However, it was somebody's comment about the pharmaceutical industry thinking that "Shoichet deserves a heroes prize" that got me looking at his work, and I quickly learnt the reasons for that quote. As we all know, one of the biggest or perhaps the biggest problem facing HTS in industry is false positives. A lot of times, molecules that are found to be active in an assay fail to be active later. If industry could weed out such nuisances ahead of time, a lot of time, money and energy could be saved.

Shoichet, after a lot of interesting initiation and investigation, came up with one simple reason for why molecules may be showing false colours; because they form colloidal aggregates that somehow inhibit the proteins in the assay. If these are broken up say with detergent, the individual molecules no longer show activity. Thus, a relatively simple physical phenomenon is responsible for these molecules showing false activities. Such molecules were detected in earlier assays by some characteristics, mainly very steep dose-response curves and flat SAR; changes in structure usually causing very small changes in activity. They are also often promiscuous inhibitors. But nobody knew what was exactly happening and all the analysis was post-"mortem".

The first step in Shoichet's lab was the elucidation of this aggregation-induced inhibition. The aggregation can be detected with dynamic light scattering (DLS). The more challenging and useful step is to be able to come up with a list of chemical scaffolds that are likely to show this phenomenon, so that one can watch out for them beforehand. Before that, one would also need to know the exact mechanism of aggregation-based inhibition. In case of some molecules, there is some structural correlation, flat aromatic dye-like molecules being prone to aggregation for example by stacking. But many other scaffolds seem more diverse and at first glance show no common functionalities. Ths phenomenon is linked by common physical forces, not chemical ones. The details are not known but continue to be worked out.

Shoichet's lab continues to make progress, and he has recently come up with a screen for detecting such aggregation-based inhibitors (DOI: 10.1021/jm061317y). There are two major conclusions from the study; first, that breaking up aggregates with detergents can be a good way of identifying them, and secondly that aggregation may be a much more common phenomenon for false positives in screens than was thought before. This fact may be extremely significant for industry and could potentially save a lot of time, money and labour beforehand.

In other quite different work (DOI: 10.1038/nature05981), Shoichet also made the cover of Nature, when he used docking and structure-based design to predict the function for an enzyme whose function was unknown, based on substrate docking and analysis. The strategy used was quite clever; docking thousands of high-energy forms of metabolites rather than the metabolites themselves to know which ones would optimally interact with the active site. In this particular case, the "optimum interaction" pointed to a deamination, and the protein of unknown function indeed experimentally turned out to be a good deaminase.

All in all, a very promising chemist and I believe one to watch out for. Unfortunately, the interview itself is published in the journal Assay and Drug Development Technologies, not one which libraries usually subscribe to (I got it through ILL). But here's the DOI anyway (DOI: 10.1089/adt.2007.9996)

Also, again, check out his Colbert-style interview on youtube.

Turning a false-positive into an active

People who deal with molecular recognition are well aware of what difference a small modification to a molecule can make. Just today I was attending a talk by a chemist who binds small molecules to RNA aptamers. He showed an aptamer that binds theophylline with 10,000 fold more affinity by caffeine- a huge difference in binding affinity for a molecule differing by only one methyl.

So it is also for medicinal agents, as demonstrated below for an example from the cited study. People who do screening must always have this nagging doubt about false positives; what if there is only a slight modification to a false positive that will convert it into an active?

Image Hosted by ImageShack.us

Bill Jorgensen's group has done a similar study for an anti-RT HIV inhibitor. He first did similarity searching with the Maybridge library based on six known NNRTI inhibs of RT. Based on this, he found a couple of molecules in the library which he then docked into the active site of RT using the program GLIDE. Along with the six known inhibitors which scored at the top as binders, he also found one from the library. GLIDE had already been benchmarked by reproducing known crystallographic conformations.

However, when they tested this GLIDE ranked molecule against HIV, it was disappointingly inactive. On the other hand, perhaps, since GLIDE had docked it up there with the known actives, there might be a small modification that one could make to it which would inject some activity in it? Jorgensen's group used a program that they have developed named BOMB, which basically docks a molecule in an active site, and then grows appendages to it to see if it would make a difference in the binding affinity. BOMB tried out combinations of different groups on the phenyl ring of the molecule, scored the resulting structures using its energy function, and finally settled on one particular modified structure- also filtered by logP values and other Lipinski considerations- that eventually gave an IC50 of 300 nM. Not a fantastic number, but good enough to pursue as a lead.

Also noteworthy in the paper is a short discussion of another publication where a similar structure was published. According to the authors, the other authors assayed the wrong compound. Heh.

Reference:
From Docking False-Positive to Active Anti-HIV Agent
Gabriela Barreiro, Joseph T. Kim, Cristiano R. W. Guimarães, Christopher M. Bailey, Robert A. Domaoal, Ligong Wang, Karen S. Anderson, and William L. Jorgensen
Web Release Date: 06-Oct-2007; (Article) DOI: 10.1021/jm070683u