Field of Science

Showing posts with label medicinal chemistry. Show all posts
Showing posts with label medicinal chemistry. Show all posts

Minority Report Meets Drug Discovery: Intelligent Gestural Interfaces and the Future of Medicine

In 2002, Steven Spielberg’s Minority Report introduced one of the most iconic visions of the future: a world where data is accessed, manipulated, and visualized through an immersive, gestural interface. The scene where Tom Cruise’s character, Police Chief John Anderton, swiftly navigates vast amounts of visual information by simply swiping his hands through thin air is not just aesthetically captivating but also hints at the profound potential of such interfaces in real-world applications—particularly in fields as complex as drug discovery. Just like detective work involves combining and coordinating data from disparate sources such as GPS, real-time tracking, historical case studies, image recognition and witness reports, drug discovery involves integrating data from disparate sources like protein-ligand interactions, patent literature, genomics and clinical trials. Today, advancements in augmented reality (AR), virtual reality (VR), and high-performance computing (HPC) offer the tantalizing possibility of a similar interface revolutionizing the way scientists interact with multifactorial biology and chemistry datasets.

This post explores what a Minority Report-style interface for drug design would look like, how the seeds of such a system already exist in current technology, and the exciting potential this kind of interface holds for the future of drug discovery.

The Haptic, Gestural Future of Drug Discovery

Perhaps one of the most memorable aspects of Minority Report is the graceful, fluid way in which Tom Cruise’s character interacts with a futuristic interface using only his hands. With a series of quick, intuitive gestures, he navigates through complex data sets, zooming in on images, isolating key pieces of information, and piecing together the puzzle at the center of the plot. The thrill of this interface comes from its speed, accessibility, and above all, its elegance. Unlike the clunky, keyboard-and-mouse-driven systems we’re used to today, this interface allows data to be accessed and manipulated as effortlessly as waving a hand.

In drug discovery, such fluid navigation would be game-changing. As mentioned above, the modern scientist deals with a staggering amount of information: genomics data, chemical structures, protein-ligand interactions, toxicity reports, and clinical trial results, all coming from different sources. The ability to sweep through these datasets with a flick of the wrist, pulling in relevant data and discarding irrelevant noise in real-time, would make the process of drug design not only more efficient but more dynamic and creative. Imagine pulling together protein folding simulations, molecular docking results, and clinical trial metadata into a single, interactive, 3D workspace—all by making precise, intuitive hand movements like Tom Cruise.

The core of the Minority Report interface is its gestural and haptic nature, which would be crucial for translating such a UI into the realm of drug design. By introducing haptic feedback into the system—using vibrations or resistance in the air to simulate touch—a researcher could "feel" molecular structures, turning abstract chemical properties into tactile sensations. Imagine "grabbing" a molecule and feeling the properties of its surface—areas of hydrophobicity, polarity, or charge density—all while rotating the structure in mid-air with a flick of your wrist. Like an octopus sensing multiple inputs simultaneously, a researcher would be the active purveyor of a live datastream of multilayered data. This tactile feedback could become a new form of data visualization, where chemists and biologists no longer rely solely on charts and numbers but also on physical sensations to understand molecular behavior. The experience would translate to an entirely new dimension of interacting with molecular data and models, making it possible to “sense” molecular conformations in ways that are impossible with current 2D screens.

Such a haptic interface would also make the process more accessible. Students and new researchers in drug discovery would quickly learn how to navigate and manipulate datasets through a gestural UI. The muscle memory developed through these natural, human movements would make the learning curve less steep, transforming the learning experience into something more akin to a hands-on laboratory session rather than an abstract, numbers-on-a-screen challenge. Drug discovery and molecular design would be democratized.

Swiping Through Multifactorial Datasets

One of the most exciting possibilities of a Minority Report-style UI in drug discovery is its ability to merge multifactorial datasets, making complex biology and chemistry data "talk" to each other. In drug discovery, researchers deal with data from various domains, — genomics, proteomics, cheminformatics, clinical data etc. — each of which exists in its own silo; any researcher in the area would relate to the pain of integrating these very different databases, an endeavor that requires a significant amount of effort and specialized software. Currently, entire IT departments are employed to these ends. A futuristic UI could change that entirely.

Imagine a scientist swiping through an assay dataset with one hand, while simultaneously bringing in chemical structure data and purification data on stereoisomers with the other. Perhaps throw in a key blocking patent and gene expression data. These diverse datasets could then be overlaid in real time, with machine learning algorithms providing instant insights into correlations and potential drug candidates. For instance, one swipe could summon a heat map of gene expression related to a disease, while another flick could display how a particular small molecule binds to a target protein implicated in that disease. A few more gestures could allow the scientist to access historical drug trials and toxicity data as well as patent data, immediately seeing if any patterns emerge. The potential here is enormous: combining these multifactorial datasets in such a seamless, visual way would enable researchers to generate hypotheses on the fly, test molecular interactions in real-time, and identify the most promising drug candidates faster than ever before.

The Seeds Are Already Here: AR, VR, and High-Performance Computing

While this vision seems futuristic, the seeds of this interface already exist in today's technology. Augmented reality (AR) and virtual reality (VR) platforms are rapidly advancing, providing immersive environments that allow users to interact with data in three dimensions. AR devices like Microsoft's HoloLens and VR systems like the Oculus Rift already provide glimpses of what a 3D drug discovery workspace might look like. For example, AR could be used to visualize molecular structures in real space, allowing researchers to walk around a protein or zoom in on a ligand-binding site as if it were floating right in front of them.

At the same time, high-performance computing (HPC) is already pushing the limits of what we can do with drug discovery. Cloud-based platforms provide immense computing power that can process large datasets, while AI-driven software accelerates the pace of molecular docking simulations and virtual screening processes. Combining these technologies with a Minority Report-style interface could be the key to fully realizing the potential of this future workspace.

LLMs as Intelligent Assistants

While the immersive interface and tactile data manipulation are powerful, the addition of large language models (LLMs) brings an entirely new layer of intelligence to the equation. In this vision of drug discovery, LLMs would serve as intelligent research assistants, capable of understanding complex natural language queries and providing context-sensitive insights. Instead of manually pulling in data or running simulations, researchers could ask questions in natural language, and the LLM would retrieve relevant datasets, compute compound properties, run analyses, and even suggest possible next steps. Even if a researcher could summon up multiple datasets by swiping in an interactive display, they would still need an LLM to answer questions pertaining to cross-correlations between these datasets.

Imagine a researcher standing in front of an immersive display, surrounded by 3D visualizations of molecular structures and genomic data. With a simple voice command or text prompt, they could ask the LLM, “Which compounds have shown the most promise in targeting this specific binding site?” or “What genetic mutations are correlated with resistance to this drug?” or even fuzzier questions like “What is the probability that this compound would bind to the site and cause side effects?”. The LLM would then comb through millions of datasets, both existing and computed, and instantly provide answers, suggest hypotheses, or even propose new drug candidates based on historical data.

Moreover, LLMs could help interpret complex, multifactorial relationships between datasets. For example, if a researcher wanted to understand how a particular chemical compound might interact with a genetic mutation in cancer cells, they could ask the LLM to cross-reference all available data on drug resistance, molecular pathways, and previous clinical trials. The LLM could provide a detailed, synthesized response, saving the researcher countless hours of manual research and allowing them to focus on making creative, strategic decisions.

This kind of interaction would fundamentally change the way scientists approach drug discovery. No longer would they need to rely solely on their own ability to manually search for and interpret data. Instead, they could work in tandem with an intelligent, AI-driven system that helps them navigate the immense complexity of modern drug design. With the right interface, researchers could manipulate massive amounts of drug discovery data in real-time, powered by already existing HPC infrastructure.

Current challenges

While this vision of an all-in-one molecular design interface sounds promising, we would be remiss in not mentioning some familiar current challenges. Data is still highly siloed, even within organizations, and inter-organizational data sharing is still bound by significant legal, business and technological challenges. While AR and VR are now being democratized through increasingly cheap headsets and software, the experience is not as smooth as we would like, and bringing in disparate data sources into the user experience remains a problem. In the future, common API formats could become a game changer. Finally, LLMs still suffer from errors and hallucinations. Having a human in the loop would be imperative in overcoming these limitations, but there is little doubt that the sheer time-saving and consolidation they enable, along with the ability to query data in natural language, would make their use not just important but inevitable.

A Future of Instant, Integrated Data at Your Fingertips

The promise of a Minority Report-style interface for drug discovery lies in its ability to make data instantly accessible, integrated, and actionable. By swiping and gesturing in mid-air, scientists would no longer be constrained by traditional input methods, unlocking new levels of creativity and efficiency. This kind of interface would enable instant access to everything from raw molecular data to advanced machine-learning models predicting the efficacy of new drug candidates.

We can image a future where a drug designer could pull up decades of research on a specific disease, instantly overlay that with genomic data, and compare it with molecular screening results—all in a 3D, immersive environment. The heightened experience would make it possible to come up with radically new hypotheses about target engagement, efficacy and toxicity in short order. Collaboration would also reach new heights, as teams across the world interacted in the same virtual workspace, manipulating the same data sets in real time, regardless of their physical location. The interface would enable instant brainstorming, rapid hypothesis generation and testing, and seamless sharing of insights. The excitement surrounding such a future is palpable. By blending AR, VR, HPC, and LLMs, we can transform drug discovery into an immersive, highly interactive, and profoundly intuitive process.

Let the symphony start playing.

Big Trouble in Little Synthetic Organic Chemistry?

Michael Rafferty who teaches in the Department of Medicinal Chemistry at the University of Kansas has a thought-provoking article in the Journal of Medicinal Chemistry in which he questions whether it's time to reinvent the model for training academic scientists in graduate programs to better equip them for the complexity and rigors of modern drug discovery. His target is the cadre of synthetic organic chemists who for decades have functioned as the indispensable backbone of the pharmaceutical industry. The title of the article - "No Denying It: Medicinal Chemistry Training Is In Big Trouble" - should be self-explanatory, in case anyone is wondering where exactly the author's sentiments lie on the topic.

Even today when you say that someone is a "medicinal chemist" it usually means someone who is trained as a synthetic organic chemist, who either goes into the lab and makes molecules himself or herself or who directs other people to do the same. Rafferty is asking whether the decades-old standard of recruitment into medicinal chemistry groups in the pharmaceutical industry - sound training in synthetic organic chemistry - might have to be revised.

Rafferty's basic point is that the kind of wisdom needed to find hits, advance them into lead compounds and finally into drug candidates does not really benefit from having a background in pure synthetic organic chemistry: it's much more about SAR analysis and understanding pharmacological properties. As he points out, the pharmaceutical industry has of course realized and maintained that all that wisdom can be learnt on the job. But Rafferty is not sure, and part of his skepticism comes from two revealing studies that basically showed two things: first, that even experienced medicinal chemists do not agree when picking good leads, and second, that most medicinal chemists even now don't really take optimum properties into account when designing compounds. The problem with lead picking is thus not synthesis, it's an ability to parse a complex landscape of multiple properties. Multiparameter optimization is still a beast whose footprints are rarely found among the thinking of medicinal chemists.

I think in general he's right. Advances in pharmacology, toxicology, computational chemistry and other fields over the past few decades have made it possible to both calculate as well as use property-based information in early stages of drug discovery. The article focuses on lipophilicity as one parameter which really should be considered on a regular basis but which isn't a lot of time. The problem is that a lot of synthesis has turned into a machine for cranking out molecules, so drug discovery scientists end up making molecules because they can be easily made. It's a theme that I and others have written about previously: making molecules is no longer the rate determining step in drug discovery: design is the important paradigm. One of the reasons is that CROs in China and India can now often make molecules as easily as in-house synthetic chemists. In one sense what the article is saying is because these CROs can now pick up the slack, chemists can use the time to more productively think about property-based optimization.

Now while I think it's cogent to include as much property information as possible in early drug discovery, it's worth noting that some of this information is dubious and some is valuable; the problem is that often it's hard to say which information would be dubious and which would be valuable. One of the reasons medicinal chemists disagree on compound selection is because gut instincts and experience can sometimes overrule what may seem like cogent limits on properties like lipophilicity. Nonetheless, having medicinal chemists who are tuned by default to thinking about properties would be a good idea. 

The second caveat I would apply to approaches like this is to not discount the value of a classical synthetic organic chemistry education. As has been amply demonstrated, making a complex molecule over a long period of time is more about handling setbacks, persisting with grit and developing the kind of character that can handle repeated failures than about making the molecule per se. And god knows we need all these qualities in drug discovery, a field which is literally a glutton for attrition and failure. In addition, even today there are molecules which often stump the best efforts of standard synthetic routes. Thus, it's always a good idea to have a core group of accomplished synthetic chemists in any program. In one sense the argument is really about degree, it's about what the size of this core should be, and the article argues that maybe it should be smaller than what has been traditionally thought.

Rafferty's main prescription is that graduate programs training chemists for drug discovery should now focus less on synthesis and more on multiparameter optimization and on other disciplines which can be used to think about properties upfront. The industry should do likewise in deemphasizing training of synthetic organic chemistry and emphasizing broader training in medicinal chemistry during recruitment. When I was in graduate school I was fortunate to study under a world-class medicinal chemist. Not only did his group teach students to think about properties at a relatively early stage, but more in line with what this article says, he also created a very good drug discovery course which gave students a solid flavor of the process and emphasized the contributions of other disciplines like pharmacology, formulation, metabolic studies and molecular modeling. Rafferty is encouraging more graduate programs to include such courses, and I definitely agree with him on this. The second prescription he has is to create more industry-academic partnerships in which industry contributes personnel, scholarships and funding to expose students to actual drug discovery and not just synthesis. A scheme like this has been in place in Europe for some time now.

Wikipedia seems to have caught up with the times when it defines medicinal chemistry as a discipline which 

"In its most common practice —focusing on small organic molecules—encompasses synthetic organic chemistry and aspects of natural products and computational chemistry in close combination with chemical biology, enzymology and structural biology, together aiming at the discovery and development of new therapeutic agents. Practically speaking, it involves chemical aspects of identification, and then systematic, thorough synthetic alteration of new chemical entities to make them suitable for therapeutic use. It includes synthetic and computational aspects of the study of existing drugs and agents in development in relation to their bioactivities (biological activities and properties), i.e., understanding their structure-activity relationships (SAR)."

Perhaps academia and industry can embrace this definition more fully.

Image: Amriglobal

The same and not the same: Carboxylic acids and tetrazoles

Here's a useful comparative study from Donna Huryn, Carlo Ballatore and others at Penn which looks at a problem encountered by almost every medicinal chemist in their careers: finding the right isosteres for certain functional groups. I certainly have faced this challenge in my work, and this work illustrates that the solutions for addressing the challenge can be more subtle than we think.

For non-medicinal chemists, an isostere is a set of atoms that can be swapped with another set of atoms in a molecule because it possesses physicochemical properties that are similar to the parent. This swap can be done for a variety of reasons: most commonly to improve physical properties like cell permeability. Carboxylic acids pose special problems for permeability due to their charge (since the cell membrane is hydrophobic and the charge renders them hydrophilic), so finding uncharged or other suitable charged replacements for them has been a routine consideration in drug design.

The current study looks at 35 different isosteric derivatives of phenylpropionic acids. It measures common properties like permeability (in PAMPA), logD, pKa and plasma protein binding (PPB) which can influence permeability. It also calculates the logD and pKa for comparison with the experimental values.

What the authors find is that by and large there are several acid substitutes like sulfonamides, hydroxamic acids and oxadiazoles which are superior to the parent acid in terms of permeability, and by and large this trend seems to correlate well with their logD and pKa values.

I say "by and large" above because there is one glaring exception which sticks out and which is perhaps not obvious: tetrazoles. Tetrazoles are very common replacements for acids; intuitively one would think that locking up the atoms in a ring and adding in a carbon would improve their ability to get past a membrane. But as the study finds, tetrazoles have permeabilities that are generally worse than those of the acids, sometimes by as much as 1.5 or 2 log units. And this is perhaps not too surprising if you think a bit more about their structure: While you may be locking up atoms in a ring, you are also adding two more polar atoms. In addition the tetrazole NH group is as ionizable as the COOH group of an acid. The paper says that these effects lead to a bigger desolvation penalty for the tetrazole, and that seems right; the tetrazole just seems to love water more, and a quantum mechanical study that studies the details of its solvation pattern relative to the acid might be interesting. Tetrazoles also seem to present more plasma protein binding.

There are some other interesting subtleties noted in the paper that are worth remembering. For instance the calculated logD and pKa values generally correlated well with the experimental values, except for two systems (oxadiazole-thiones and 1,3 dicarbonyls). They also noted that bigger desolvation penalties aren't always correlated with lower logD values. It's another slightly counterintuitive observation which I have seen myself in a kinase inhibitor study before.

Overall these are nice lessons to remember. Medicinal chemistry is full of useful general rules, but exceptions such as these show that the field is far from explored and will always stay interesting; there will always be enough subtleties to keep us occupied. More generally this comparison illustrates one of my favorite principles in chemistry: similarity can be both useful and deceptive.

The rise of translational research and the death of organic synthesis (again)?

The journal ACS Neuroscience has an editorial lamenting the shortage of qualified synthetic organic chemists and pharmacologists in the pharmaceutical and biotech industries. The editorial lays much of the blame at the feet of flagging support for these disciplines at the expense of the fad of 'translational research'. It makes the cogent point that historically, accomplished synthetic organic chemists and pharmacologists have been the backbone of the industry; whatever medicinal chemistry and drug design they learnt was picked up on the job. The background and rigor that these scientists brought to their work was invaluable in discovering some of the most important drugs of our time, including ones against cancer, AIDS and heart disease.
The current fascination of applied basic science, i.e., translational science, to funding agencies, due in large part to the perception of a more immediate impact on human health, is a harbinger of its own doom. Strong words? It is clear in the last 10 years that research funding for basic organic chemistry and/or molecular pharmacology is in rapid decline. However, the quality of translational science is only as strong as the basic science training and acumen of its practitioners—this truth is lost in the translational and applied science furor. A training program that instills and trains the “basics” while offering additional research in applied science can be a powerful combination; yet, funding mechanisms for the critical first steps are lacking. 
Historically, the pharmaceutical industry hired the best classically trained synthetic chemists and pharmacologists, and then medicinal chemistry/drug discovery was taught “on the job”. These highly trained and knowledge experts could tackle any problem, and it is this basic training that enabled success against HIV in the 1990s. When the next pandemic arises in the future, we will have lost the in-depth know-how to be effective. Moreover, innovation will diminish.
I have a problem pushing translational research at the expense of basic research myself. As I wrote in a piece for the Lindau Nobel Laureate meeting a few years ago, at least two problems riddle this approach:
The first problem is that history is not really on the side of translational research. Most inventions and practical applications of science and technology which we take for granted have come not from people sitting in a room trying to invent new things but as fortuitous offshoots of curiosity-driven research...For instance, as Nobel Laureates Joseph Goldstein and Michael Brown describe in a Science opinion piece, NIH scientists in the 60s focused on basic questions involving receptors and cancer cells, but this work had an immense impact on drug discovery; as just one glowing example, heart disease-preventing statins which are the world’s best-selling drugs derive directly from Goldstein and Brown’s pioneering work on cholesterol metabolism. Examples also proliferate other disciplines; the Charged-Coupled Device (CCD), lasers, microwaves, computers and the World Wide Web are all fruits of basic and not targeted research. If the history of science teaches us anything, it is that curiosity-driven basic research has paid the highest dividends in terms of practical inventions and advances. 
The second more practical but equally important problem with translational research is that it puts the cart before the horse. First come the ideas; then come the applications. There is nothing fundamentally wrong with trying to build a focused institute to discover a drug, say, for schizophrenia. But doing this when most of the basic neuropharmacology, biochemistry and genetics of schizophrenia are unknown is a great diversion of focus and funds. Before we can apply basic knowledge, let's first make sure that the knowledge exists. Efforts based on incomplete knowledge would only result in a great squandering of manpower, intellectual and financial resources. Such misapplication of resources seems to be the major problem for instance with a new center for drug discovery that the NIH plans to establish. The NIH seeks to channel the newfound data on the human genome to discover new drugs for personalized medicine. This is a laudable goal, but the problem is that we still have miles to go before we truly understand the basic implications of genomic data.

It is only recently that we have started to become aware of the "post-genomic" universe of epigenetics and signal transduction. We have barely started to scratch the surface of the myriad ways in which genomic sequences are massaged and manipulated to produce the complex set of physiological events involved in disease and health. And all this does not even consider the actual workings of proteins and small molecules in mediating key biological events, something which is underlined by genetics but which constitutes a whole new level of emergent complexity. In the absence of all this basic knowledge which is just emerging, how pertinent is it to launch a concerted effort to discover new drugs based on this vastly incomplete knowledge? It would be like trying to construct a skyscraper without fully understanding the properties of bricks and cement.
As an aside, that piece also mentions NIH's NCATS translational research center that has been the brainchild of Francis Collins. It's been five years since that center was set up, and while I know that there are some outstanding scientists working there, I wonder if someone has done a quantitative analysis of how much the work done there has, well, translated into therapeutic developments.

The editorial also has testimonials from leading organic chemists like Phil Baran, E J Corey and Stephen Buchwald who attest to the power of basic science that they discovered in their academic labs, power that they see almost disappearing from today's labs and funding agencies. This basic science which they have pioneered unexpectedly found use in industry. Buchwald's emphasis on C-N cross-coupling reactions is especially noteworthy since it was these kinds of reactions which really transformed drug synthesis and which led to Nobel Prizes for their inventors.

Baran's words are worth noting:
“It is ironic that a field with such an incredible track record for tangible contributions to the betterment of society is under continual attack. Fundamental organic synthesis has been defending its existence since I was a graduate student. If the NIH continues to disproportionally cut funding to this area, progress in the development of medicines will slow down and a vital domestic talent pool will evaporate leaving our population reliant on other countries for the invention of life saving medicines, agrochemicals, and materials.”
Baran is right that fundamental organic synthesis has been defending its existence for the last twenty years or so, but as has been discussed on this blog and in other sources, it's probably because it worked so well that it became a victim of its own success. The NIH is simply not interested in funding more total synthesis for its own sake. To some extent this is a mistake since the training that even a non-novel total synthesis imparts is valuable, but it's also hard to completely blame them. The number of truly novel reactions that have been invented in the last thirty years or so can be counted on one hand, and while chemists like Baran continue to perform incredibly creative feats in the synthesis of complex organic molecules, what they are doing is mostly applying known chemistry in highly imaginative new ways. I have no doubt that they will also invent some new chemistry in the next few years, but how much of it will compare to the fundamental explosion of new reactions and syntheses in the 1960s and 70s? I don't think this blog as well as others have denied the kind of training that synthetic organic chemistry provides, but I have certainly questioned the aura that sometimes continues to surround it (although it has declined in the last few decades) as well as the degree to which the pharmaceutical industry truly needs it.

To some extent the argument is simply about degree. The biggest challenge in most of the pharmaceutical company's postwar history was figuring out the synthesis of important drugs like penicillin, niacin and avermectin. In the era of massive screening of natural products, design wasn't really a major consideration. Contrast this period to today. The general problem of synthesis is now solved, and the major challenge facing today's drug discovery scientists is design. The big question today is not "How do I make this molecule?" but rather "How do I design this molecule within multiple constraints (potency, stability, toxicity etc.) all at the same time?" Multiparameter optimization has replaced synthesis as the holy grail of drug discovery. There are still undoubtedly tough synthetic puzzles that would benefit from creative problem-solving, but nobody thinks these puzzles won't yield to enough manpower or resources or would necessitate the discovery of fundamental new chemical principles. We of course still need top-notch synthetic organic chemists trained by top-notch academic chemists like Corey and Baran, but we equally (or even more) need chemists who are trained in solving such multiparameter design problems. Importantly, the solution to these problems is not going to come only from synthesis but also from other fields like pharmacokinetics, statistics and computer-aided design.

Another major point which I think the editorial does not touch on is the massive layoffs and outsourcing in industry which have bled it dry of deep and hard-won institutional knowledge. Drug discovery is not theoretical physics, and you cannot replenish lost talent and discover new drugs simply by staffing your organization with smart twenty-five year old wunderkinds from Berkeley or Harvard. Twenty or thirty years' experience counts for a hell of a lot in this industry; far from being a fever chill, age is a unique asset in this world. To me, this loss of institutional knowledge is a tragedy that is gargantuan compared to the lack of support for training synthetic organic chemists, and one that may have likely hobbled pharmaceutical chemistry for decades to come, if not longer.

Other than that the editorial gets it right. Too much emphasis on translational research can detract from the kind of rigorous, character-building experience that organic synthesis and classical pharmacology provide. As with many other things we need a bit of both, and some moderation seems to be in order here.

Hit picking parties, eviscerating weaknesses and other vistas from the world of molecular docking

Here’s a good review of the pitfalls and promises of molecular docking by John Irwin and Brian Shoichet (UCSF) which is worth your time, especially if you are a non-specialist who wants a summary of what’s happening in the field. As the review notes, there are many first principles-based reasons why docking should not work – poor calculation of ligand conformations, poor treatment of protein and ligand electrostatics and desolvation, non-existent consideration of protein movement, wishful treatment of water molecules, sloppy representation of x-ray structures...the list just goes on. When docking 107 or so ligands involving 1013 total configurations, any one of these “maddening details” can doom your study.

And yet as the authors note both through general considerations and a few case studies, incremental but steady improvements in docking methodology have now made the technique respectable in most structure-based drug design campaigns (which interestingly have provided more drug candidates than HTS). Several reasons have contributed to this respectability. The first is the sheer throughput; no experimental technique can possibly screen 10 million ligands in a few days or weeks, so even with its flaws docking rises up at least as a potential complement to experiment. As the review notes, even a 10% success rate in finding new binders to a good target would be an improvement, and with well-defined binding sites the success rate can surpass high throughput screening (thus, the correct question to ask is not whether your method gives you false positives and negatives but whether this error rate is enough to overwhelm the experimental discovery rate for true positives). The second factor is the existence of massive databases like ZINC and ChEMBL containing millions of annotated ligands which could serve as starting points for ligand discovery.

Thirdly, while the holy grail of docking would be to correctly predict the absolute affinities of your ligands or at least to rank them, the more modest goal is to try to separate binders from non-binders and to discover novel chemotypes. Docking has been reasonably successful in meeting this goal, and the review presents several case studies that discovered interesting ligands which were dissimilar to known ones. In turn however, what really makes the discovery of novel chemotypes interesting is that it could lead to novel biology. It is this ability to potentially “break out of medicinal chemistry boxes” that makes docking attractive. For instance you could potentially find agonists by docking when you only found antagonists before, or – in what is one of the more interesting examples illustrated in the article – you could ‘deorphanize’ an enzyme by docking potential substrates to it and predicting its reaction profile. I still find this evidence anecdotal, but it's at least a good starting point for trying out things.

The rest of the review is also useful, not in the least because - given that it's from the Shoichet lab - there's also an instructive checklist of caveats to keep in mind while experimentally screening (PAINS, aggregators etc.) ligands. There is also a discussion of using homology models for docking. Using homology models is tricky even for lead optimization, so I would be wary in applying them too widely for high-throughput docking. While the review does illustrate an interesting case involving a GPCR, I want to note that I did blog about this case when it was published and described how – new ligands notwithstanding – the calculation seemed to use enough computing power and models to light up a startup, along with copious expert input.

It’s this last point in the review which is really the crux of the matter. When the servers have cooled down and the electrons have stopped flowing, the most important equipment that one can bring to bear on a docking study is a good pair of eyes connected to an experienced brain. Even with small error rates "the scum can rise to the top" ("The scum is out there" could be a good tagline for a X-Files episode about HTS). There is little substitute for careful inspection of top docked hits and looking at things like strain, abnormal charged interactions and wrong tautomeric states; otherwise staring down a computer screen would present the same risks as staring down a gun barrel.

The Shoichet lab has occasional "hit-picking parties" where teams of medicinal and computational chemists examine docked structures, and I suspect these parties are more common in other places than you think (although probably not as common as they should be). It’s only when the high throughput-low accuracy domain of docking meets the low throughput-high accuracy domain of the human mind that docking will continue to be successful. Given what we have seen so far I think there are grounds for hope.

The death of new reactions in medicinal chemistry?

JFK to medicinal chemists: Get out of your comfort zone and
try out new reactions; not because it's easy, but because it's hard
Since I was discussing the "death of medicinal chemistry" the other day (what's the use of having your own blog if you cannot enjoy some dramatic license every now and then), here's a very interesting and comprehensive analysis in J. Med. Chem. which has a direct impact on that discussion. The authors Dean Brown and Jonas Boström who are from Astra Zeneca have done a study of the most common reactions used by medicinal chemists using a representative set of papers published in the Journal of Medicinal Chemistry in the years 1984 and 2014. Their depressing conclusion is that about 20 reactions populate the toolkit of medicinal chemists in both years. In other words, if you can run those 20 chemical reactions well, then you could be as competent a medicinal chemist in the year 2015 as in 1984, at least on a synthetic level.

In fact the picture is probably more depressing than that. The main difference between the medicinal chemistry toolkit in 1984 vs 2014 is the presence of the Suzuki-Miyaura cross-coupling reaction and amide bond formation reactions; these seem to exist overwhelmingly in modern medicinal chemistry. The authors also look at overall reactions vs "production reactions", that is, the final steps which generate the product of interest in a drug discovery project. Again, most of the production reactions are still dominated by the Suzuki reaction and the Buchwald-Hartwig reaction. Reactions like phenol alkylation which is used more frequently in 2014 and not in 1984 partly point to the fact that we are now more attuned to unfavorable metabolic reactions like glucuronidation which necessitate the capping of free phenolic hydroxyl groups.

There is a lot of material to chew upon in this analysis and it deserves a close look. Not surprisingly, there is a horde of important and interesting factors like reagent and raw material availability, ease of synthesis (especially outsourcing) and better (or flawed and exaggerated) understanding of druglike character that has dictated the relatively small differences in reaction use in the last thirty years. In addition there is also a thought-provoking analysis of differences in reactions used for making natural products vs druglike compounds. Surprisingly, the authors find that reactions like cross-coupling which heavily populate the synthesis of druglike compounds are not as frequently encountered in natural product synthesis; among the top 20 reactions used in medicinal chemistry, few are used in natural product synthesis.

There is a thicket of numbers and frequency analysis of changes reaction type and functional group type showcased in the paper. But none of that should blind us to the central take home message here: in terms of innovation, at least as measured by new reaction development and use, medicinal chemistry has been rather stagnant in the last twenty years. Why would this be so? Several factors come to mind and some of them are discussed in the paper, and most of them don't speak well of the synthetic aspects of the drug discovery enterprise. 

As the authors point out, cross-coupling reactions are easy to set up and run and there is a wide variety of catalytic reagents that allows for robust reaction conditions and substrate variability. Not surprisingly, these reactions are also disproportionately easy to outsource. This means that they can produce a lot of molecules fast, but as commonsense indicates and the paper confirms, more is not better. In my last post I talked about the fact that one reason wages have been stagnant in medicinal chemistry is precisely because so much of medicinal chemistry synthesis has become cheap and easy, and coupling chemistry is a good reason why this is so.

One factor that the paper does not explicitly talk about but which I think is relevant is the focus on certain target classes which has dictated the choice of specific reactions over the last two decades or so. For example, a comprehensive and thought-provoking analysis by Murcko and Walters from 2012 speculated that a big emphasis on kinase inhibitors in the last fifteen years or so has led to a proliferation of coupling reactions, since biaryls are quite common among kinase inhibitor scaffolds. The current paper validates this speculation and in fact demonstrates that para-disubstituted biphenyls are among the most common of all modern medicinal chemistry compounds.

Another damning critique that the paper points to in its discussion of the limited toolkit of medicinal chemistry reactions is our obsession with druglike character and with this rule and that metric for defining such character; a community pastime which we have been collectively preoccupied with roughly since 1997 (when Lipinski published his paper). The fact of the matter is that the 20 reactions which medicinal chemists hold so dear are quite amenable to producing their favorite definition of druglike molecules; flat, relatively characterless, high-throughput synthesis-friendly and cheap. Once you narrowly define what your target or compound space is, then you also limit the number of ways to access that space.

That problem becomes clear when the authors compare their medicinal chemistry space to natural product space, both in terms of the reactions used and the final products. It's well known that natural products have more sp3 characters and chiral centers, and reactions like Suzuki coupling are not going to make too many of those. In addition, the authors perform a computational analysis of 3D shapes on their typical medicinal chemistry dataset. This analysis can have a subjective component to it, but what's clear not just from this calculation but from other previous ones is that what we call druglike molecules occupy a very different shape distribution from more complex natural products. 

For instance, a paper also from AZ that just came out demonstrated that many compounds occupying "non-Lipinski" space have sphere and dislike shapes that are not seen in more linear compounds. In that context, the para bisubstituted biphenyls which dot the landscape of modern druglike molecules are the epitome of linear compounds. As the authors show us, there is thus a direct correlation between the kinds of reactions used commonly at the bench today and the shapes and character of compounds which they result in. And all this constrained thinking is producing a very decided lack of diversity in the kinds of compounds that we are shuttling into clinical trials. The focus here in particular may be on synthetic reactions but it's affecting all of us and is at least a part of the answer to why medicinal chemists don't seem to see better days.

Taken together, the analyses in this review throw the gauntlet at the modern medicinal chemist and ask a provocative question: "Why are you taking the easy way out and making compounds that are easy to make? Why aren't you trying to expand the scope of novel reactions and trying to explore uncharted chemical space"? To which we may also add, "Why are you letting your constrained views of druglike space and metrics dictate the kind of reactions you use and the molecules they result in"? 

As they say however, it's always better to light a candle than to just curse the darkness (which can be quite valuable in itself). The authors showcase several new and interesting reactions - ring-closing cross-metathesis, C-H arylation, fluorination, photoredox catalysis - which can produce a wide variety of interesting and novel compounds that challenge traditional druglike space and promise to interrogate novel classes of targets. Expanding the scope of these reaction is not easy and will almost certainly result in some head scratchers, but that may be the only way we innovate. I might also add that the advent of new technology such as DNA encoded library technology also promises to change the fundamental character of our compounds. 

This paper is clearly a challenge to medicinal chemists and in fact is pointing out an embarrassing truth for our entire community: cost, convenience, job instability, poor management and plain malaise have made us take the easy way out and keep on circling back to a limited palette of chemical reactions that ultimately impact every aspect of the drug discovery enterprise. Some of these factors are unfortunate and understandable, but others are less so, especially if they're negatively affecting our ability to hit new targets, to explore novel chemical space, and ultimately to discover new drugs for important diseases which kill people. What the paper is saying is that we can do better.

More than fifty years ago John F. Kennedy made a plea to the country to work hard on getting a man to the moon and bring him back, "not because it's easy, but because it's hard". I would say analyses like this one ask the same of medicinal chemists and drug discovery scientists in general. It's a plea we should try to take to heart - perhaps the rest of JFK's exhortation will motivate us:

"We choose to go to the moon. We choose to go to the moon in this decade and do the other things, not because they are easy, but because they are hard, because that goal will serve to organize and measure the best of our energies and skills, because that challenge is one that we are willing to accept, one we are unwilling to postpone, and one which we intend to win."

The death of medicinal chemistry?

E. J. Corey's rational methods of chemical synthesis
revolutionized organic chemistry, but they also may have been
responsible for setting off unintended explosions in the medicinal
chemistry job market
Chemjobber points us to a discussion hosted by Michael Gilman, CEO of Padlock Therapeutics, on Reditt in which he laments the fact that medicinal chemistry has now become so commoditized that it's going to be unlikely for wages to rise in that field. Here's what he had to say.
"I would add that, unfortunately, medicinal chemistry is increasingly regarded as a commodity in the life sciences field. And, worse, it's subject to substantial price competition from CROs in Asia. That -- and the ongoing hemorrhaging of jobs from large pharma companies -- is making jobs for bench-level chemists a bit more scarce. I worry, though, because it's the bench-level chemists who grow up and gather the experience to become effective managers of out-sourced chemistry, and I'm concerned that we may be losing that next general of great drug discovery chemists."
I think he's absolutely right and that's partly what has been responsible for the woes of the pharmaceutical and biotech industry over the last few years. But as I noted in the comments section of CJ's blog, the historian of science in me thinks that this is ironically the validation of the field of organic synthesis as a highly developed field whose methods and ideas have now become so standardized that you need very few specialized practitioners to put them into practice. 

I have written about this historical aspect of the field before. The point is that synthesis was undeveloped, virgin territory when scientists like R B Woodward, E J Corey and Carl Djerassi worked in it in the 1950s and 60s. They were spectacularly successful. For instance, when Woodward synthesized complex substances like strychnine (strychnine!) and reserpine, many chemists around the world did not believe that we could actually make molecules as complicated as these. Forget about standardization, even creative chemists found it quite hard to make molecules like nucleic acids and peptides which we take for granted now.

It was a combination of hard work by brilliant individuals like Woodward combined with the amazing proliferation of techniques for structure determination and purification (NMR, crystallography, HPLC etc.) that led to the vast majority of molecules falling under the purview of chemists who were distinctly non-Woodwardian in their abilities and creative reach. Corey especially turned the field into a more or less precisely predictive science that could succumb to rational analysis. In the 1990s and 2000s, with the advent of palladium-catalyzed coupling chemistry, more sophisticated instrumentation and combinatorial chemistry, even callow chemists could make molecules which would have taken their highly capable peers months or years to make in the 60s. As just an example, today in small biotech companies, interns can learn to make in three months the same molecules that bench chemists with PhDs are making. The bench PhDs presumably have better powers of critical thinking and planning, but the gap has still significantly narrowed. The situation may reach a fever pitch with the development of automated methods of synthesis. The bottom line is that synthesis is not longer the stumbling block for the discovery of new drugs; it's largely an understanding of biology and toxicity.

Because organic synthesis and much of medicinal chemistry have now become victims of their own success, tame creatures which can be harnessed into workable products even by modestly trained chemists in India or China, the more traditional scenario as pointed out by Dr. Gilman now involves a few creative and talented medicinal chemists at the top directing the work of a large number of less talented chemists around the world (that's certainly the case at my workplace). From an economic standpoint it makes sense that only these few people at the top command the highest wages and those under them make a more modest living; the average wage has thus been lowered. That's great news for the average bench chemist in Bangalore but not for the ambitious medicinal chemist in Boston. And as Dr. Gilman says, considering the layoffs in pharma and biotech it's also not great news for the field in general.

It's interesting to contemplate how this situation mirrors the situation in computer science, especially concerning the development of the customized code that powers our laptops and workstations; it's precisely why companies like Microsoft and Google can outsource so much of their software development to other countries. Coding has become quite standardized, and while there will always be a small niche demand for novel code, this will be limited to a small fraction at the top who can then shower the rest of the hoi polloi with the fruits of their labors. The vast masses who do coding meanwhile will never make the kind of money which the skill set commanded fifteen years ago. Ditto for med chem. Whenever a discipline becomes too mature it sadly becomes a victim of its own success. That's why it's best to enter a field when the supply is still tight and the low hanging fruit is still ripe for the taking. In the tech sector data science is such a field right now, but you can bet that even the hallowed position of data scientist is not going to stay golden for too long once that skill set too becomes largely automated and standardized.

What, then, will happen to the discipline of medicinal chemistry? The simple truth is that when it comes to cushy positions that pay extremely well, we'll still need medicinal chemists, but only a few. In addition, medicinal chemists will have to shift their focus from synthesis to a much more holistic approach; thus medicinal chemistry, at least as traditionally conceived with a focus on synthesis and rapid access of chemical analogs, will be seeing its demise soon. Most medicinal chemists are still reluctant to think of themselves as anything other than synthetic chemists, but this situation will have to change. Ironically Wikipedia seems to be ahead of the times here since its entry on medicinal chemistry seems to encompass pharmacology, toxicology, structural and chemical biology and computer-aided drug design. It would be a good blueprint for the future.

In particular, medicinal chemistry in its most common practice —focusing on small organic molecules—encompasses synthetic organic chemistry and aspects of natural products and computational chemistry in close combination with chemical biology, enzymology and structural biology, together aiming at the discovery and development of new therapeutic agents. Practically speaking, it involves chemical aspects of identification, and then systematic, thorough synthetic alteration of new chemical entities to make them suitable for therapeutic use. It includes synthetic and computational aspects of the study of existing drugs and agents in development in relation to their bioactivities (biological activities and properties), i.e., understanding their structure-activity relationships (SAR). Pharmaceutical chemistry is focused on quality aspects of medicines and aims to assure fitness for purpose of medicinal products.

To escape the tyranny of the success of synthetic chemistry, the accomplished medicinal chemist of the future will thus likely be someone whose talents are not just limited to synthesis but whose skill set more broadly encompasses molecular design and properties. While synthesis has become standardized, many other disciplines in drug discovery like computer-aided drug design, pharmacology, assay development and toxicology have not. There is still plenty of scope for original breakthroughs and standardization in these unruly areas, and there's even more scope for traditional medicinal chemists to break off chunks of those fields and weave them into the fabric of their own philosophy in novel ways, perhaps by working with these other practioners to incorporate "higher-level" properties like metabolic stability, permeability and clearance into their own early designs. This takes me back to a post I wrote on an article by George Whitesides which argued that chemists should move "beyond the molecule" and toward uses and properties: Whitesides could have been talking about contemporary medicinal chemistry here.

The integration of downstream drug discovery disciplines into the early stages of synthesis and hit and lead discovery will itself be a novel kind of science and art whose details need to be worked out; that art by itself holds promising dividends for adventurous explorers. But the mandate for the 20th century medicinal chemist in the 21st still rings true. Medicinal chemists who can borrow from myriad other disciplines and use that knowledge in their synthetic schemes, thus broadening their expertise beyond the tranquil waters of pure synthesis into the roiling seas of biological complexity will be in far more demand both professionally and financially. Following Darwin, the adage they should adopt is to be the ones who are not the strongest or the quickest synthetically but the ones who are most adaptable and responsive to change. 

For medicinal chemistry to thrive, its very definition will have to change.

Agonists and antagonists, and why drug discovery is hard (again)

Here's a valuable and comprehensive review on one of the most glaring pieces of evidence for why drug discovery is so hard - the fact that very small structural changes in molecules can lead to drastic changes in their biological activity.

I particularly like this review because it's absolutely chock-full of examples of small structural changes which not only impact the magnitude of binding of a small molecule to a receptor protein but invert it - that is, change an agonist into an antagonist. And the receptor family in this case is GPCRs, so it's not like we're talking about a minor rash of examples in a scientifically insignificant and financially paltry domain.

Here's one of my favorite examples from the dozens showcased in the piece; in this case a set of small molecules targeting the nociceptin receptor which is being studied as a potential target in treating heart failure and depression.



At first sight it's compelling how such similar groups as a cyclooctyl, a cyclooctyl-methyl and a phenyl can lead to complete inversion of activity, from 200 nM agonism to 1.5 nM antagonism. Thinking in 3D however makes the observation a bit more comprehensible. The N-cyclooctyl on the right is going to have a very well-defined conformational preference - pointing pretty much straight in one direction. The cyclooctyl-methyl on the other hand is going to have much more conformational freedom. It's also going to occupy much more space than the phenyl group on the right.

Now this kind of retrospective analysis may well be the explanation, but very few medicinal chemists would have been able to predictive this complete inversion in activity at the outset (as a medicinal chemist recently quipped at a Gordon Conference, "We medicinal chemists are very good at predicting the past.")

Here's a more diabolical example that would have been even harder to predict. In this case the target concerns two suptypes of the mGlu (metabotropic glutamate) receptor which is involved among other things in Parkinson's and anxiety.



In this case, not only does that 'magic' methyl group and its precise stereochemistry change an antagonist into an agonist but it even changes the agonism/antagonism mix at two separate receptors. Try explaining that, even in retrospect.

These kinds of well-known activity cliffs reinforce the essentially non-linear nature of medicinal chemistry, a quality that is essentially emergent since it arises from the interaction of small molecules with a highly non-linear biological system. Neither experimental chemistry nor computational modeling would allow us to predict activity cliffs like these because of the lack of sensitivity in such techniques.

It's things like these which I always think really need to be communicated to laymen to impress the staggering difficulty of drug design to them - most of the times we are simply ignorant when it comes to designing molecules like the ones above with any kind of predictive power and we can only find out about their fickle properties in retrospect. Perhaps then we will get less heat from the public for why we sometimes have to spend (and charge) so much money for our products.

Prof. Erland Stevens's edX med chem class

Just as he did previously, Prof. Erland Stevens of Davidson College is teaching a comprehensive med chem edX class that would be useful for anyone wanting to dive into the field. The course attracted 25,000 students last year. Here's the syllabus and list of topics - it certainly looks like something I would enthusiastically check out if I were starting out in the field, or even if I had been in it for a while.

Information on the course:

  • The course: D001x Medicinal Chemistry
  • Host platform: edX
  • Date: Starts 10/5/15, but enrollment is open until 12/11/15
  • Length: 8 weeks
  • Cost: free
  • Time: 6-8 h/wk for all content, 1 h/wk to peruse the video lectures
  • Pre-req: chem (organic functional groups, line-angle structures), biology (parts of cell), math (logarithms & exponents)
  • Topics (~1 wk per topic)
  • Drug Approval (early drugs, regulatory process, cost, IP concerns)
  • Enzymes & Receptors (inhibition, Ki, ligand types, Kd)
  • Pharmacokinetics (Vd, clearance, compartment models)
  • Metabolism (phase I & II reactions, CYP450 isoforms, prodrugs)
  • Molecular Diversity (binding, drug space, combi chem, libraries)
  • Lead Discovery (screening, filtering hits, drug metrics)
  • Lead Optimization (functional group replacements, isosteres, peptidomimetics)
  • Case Studies on Selected Drug Classes
  • Bonus features
  • Interviews with pharma professionals, including scientists from Novartis (a partner on the course)
  • Virtual labs involving online tools for predicting drug-relevant activity
  • Target audience: pre-med students, graduate students, recent pharma hires, research assistants


What would be a "non-intuitive" prediction in medicinal chemistry?

Over the last two decades when computer-aided drug design was in development, one of the most common refrains you heard from medicinal chemists about its utility was that it did a poor job predicting  “non-intuitive” structural modifications to molecules. But the term is not always easy to define, and while the charge is often valid, it’s also sometimes unfair since what’s non-intuitive can be highly subjective and constitute a moving target. Also, the bar for non-intuitive ideas can be rather high based on the experience of particular medicinal chemists; it’s not always fair to hold a simple computational method up to the same standards as a chemist with thirty years experience (that’s not exactly what the software has been designed for…).

Nonetheless, it’s rather refreshing to hear modelers level the same charge against themselves, which is perhaps a sign that the entire field is now seeking higher standards than before. I was pleased to hear this sentiment noted several times in the ACS meeting in Boston which I just attended. But the question still stands: What prediction could a modeler make that would be deemed ‘non-intuitive’ or 'novel' by a fairly experienced medicinal chemist? There are at least a few cases that come to mind:

Binding and solution conformation prediction: Chemists are used to looking at 2D structures, and even a highly experienced chemist won’t be able to predict most of the time how a complicated-looking molecule will bind in a protein binding pocket. That is what docking is for, and it's one of the few areas of modeling which can claim a modest but solid degree of success. What is still a non-trivial problem is to predict the ensemble of conformations in solution which converge to a single conformation in the protein pocket. I worked on this problem myself in grad school, and it took up the majority of my half-decade or so spent there. The problem is in both determining the population of solution conformations and estimating the binding energy going from multiple to one conformation, and the general solution is still tedious and complicated.

The prediction of conformational changes in general is a problem that cannot be easily visualized by medicinal chemists without some kind of computational or experimental (especially NMR) support. Subtle structural additions like methyl groups or halogens can sometimes cause significant changes in solution conformational populations, which in turn may impact the binding conformation. Generally speaking it is impossible to understand these effects using intuition alone. Intramolecular hydrogen bonds which can stabilize conformations and improve membrane permeability are also hard to visualize or predict without some kind of computational analysis, especially for larger molecules like macrocycles.

Scaffold hopping: Another attractive idea which is not obvious to medicinal chemists. Scaffold hopping involves essentially locating the binding pharmacophore for a molecule and then finding (ideally) a completely different set of bonds and connectivities that would map on to the same pharmacophore. It is especially useful for transforming ring systems to one another or constraining an acyclic system in a ring. One utility of scaffold hopping is to locate bioisosteres. Computational techniques can be very useful here in principle, although pharmacophore detection can be spotty because of problems with false positives and negatives. Scaffold hopping is not just non-intuitive but is also a boon to getting around intellectual property which is usually every chemist’s nemesis.

Calculating desolvation penalties: An experienced medicinal chemist may be able to look at a compound and make a guess about its size or lipophilicity, but guessing desolvation penalties is intuitively quite hard except in obvious cases (as in the case of a positively or negatively charged group – even then, guessing the sum of all the interactions is challenging). One of the reasons is that desolvation being a point charge-dipole interaction, its energy goes up as the square of the charge (instead of just inversely as in Coulomb's law): small changes in heteroatom distributions can thus have significant and non-obvious effects on solvation/desolvation. 

Unfortunately calculating solvation energies is also still hard in a general sense for computational chemists, but progress continues to be made. One of the most successful predictions of modeling would be a case where a highly charged group compensates for all its desolvation by making perfectly formed hydrogen bonds with a protein - this is a very hard thing to predict as of now. Generally speaking though, desolvation penalties would, at least in principle, fall into the category of things that medicinal chemists wouldn’t often be able to guess.

Calculating strain energies: This is another case where even experienced medicinal chemists may not be able to make intuitive statements. Sometimes it’s obvious in an x-ray crystal structure that ligands are strained (manifested for instance in the form of bent amides, bent phenyl rings or any kind of non-planar conjugated systems). But other times the effects of strain can be invisible to the naked eye. The problem is that bond length changes of as little as tenths of an angstrom can translate into significant strain energies of several kcals/mol, and it is hard if not impossible for even seasoned medicinal chemists to actually see these strain-inducing elements without some kind of calculation. That’s where modeling can help.

Water molecules: We know well by now how complicated the behavior of water molecules in protein binding sites can be. Sometimes the kinds of predictions that modelers make about easily and productively displaced water molecules are rather obvious, such as when they are talking about a water molecule in a nice hydrophobic cavity. The subtle cases are harder to intuitively predict. For instance crystallographic waters may be firmly bound and therefore may have good enthalpy but may still be unhappy and displaceable because of an unfavorable entropy. Similarly as detailed in the link above, water molecules at ligand-solvent interfaces may have unexpected thermodynamic features. Unhappy water prediction methods like WaterMap and SZMAP are promising, but only when they can predict non-intuitive scenarios that are refractory to easy analysis by medicinal chemists.

Data analysis: Generally speaking the word ‘non-intuitive’ may also mask more mundane but useful goals like being able to analyze large amounts of data and suggest useful trends. In fact that’s a task that’s usually quite unsuited to the skills of a medicinal chemist because of its reliance on numbers and statistics and opacity to easy structural visualization.

Feel free to note others in the comments section which I might have left out. Even better are cases where modelers can make suggestions that aren't just non-intuitive but counterintuitive. For instance, if you can predict that a methyl group filling a pocket would lead to a drop in potency (steric reasons? trapped water?) that would be a good counterintuitive prediction. Or if you could predict that cyclization of a molecule would actually increase conformational flexibility because of alleviation of syn-pentane interactions (as I found out in my comparison of cyclic dictyostatin with acyclic discodermolide), that prediction would also fall into the same category. Counterintuitive predictions also provide acid tests of any model because of their emphasis on falsifiability.

I don’t claim that all the goals listed above are well within the purview of molecular modeling. What I am claiming is that there are several challenging tasks on which modeling has started to make inroads. And a good number of these could be called “non-intuitive” or even "counterintuitive". A decade or two down the line I don't think such predictions from modeling will be as rare as we currently think they are, and that's something that we all should look forward to.