Field of Science

Showing posts with label genomics. Show all posts
Showing posts with label genomics. Show all posts

Leroy Hood and the tool-driven revolution in biology

The Galisonian view of science - named after historian of science Peter Galison - says that science is driven as much or even more by new techniques and instruments as by new ideas. Sadly most people have always placed theoretical ideas at the forefront of scientific revolutions, a view enforced by Thomas Kuhn's famous book "The Structure of Scientific Revolutions". But a study of the history of science shows that new tools have been as instrumental in opening up whole new areas of science as new ideas. In fact one may argue that ideas allow you to largely explain while novel tools allow you to largely discover new things.
From the viewpoint of tool-based science, scientists like Faraday, Rutherford, Woodward, and Lamb are as important as Newton, Dirac, Heisenberg and Pauling. To this list of tool-builders and users must be added the name of Leroy Hood. Hood is one of the most important pioneers of the genomics revolution. Seeing far ahead of most biologists in the 1980s when he was at Caltech, he invented four tools that were to revolutionize the theory and practice of genomics: the protein sequencer, the protein synthesizer, the DNA synthesizer and the DNA sequencer. At a time when most biologists positively looked down upon technology development and engineers, Hood blazed new paths in combining chemistry, instrumentation and biology. His tools not only allowed biologists to do things better, but allowed them to discover new things which they hadn't imagined before.
Luke Timmerman has written a valuable biography of Hood which would be of interest to anyone interested in the recent history of the gene. I picked it up encouraged by Keith's favorable review (http://omicsomics.blogspot.com/…/veteran-biotech-reporter-l…) and am glad I did. My only reservation is that Timmerman could have done a much better job embedding Hood's inventions in the bigger story of genetics and molecular biology. There were parts of the book where I thought the science could have been fleshed out much more, so if you are looking for a concomitant work of popular science along with a biography, this is not really it.
Hood's essential qualities were ingrained during a vigorous upbringing in rural Montana. His father was a peripatetic telephone engineer who did not give praise easily. He and Hood's mother taught their children to be self-reliant, resilient and hard-working. Throughout his career Hood has been a force of nature, displaying these qualities to an unprecedented extent and leaving behind some of his more talented competitors by sheer tenacity and dedication. As he recounts, the most valuable class for him in high school was not math or science but debating. He was also his high school's star quarterback. Even now, at the age of 75, he runs 3 miles every day and does a hundred push ups. He has also combined great scientific talent with a passion for public speaking and entrepreneurship; through these skills he has raised hundreds of millions of dollars from universities, funding agencies and wealthy philanthropists and made millions of his own. He has given generously to the cause of middle and high school education. No obstacle has been daunting for him, and by any of the usual metrics his career has been stunningly successful; as his website points out, "in addition to his ground-breaking research, Hood has published 750 papers, received 36 patents, 17 honorary degrees and more than 100 awards and honors, and has founded or co-founded 15 biotechnology companies including Amgen and Applied Biosystems."
Hood got his undergraduate and graduate degrees from Caltech along with an MD from Johns Hopkins. Caltech sought him out as an assistant professor right after graduation. Hood's early contributions were to immunology where he figured out the basis of antibody diversity. But soon he began to broaden his horizons and became one of the first biologists to truly appreciate the impact of new technology on biology. He had an amazing talent to spot big picture problems, drive himself mercilessly to crack them and recruit world class people to solve them. Using his unique skill set he built the first protein sequencer and DNA sequencer and licensed them out to the company Applied Biosystems. The DNA sequencer is at the very heart of the genomics revolution. Gene sequencing is no longer just a tool for faster and more efficient molecular biology, but it has transformed itself into a formidable instrument to explore stunning new domains of biology, from the creation of new organisms to the cracking of the genetic code for all kinds of diseases to the exploration of the world's biodiversity. Hood's work showed that not only can technology enable science but it can actually give rise to new science.
Unfortunately Hood's grand visions and the size of his lab and research projects (at one point his lab numbered more than a hundred people) soon ran afoul of Caltech's desire to stay a small, tightly knit school. Very soon he had a falling out with the faculty. One of his students who is now the head of research at Merck was then a professor at the University of Washington. He persuaded the medical school at UW to invite Hood for a few lectures. The chairman of the department in turn persuaded Bill Gates to attend those lectures. Gates who had started taking an interest in biology in the late 90s was entranced by Hood and immediately agreed to endow a $12 million dollar faculty position at UW for Hood. Hood's moved to UW was accompanied by breathless press releases proclaiming that his appointment was one of the most momentous events in the history of the university.
At UW Hood became the father of a new science: systems biology. He was no longer content to just explore genes and whole organisms, instead he wanted to bring about a completely unified view of biology by connecting atoms to molecules to cells, all the way to whole organisms and ecosystems. It was a grand vision, and one which only someone like Hood could pull off. Systems biology is now a mainstay of cutting edge biological science, bringing together biologists, mathematicians, computer scientists and other. But Hood got there first, being one of the first scientists to bring together interdisciplinary subject experts.
Sadly it was here that Hood's failings become clear, and Timmerman pulls no punches in narrating them. Hood was a big picture thinker, not a detail-oriented person. He left the day to day running of his labs to postdocs and research associates. More importantly, he was terrible at interpersonal relationships. He almost never took interest in his students' lives, never picked up the check when he "took them out" for lunch and regularly played favorites. He was not an unkind person, but he was simply too busy, driven to succeed and tone deaf to the everyday human relationships that make any endeavor successful. He was not above claiming credit for others' discoveries, not intentionally but because of his relentless drive to finish that simply left him clueless about such things. He rubbed people the wrong way at Caltech and UW and found even the generous support at UW insufficient for his systems biology vision. Predictably enough, when some of his key allies passed away, he had a falling out at UW too after he tried to sell them a plan for an independent new institute. Confident that his friend Bill Gates would fund it, he went to see Gates at Microsoft, only to be turned away with an icy dismissal (Gates: "I never fund things that I think are going to fail."). Undaunted, Hood poured $5 million of his own money into the institute. Personally too he faced a tragedy: his wife Valerie who he had married out of college succumbed to Alzheimer's disease.
Since then, the Institute for Systems Biology in Seattle has become a thriving research institute that is at the forefront of investigating both basic and applied genetics. Hood continues to be a powerhouse, crisscrossing the world giving talks about how biology is going to revolutionize human life. The system's research may or may not help discover new cures for important diseases, but what's more important is the vision and accomplishment of one man in achieving all that: Lee Hood. Hood is a fantastic example of what happens when passionate tenacity for a cause, a deep appreciation of the impact of technology on science, a passion for entrepreneurship and a relentless pursuit of the big picture come together to create an explosive mix. In the DNA sequencers that are humming softly in hundreds of thousands of industrial and academic laboratories and hospitals around the world, reading and rewriting the code of life, Lee Hood's legacy keeps humming on too.

Three reasons why nobody should be surprised by junk DNA

Science writer Carl Zimmer just wrote an article in the New York Times about the junk DNA wars and it's worth a read. Zimmer presents conflicting expert opinions about the definition of 'junk' and the exact role that junk DNA plays in various genetic processes. As he makes clear, the conversation is far from over. But personally I have always thought that the whole debate has been a bit of an unnecessary dustup (although the breathless media coverage of ENCODE didn't help). What surprises me especially is that people are surprised by junk DNA.
Even to someone like me who is not an expert, the existence of junk DNA appears perfectly normal. Let me explain. I think that junk DNA shouldn’t shock us at all if we accept the standard evolutionary picture. The standard evolutionary picture tells us that evolution is messy, incomplete and inefficient. DNA consists of many kinds of sequences. Some sequences have a bonafide biological function in that they are transcribed and then translated into proteins that have a clear physiological role. Then there are sequences which are only transcribed into RNA which doesn’t do anything. There are also sequences which are only bound by DNA-binding proteins (which was one of the definitions of “functional” the ENCODE scientists subscribed to). Finally, there are sequences which don’t do anything at all. Many of these sequences consist of pseudogenes and transposons and are defective and dysfunctional genes from viruses and other genetic flotsam, inserted into our genome through our long, imperfect and promiscuous genetic history. If we can appreciate that evolution is a flawed, piecemeal, inefficient and patchwork process, we should not be surprised to find this diversity of sequences with varying degrees of function or with no function in our genome.
The reason why most of these useless pieces have not been weeded out is simply because there was no need to. We should remember that evolution does not work toward a best possible outcome, it can only do the best with what it already has. It’s too much of a risk and too much work to get rid of all these defective and non-functional sequences if they aren’t a burden; the work of simply duplicating these sequences is much lesser than that of getting rid of them. Thus the sequences hung around in our long evolutionary history and got passed on. The fact that they may not serve any function at all would be perfectively consistent with a haphazard natural mechanism depending on chance and the tacking on of non-functionality to useful functions simply as extra baggage.
There are two other facts in my view which should make it very easy for us to accept the existence of junk DNA. Consider that the salamander genome is ten times the size of the human genome. Now this implies two possibilities; either salamanders have ten times functional DNA than we do, or that the main difference between us and salamanders is that they have much more junk DNA. Wouldn’t the complexity of salamander anatomy of physiology be vastly different if they really had so much more functional DNA? On the contrary, wouldn’t the relative simplicity of salamanders compared to humans be much more consistent with just varying degrees of junk DNA? Which explanation sounds more plausible?
The third reason for accepting the reality of junk DNA is to simply think about mutational load. Our genomes, as of other organisms, have undergone lots of mutations during evolution. What would be the consequences if 90% of our genome were really functional and had undergone mutations? How would we have survived and flourished with such a high mutation rate? On the other hand, it’s much simpler to understand our survival if we assume that most mutations that happen in our genome happen in junk DNA.
As a summary then, we should be surprised to find someone who says they are surprised by junk DNA. Even someone like me who is not an expert can think of at least three simple reasons to like junk DNA:
1. The understanding that evolution is an inherently messy and inefficient process that often produces junk. This junk may be retained if it’s not causing trouble.
2. The realization that the vast differences in genome sizes are much better explained by junk DNA than by assuming that most DNA is truly functional.
3. The understanding that mutational loads would be prohibitive had most of our DNA not been junk.
Finally as a chemist, let me say that I don’t find the binding of DNA-binding proteins to random, non-functional stretches of DNA surprising at all. That hardly makes these stretches physiologically important. If evolution is messy, chemistry is equally messy. Molecules stick to many other molecules, and not every one of these interactions has to lead to a physiological event. DNA-binding proteins that are designed to bind to specific DNA sequences would be expected to have some affinity for non-specific sequences just by chance; a negatively charged group could interact with a positively charged one, an aromatic ring could insert between DNA base pairs and a greasy side chain might nestle into a pocket by displacing water molecules. It was a pity the authors of ENCODE decided to define biological functionality partly in terms of chemical interactions which may or may not be biologically relevant.
The dustup from the ENCODE findings suggests that scientists continue to find order and purpose in an orderless and purposeless universe which can nonetheless produce structures of great beauty. They would like to find a purpose for everything in nature and are constantly looking for the signal hidden in the noise. Such a quest is consistent with our ingrained sense of pattern recognition and has often led to great discoveries. But the stochastic, contingent, haphazard meanderings of nature mean that sometimes noise is just that, noise. It’s a truth we must accept if we want to understand nature as she really is.

Adapted from a previous post on Scientific American Blogs.

The data junkie and the lamp post: Cancer, genomics and technological solutionism

The Cancer Genome Atlas
The other day I was having a nice discussion with a very knowledgable colleague about advances in genomic sequencing and how they are poised to transform the way we acquire and process genetic information in biology, biochemistry and medicine. The specific topic under consideration was the bevy of sequencing companies who are showcasing their wares at the Advances in Genome Biology and Technology 2015 conference. Many companies like Solexa and Oxford Nanopore are in an intense race to prove whose sequencing technology can become the next Illumina, and it's clear that much fame and fortune lies ahead for whoever gets there first. It's undoubtedly true that technological developments in this field are going to have enormous and uncertain ramifications for all kinds of disciplines as well as potentially for our way of life.
And yet as I mull over these issues I am reminded of an article by MIT biologist Michael Yaffe from the journal Science Signaling which warns against the quick wielding of what the philosopher of technology Yevgeny Morozov has called “technological solutionism”. Technological solutionism is the tendency to define problems primarily or purely based on whether or not a certain technology can address them. This is a concerning trend since it foreshadows a future where problems are no longer prioritized by their social or political importance but instead by how easily they would succumb under the blade of well-defined and easily available technological solutions. Morozov’s solutionism is a more sophisticated version of the adage about everything looking like a nail when you have a hammer. But it’s all too real in this age of accelerated technological development, when technology advances much faster than we can catch up with its implications. It’s a problem that only threatens to grow.
In his piece Yaffe alerts us to the pitfalls of somewhat mindlessly applying genomic sequencing to discovering the basis and cure for cancer and succumbing to such solutionism in the process. One of the great medical breakthroughs of the twentieth century was the finding that cancer is in its heart and soul a genetic disease. This finding was greatly bolstered by the discovery of specific genes (oncogenes and tumor suppressor genes) which when mutated greatly increase the probability and progress of the disease. The availability of cheap sequencing techniques in the latter half of the century gave scientists and doctors what seemed to be a revolutionary tool for getting to the root of the genetic basis of cancer. Starting with the great success of the human genome project, it became increasingly easier to sequence entire genomes of cancer patients to discover the mutations that cause the disease. Scientists have been hopeful since then that sequencing cancer cells from hundreds of patients would enable them to discover new mutations which in turn would point to new potential therapies.
But as Yaffe points out, this approach has often ended up relegating true insights into cancer to the application of one specific technology – that of genomics – to probe the complexities of the diseases. And as he says, this is exactly like the drunk looking under the lamppost, not because that’s where his keys really are but that’s where the light is. In this case the real basis for cancer therapy constitutes the keys, sequencing is the light. During the last few years there have been several significant studies on major cancers like breast, colorectal and ovarian cancer which have sought to sequence cancer cells from hundreds of patients. This information has been incorporated into The Cancer Genome Atlas, an ambitious effort to chart and catalog all the significant mutations that every important cancer can possibly accrue.
But these efforts have largely ended up finding more of the same. The Cancer Genome Atlas is a very significant repository, but it may end up accumulating data that’s irrelevant for actually understanding or curing cancer. Yaffe acknowledges this fact and expresses thoughtful concerns about the further expenditure of funds and effort on massive cancer genome sequencing at the expense of other potentially valuable projects.
So far, the results have been pretty disappointing. Various studies on common human tumors, many under the auspices of The Cancer Genome Atlas (TCGA), have demonstrated that essentially all, or nearly all, of the mutated genes and key pathways that are altered in cancer were already known…Despite the U.S. National Institutes of Health (NIH) spending over a quarter of a billion dollars (and all of the R01 grants that are consequently not funded to pay for this) and the massive data collection efforts, so far we have learned little regarding cancer treatment that we did not already know. Now, NIH plans to spend millions of dollars to massively sequence huge numbers of mouse tumors!
It’s pretty clear that while there has been valuable data gathered from sequencing these patients, almost none of it has led to novel insights. Why, then, do the NIH and researchers continue to focus on raw, naked sequencing? Enter the data junkie and the lamppost:
I believe the answer is quite simple: We biomedical scientists are addicted to data, like alcoholics are addicted to cheap booze. As in the old joke about the drunk looking under the lamppost for his lost wallet, biomedical scientists tend to look under the sequencing lamppost where the “light is brightest”—that is, where the most data can be obtained as quickly as possible. Like data junkies, we continue to look to genome sequencing when the really clinically useful information may lie someplace else.
The term “data junkie” conjures up images of the quintessential chronically starved, slightly bug-eyed nerd hungry for data who does not quite realize the implications or the wisdom of simply churning information out from his fancy sequencing machines and computer algorithms. The analogy would have more than a shred of truth to it since it speaks to something all of us are in danger of becoming; data enthusiasts who generate information simply because they can. This would be technological solutionism writ large; turn every cancer research and therapeutics problem into a sequencing problem because that’s what we can do cheaply and easily.
Clearly this is not a feasible approach if we want to generate real insights into cancer behavior. Sequencing will undoubtedly continue to be an indispensable tool but as Yaffe points out, the real action takes place at the level of proteins, in the intricacies of the signaling pathways involving hundreds of protein hubs whose perturbation is key to a cancer cell’s survival. When drugs kill cancer cells they don’t target genes, they directly target proteins. Yaffe mentions several recent therapeutic discoveries which were found not by sequencing but by looking at the chemical reactions taking place in cancer cells and targeting their sources and products; essentially by adopting a protein-centric approach instead of a gene-centric one. Perhaps we should re-route some of those resources which we are using for sequencing into studying these signaling proteins and their interdependencies:
These therapeutic successes may have come even faster, and the drugs may be more effectively used in the future, if cancer research focuses on network-wide signaling analysis in human tumors (20), particularly when coupled with insights that the TCGA sequencing data now provide Currently, signaling measurements are hard, not particularly suited for high-throughput methods, and not yet optimized for use in clinical samples. Why not invest in developing and using technologies for these signaling directed studies?
In other words, why not ask the drunk to buy a lamp and install it in another part of town where his keys are more likely to located? It’s a cogent recommendation. But it’s important not to lose sight of the larger implications of Yaffe’s appeal to explore alternative paradigms for finding effective cures for cancers. In one sense he is directly speaking to the love affair with data and new technology that seems to be increasingly infecting the minds and hearts of the new generation. Whether it’s cancer researchers hoping that sequencing will lead to breakthroughs or political commentators hoping that Twitter and Facebook will help bring democracy in the Arab world, we are all in danger of being sucked into the torrent of technological solutionism. Of this we must be eternally vigilant.

Adapted from a previous post on Scientific American Blogs.

ENCODE, Apple Maps and function: Why definitions matter


ENCODE (Image: Discover Blogs)
Remember that news-making ENCODE study with its claims that “80% of the genome is functional”? Remember how those claims were the starting point for a public relations disaster which pronounced (for the umpteenth time) the "death of junk DNA"? Even mainstream journalists bought into this misleading claim. I wrote a post on ENCODE where I expressed surprise at why anyone would be surprised by junk DNA to begin with.

Now Dan Graur and his co-workers from the University of Houston have published a meticulous critique of the entire set of interpretations from ENCODE. Actually let me rephrase that. Dan Graur and his co-workers have published a devastating takedown of ENCODE in which they pick apart ENCODE’s claims with the tenacity and aplomb of a vulture picking apart a wildebeest carcass. Anyone who is interested in ENCODE should read this paper, and it’s thankfully free.

First let me comment a bit on the style of the paper which is slightly different from that in your garden variety sleep-inducing technical article. The title – On the Immortality of Television Sets: Function in the Human Genome According to the Evolution-Free Gospel of ENCODE – makes it clear that the authors are pulling no punches, and this impression carries over into the rest of the article. The language in the paper is peppered with targeted sarcasm, digs at Apple (the ENCODE results are compared to AppleMaps), a paean to Robert Ludlum and an appeal to an ENCODE scientist to play the protagonist in a movie named "The Encode Incongruity". And we are just getting warmed up here. The authors spare little expense in telling us what they think about ENCODE, often using colorful language. Let me just say that if half of all papers were this entertainingly written, the scientific literature would be so much more accessible to the general public.

On to the content now. The gist of the article is to pick apart the extremely liberal, misleading and scarcely useful definition of “functional” that the ENCODE group has used. The paper starts by pointing out the distinction between function that’s selected for and function that’s merely causal. The former definition is evolutionary (in terms of conferring a useful survival advantage) while the latter is not. As a useful illustration, the function of the human heart that is selected for is to pump blood while the function that’s causal is an additional weight of 300 grams and a capacity for producing thumping sounds.

The problem with the ENCODE data is that it features causal functions, not selected ones. Thus for instance, ENCODE assigns function to any DNA sequence that displays a reproducible signature like binding to a transcription factor protein. As this paper points out, this definition is just too liberal and often flawed. For instance a DNA sequence may bind to a transcription factor without inducing transcription. In fact the paper asks why the study singled out transcription as a function: “But, what about DNA polymerase and DNA replication? Why make a big fuss about 74.7% of the genome that is transcribed, and yet ignore the fact that 100% of the genome takes part in a strikingly “reproducible biochemical signature” – it replicates!”

Indeed, one of the major problems with the ENCODE study seems to be its emphasis on transcription as a central determinant of “function”. This is problematic, since as the authors note, there's lots of sequences that are transcribed which are known to have no function. But before we move on to this, it’s worth highlighting what the authors call “The Encode Incongruity” in homage to Robert Ludlum. The Encode Incongruity points to an important assumption in the study; the implication that a biological function can be maintained without selection and that the sequences with “causal function” identified by ENCODE will not accumulate deleterious mutations. This assumption is unjustified.

The paper then revisits the five central criteria used by ENCODE to define “function” and carefully takes them apart:

1. “Function” as transcription.
This is perhaps the biggest bee in the bonnet. First of all, it seems that ENCODE used pluripotent stem cells and cancer cells for its core studies. The problem with these cells is that they display a much higher level of transcription than other cells, so any deduction of function from transcription in these cells would be exaggerated to begin with. But more importantly as the article explains, we already know that there are three classes of sequences that are transcribed without function; introns, pseudogenes and mobile elements (“jumping genes”). Pseudogenes are an especially interesting example since they are known to be inactive copies of protein-coding genes that have been rendered dead by mutation. Over the past few years as experiments and computational algorithms have annotated more and more genes, the number of pseudogenes has gone up even as the number of protein-coding genes has gone down. We also know that pseudogenes can be transcribed and even translated in some cells, especially of the kind used in ENCODE, just as we know that they are non-functional by definition. Similar arguments apply to introns and mobile elements, and the article cites papers which demonstrate that knocking these genes out doesn't impair function. So why would any study label these three classes of sequences as functional just because they are transcribed? This seems to be a central flaw in ENCODE.

A related point made by the authors is statistical in which they say that the ENCODE project has sacrificed selectivity for sensitivity. There are some simple numerical arguments that point to the large number of false positives inherent in sacrificing selectivity for sensitivity. In fact this is a criticism that goes to the heart of the whole purpose of the ENCODE study:
“At this point, we must ask ourselves, what is the aim of ENCODE: Is it to identify every possible functional element at the expense of increasing the number of elements that are falsely identified as functional? Or is it to create a list of functional elements that is as free of false positives as possible. If the former, then sensitivity should be favored over selectivity; if the latter then selectivity should be favored over sensitivity. ENCODE chose to bias its results by excessively favoring sensitivity over specificity. In fact, they could have saved millions of dollars and many thousands of research hours by ignoring selectivity altogether, and proclaiming a priori that 100% of the genome is functional. Not one functional element would have been missed by using this procedure.”
2. “Function” as histone modification
Histones are proteins that pack DNA into chromatin. The histones then undergo certain chemical modifications called post-translational modifications that cause the DNA to unpack and be expressed. ENCODE used the presence of 12 histone modifications as evidence of “function”. This paper cites a study that found a very small proportion of possible histone modifications associated with function. Personally I think this is an evolving area of research but I too question the assumption of having a function associated with most histone modifications.

3. “Function” as proximity to regions of open chromatin
In contrast to histone-packaged DNA, open chromatin regions are not bound by histones. ENCODE found that 80% of transcription sites were within open chromatin regions. But then they seem to have committed the classic logical fallacy of inferring the opposite, that most open chromatin regions are functional transcription sites (there’s that association between transcription and function again). As the authors note, only 30% or so of open chromatin sites are even in the neighborhood of transcription sites, so associating most open chromatin sites with transcription seems to be a big leap to say the least.

4. “Function” as transcription-factor binding.
This to me is another huge assumption inherent in the ENCODE study, especially as a chemist. As I mentioned in my earlier post, there are regions of DNA that might bind transcription factors (TFs) just by chance through a few weak chemical interactions. The binding might be extremely weak and may be a quick association-dissociation event. To me it seemed that in associating any kind of transcription-factor binding with function, the ENCODE team had inferred biology from chemistry. The current analysis gives voice to my suspicions. As the authors say, transcription sites are usually very short which means that TF-binding “look-alikes” may arise in a large genome purely by chance. Any binding to these sites may be confused with real TF-binding sites. The authors also cite a study in which only 86% of TF-binding sites in a small sample of 14 sites showed experimental binding to a TF. Extrapolating to the entire genome, it could mean that a fraction of the conjectured TF-binding sites may actually bind TFs.

5. “Function” as DNA methylation.
This is another instance in which it seems to me that biology is being inferred from chemistry. DNA methylation is one of the dominant mechanisms of epigenetics. But by itself DNA methylation is only a chemical reaction. The ENCODE team built on a finding that negatively correlated gene expression with methylation in CpG (cytosine-guanine) sites.  Based on this they concluded that 96% of all CpGs in the genome are methylated, and therefore functional. But again, in the absence of explicit experimental verification, CpG methylation cannot be equated with gene expression. At the very least this indicates follow-up work which will need to confirm the relationship. Until then the hypothesis that CpG methylation implies function will have to remain a hypothesis.
So what do we make of all this? It’s clear that many of the conclusions from ENCODE have been extrapolations devoid of hard evidence. 

But the real fly in the ointment is the idea of “junk DNA” which seems to have evoked rather extreme opinions that have ranged from proclaiming junk DNA as extinct to proclaiming it as God. Both these opinions perform a great disservice to the true nature of the genome. The former reaction virtually rolls the red carpet for “designer” creationists who can now enthusiastically remind us of how each and every base pair in the genome has been lovingly designed. At the same time, asserting that junk DNA must be God is tantamount to declaring that every piece of currently designated junk DNA must forever be non-functional. While the former transgression is much worse, it’s important to amend the latter belief. To do this the authors remind us of a distinction made by Sydney Brenner between “junk DNA” and “garbage DNA”. There’s the rubbish we keep and the rubbish we discard, but some rubbish may potentially turn useful in the future. At the same time, rubbish that may be useful in the future is not rubbish that’s useful in the present. Just because some “junk DNA” may turn out to have a function in the future does not mean most junk DNA will be functional. In fact as I mentioned in my post, the presence of large swathes of non-functional DNA in our genomes is perfectly consistent with standard evolutionary arguments.

The paper ends with an interesting discussion about “small” and “big” science that may explain some of the errors in the ENCODE study. The authors point out that big science has generally been in the business of generating and delivering data in an easy-to-access format. Small science has been much more competent in then interpreting the data. This does not mean that scientists working on big science are incapable of data interpretation; what it means is that the very nature of big data (and the time and resource allocation inherent in it) may make it very difficult for these scientists to launch the kinds of targeted projects that would do the job of careful data interpretation. Perhaps, the paper suggests, ENCODE’s mistake was in trying to act as both the deliverer and the interpreter of data. In the authors’ considered opinion, ENCODE “tried to perform a kind of textual hermeneutics on the 3.5 billion base-pair genome, disregarded the rules of scientific interpretation and adopted a position of theological hermeneutics, whereby every letter in a text is assumed a priori to have a meaning”. In other words, ENCODE seems to have succumbed to an unfortunate case of ubiquitous pattern seeking from which humans often suffer.

In any case, there are valuable lessons in this whole episode. The mountains of misleading publicity it generated, even in journals like Science and Nature, were a textbook study in media hype. As the authors say:
“The ENCODE results were predicted by one of its lead authors to necessitate the rewriting of textbooks (Pennisi 2012). We agree, many textbooks dealing with marketing, mass-media hype, and public relations may well have to be rewritten.”
From a scientific viewpoint, the biggest lesson here may be to always keep fundamental evolutionary principles in mind when interpreting large amounts of noisy biological data under controlled laboratory conditions. It’s worth remembering the last line of the paper:
“Evolutionary conservation may be frustratingly silent on the nature of the functions it highlights, but progress in understanding the functional significance of DNA sequences can only be achieved by not ignoring evolutionary principles…Those involved in Big Science will do well to remember the depressingly true popular maxim: “If it is too good to be true, it is too good to be true.”
The authors compare ENCODE to AppleMaps, the direction-finding app in the iPhone that notoriously bombed when it came out. Yet AppleMaps also provides a useful metaphor. Software can evolve into a useful state. Hopefully, so will our understanding of the genome.

First published on the Scientific American Blog Network.

Thoughts on personalized medicine


This is a piece I had written up for the annual report of this year's Lindau Meeting of Nobel Laureates. The final version had to be significantly edited because of space limitations so I thought I would post the full version here.

The future of personalized medicine

In this year’s Lindau meeting, the Israeli biochemist Ciechanover expressed great hope for the future of personalised medicine, an age in which medical treatments are customized and tailored to individual patients based on their specific kind disease.

In some ways personalised medicine is already here. Over centuries of medical progress, astute doctors have fully recognized the diversity of patients who are suffering from what appears to be the same disease. Based on their rudimentary knowledge of disease processes, empirical data and experience, physicians would then prescribe different combinations of medicines for different patients. But in the absence of detailed knowledge of disease at the genetic and molecular level, this kind of approach was naturally subjective; it continued to rely on extensive personal experience and ad hoc interpretations of incompletely documented empirical data.

This approach saw a paradigm shift in the latter half of the twentieth century as our knowledge of DNA and genetics revealed to us the rich diversity and uniqueness of individual genomes. Concomitantly, our knowledge of the molecular basis of disease led us to recognize molecular determinants unique to every individual. We are already taking advantage of this knowledge and harnessing it to personalize therapy.

Take the case of the anticancer drug temozolomide for instance. Temozolomide is prescribed for patients with a particularly pernicious form of brain cancer with poor prognosis. The drug belongs to a category of compounds called alkylating agents, a common class of anticancer drugs in which a reactive chemical group is transferred onto DNA in cancer cells, rendering them incapable of efficient cell division and causing their death. The problem is that because of its key role in sustaining life processes, DNA division is tightly controlled. Any kind of modification of the kind caused by temozolomide is treated as DNA damage and- for good reason- life has evolved multiple mechanisms to reverse such damage. In this case the body produces an enzyme that strips DNA of the reactive functionality attached by the drug. Thus the body unwillingly helps cancer cells by reversing the drug’s action. The understanding of this mechanism has led doctors to personalize temozolomide treatment only for individuals who have low levels of the drug-resisting enzyme. For other patients that produce high levels of the enzyme, temozolomide will unfortunately not be effective and doctors will have to turn to other drugs.

We will undoubtedly witness the proliferation of such advances in personalizing individual treatments in the future. But what appears to be an even more promising approach is to start at the source, at the fundamental genomic sequences that dictate the phenotypical changes associated with enzymes and proteins. The deciphering of the human genome has opened up exciting and promising new avenues for mapping differences in individual genomes and harnessing these differences in drug discovery. The most important strategy has been to compare genomes of individuals for single nucleotide polymorphisms (SNPs) which are changes in single base pairs in the DNA sequence. In fact much of the genetic variation between individuals and populations arises from these single nucleotide changes. SNPs have been of enormous value in tracing genetic diseases and generally categorizing variations in our species. They are typically utilised in genome-wide association studies in which the genomes of members of a certain homogeneous population with and without a disease are compared. Knowing the differences can enable scientists to pinpoint genetic markers responsible for the disease. These genetic markers can then be linked to phenotypes like enzyme overproduction or deficiency that are more directly related to the disease. In addition SNPs are unusually stable and remain constant between generations, providing scientists with a relatively time-invariant handle to study genetic disorders. One of the most notable instances of using SNPs to determine propensity toward disease involves the so-called ApoE gene in Alzheimer’s disease. Two SNPs in this gene lead to three alleles- E2, E3 and E4. Each individual inherits one maternal and one paternal copy of the ApoE gene and there is now solid evidence that the inheritance of the E4 allele leads to a greatly increased risk of Alzheimer’s disease.

In the long run, SNP’s may provide the foundation for much of personalised medicine. This is because SNPs also often dictate individuals’ propensity toward drugs, pathogens and vaccines. Thus in an ideal scenario, one might be able to predict a patient’s response to a whole battery of drugs using knowledge of specific SNPs associated with his or her disease.

Unfortunately this ideal scenario may be much farther than imagined. For one thing, we have still only scraped the surface of all possible SNPs, and there are already an estimated three million out there. But more importantly, the difference between knowing all the SNPs and knowing their causal connections to various diseases is almost like the difference between a list of all human beings on the planet on one hand and everything about their lives on the other; their professions, origins, hobbies, political views, family lives. Knowing the former is far from understanding the latter.

In this sense the problem with SNPs illustrates the problems with all of personalised medicine. In fact it’s a problem that plagues scientific research in general, and that’s the dilemma of separating correlation from causation. The problem is even more acute in a complex biological system like a human being where the ratio of extraneous unrelated correlations to genuinely causative factors is especially high. Simply knowing the SNP variations between a healthy and diseased individual is very different from being able to pinpoint the SNP that is directly connected to the disease. The situation is made exponentially more complex by the fact that these putative determinants usually act in combination with each other. Thus one has to now account not only for the effect of an individual SNP but also for the differential effects of its combination with other SNPs. And as if this complexity were not enough, there’s also the fact that many SNPs occur in non-coding regions of the human genome, leading to even bigger questions about their exact relevance. Sophisticated computers and statistical methods are enabling us to sort through this jungle of data, but as of now the data itself clearly outnumbers our ability to intelligently analyse it. We need to become far more capable at distinguishing signal from noise if we are to translate genetic understanding into practical therapeutic strategies.

In addition, while a certain kind of SNP may be able to determine disease tendency, there are also many false positives and negatives. Only a small percentage of SNPs are typically linked to a condition, especially when it comes to complex conditions like cancer, diabetes and psychological disorders. Many SNPs may simply be surrogate SNPs that have little to do with the disease themselves but which have come along for the ride with other SNPs. It is a difficult task to say the least to separate the wheat from the chaff and hone in on the few SNPs that are truly serving as disease determinants or markers. In such cases it is instructive to borrow from the example of temozolomide and remember that ultimately we will be able to untangle cause and effect only by looking at the molecular level interaction of drugs and biomolecules. No amount of data sequencing and analysis can really be a substitute for a robust study designed to directly demonstrate the role of a particular enzyme or protein in the etiology of a disease. It’s also worth noting that such studies have always benefited from the tools of classical biochemistry and pharmacology, and thus practitioners of these arts will continue to stand on an equal footing with the new genomics experts and computational biologists in unraveling the implications of genetic differences.

Finally, there’s the all-pervasive question of nature versus nurture. Along with genomics, one of the most important advances of the last decade has been the development of epigenetics. Epigenetics refers to changes in the genome that are induced by the environment and not hard-coded in the DNA sequence. An example includes the environmentally stimulated silencing or activation of genes by certain classes of enzymes. Epigenetic factors are now known to be responsible for a variety of processes in diseases and health. Some of these factors can even operate in the fetal stage and influence physiological responses in later life. While epigenetics has revealed a fascinating new layer of biological control and has much to teach us, it also adds another layer of complexity to the determination of individual responses to therapy. We have a long way to go before we can perfect the capability to clearly distinguish genetic from epigenetic factors as signposts for individualized therapy.

The future of personalised medicine is therefore both highly exciting as well as extremely challenging. There is much promise to be had in mapping the subtle genetic differences that make us react differently to diseases and their cures, but we will also have to be exceedingly careful in not leading ourselves astray with incomplete data, absence of causation and confirmation bias. It is a tough, but ultimately rewarding problem which will lead to both fundamental understanding and new medical advances. It deserves our attention in every way.

Image source

The leap forward

In 5 to 10 years, I could walk into a doctor's office and for less than 1000$ have my genome sequenced. This could tell me how likely I am to get Alzheimer's disease, cancer or heart disease. Such possibilities open up enormous technical, legal and moral challenges.

That this would be possible in 5 to 10 years is what Craig Venter says. Coming from a lesser man this would seem like wishful speculation, but Venter is the man who single-handedly raced the government to "shotgun" sequencing the human genome and he has a knack of turning dreams into reality. If there's one word to describe Venter it's "big". Even after the human genome the man has not rested on his laurels. Over the last few years he has trawled the seven seas in his boat, gathering marine water samples everywhere and analyzing them for interesting bacteria and unusual DNA sequences. Genetically engineering these bugs could give us organisms that could mop up CO2, that would produce biofuels. Recently Venter has teamed up with Exxon (who could imagine!) for investigating genetically engineered algae that would produce hydrocarbons for transportation fuel.

Venter aims to have a unique genomic sequencing facility where he can sequence the genome of almost any organism one can think of. The computational and sequencing power in this facility is incredible and it seems to get bigger and more efficient every day. Here he shows the facility and has a conversation about it with Richard Dawkins. Venter says that mammalian genes are no longer very interesting because they don't show that much variation compared to some other genomes such as insect and bacterial genomes. Right now Venter has a room full of sequencers and behind a secret curtain he already has a large sequencer that would replace this entire room. The rate at which the technology is growing is breathtaking. This is a one of a kind tour and the entire video is worth watching. When Dawkins asks Venter if he is trying to "play God", Venter quips that he "does not play mythological characters"...



I would very strongly recommend Venter's autobiography. His journey from a careless, unfocused, fun-loving teenager who, shaped by harrowing experiences as a medic in Vietnam brought pin point focus to his career is amazing. Venter has been called reckless, audacious, arrogant, egoistic, curt and apathetic, and to be honest all these qualities shine forth on many pages of his book (however he appears pretty modest and engaging in most of his interviews including the one above). He has managed to alienate many people who he worked with (as someone quipped, he will never win a rightly deserved Nobel prize because of his personality traits). But it doesn't matter; he is a born entrepreneur and he is nothing but brilliant and ambitious, an intrepid risk taker. He is one of those people who will bulldoze his way through problems and not care what other people think. Nobody has accomplished what he has.