Schizophrenia is another one of those conditions (like alcoholism) that we're constantly being told is highly heritable, based on the results of twin studies and various other findings (from studies that pre-date the DNA-sequencing era). Studies "proving" the heritability of schizophrenia are, in general, rarely questioned, or even reviewed critically, due to the fact that a genetic basis for schizophrenia makes so much intuitive sense and agrees so well with ordinary experience. Everybody knows someone who "has schizophrenia in their family," and it very much does seem to travel in families.
The fact that something travels in families is not automatically indicative of a genetic connection, but most people aren't ready to believe such a thing since it goes against common sense. Nevetheless, in the past few decades, a great deal of work (much of it quite compelling) has gone into the study of things like intergenerational trauma, with researchers finding, for example, that the course of PTSD in modern-day Israeli military veterans is much different for soldiers whose parents or grandparents were Holocaust survivors. (I talk about intergenerational trauma in Of Two Minds. In fact, there's a whole chapter on Trauma.)
Social workers and criminologists have known for years that childhood physical and sexual abuse tend to be intergenerational. Yet no one seriously suggests there's a "child abuse gene" or a "child molestation gene." A lot of things that aren't purely genetic are trans-generational; familial. Schizophrenia itself may be in this category. The links between childhood trauma and schizophrenia are quite strong and well known. Indeed, the literature shows that the associations between childhood trauma and schizophrenia are actually as strong as or stronger than the associations between childhood trauma and other disorders that we know have a high correspondence to childhood trauma (including depression, PTSD, anxiety disorders, dissociative disorders, eating disorders, personality disorders, substance abuse, and sexual dysfunction, among others).
Still, the search for "schizophrenia genes" goes on, and every so often we hear claims for this or that group of scientists having discovered some new genetic feature that confers high risk for schizophrenia. The latest of these reports comes from the Schizophrenia Working Group of the Psychiatric Genomics Consortium, which published a piece in Nature Letters in July 2014 called "Biological insights from 108 schizophrenia-associated genetic loci." The report trumpets the finding of "108 conservatively defined loci that meet genome-wide significance, 83 of which have not been previously reported." This is a genome-wide association study, subject to the usual GWAS caveats.
Of the 108 loci, 75% were associated with protein-coding genes, a very high number for a GWAS study, and quite a few genes involve physiologically relevant functions.
On the surface, it sounds promising.
Where things start to fall down is in determining how much heritability risk the 108 loci can account for. The authors said: "Assuming a liability-threshold model, a lifetime risk of 1%, independent SNP effects, and adjusting for case-control ascertainment, RPS" (risk profile scores) "now explain about 7% of variation on the liability scale to schizophrenia across the samples, about half of which (3.4%) is explained by genome-wide significant loci." To parse that: We know the lifetime risk of schizophrenia is 1%. If we total up the disease risk attributable to all of the 128 markers found in the study, we can explain 7% of schizophrenia risk. But if we limit ourselves to the 108 markers that met criteria for genome-wide significance (a stringent test intended to reduce the risk of false positives), we can explain only a little more than 3% of schizophrenia risk.
So once again, it's a mixed bag: a genome-wide association study (GWAS) that's filled with interesting findings (some of which may prove to be important), but that fails, quite miserably, to explain heritability.
Many people have commented on the possible reasons why genome-wide association studies have been hit-or-miss in their ability to explain heritability, why they so often find SNPs (single nucleotide polymorphisms) that don't map to genes, why the results often don't replicate well, etc. I think if we're honest, we have to admit that GWAS isn't always the right tool for the job. It's a great tool, but you have to remember how it works. GWAS attempts to match fairly common (i.e., occurring in 1% or more of the population) single nucleotide polymorphisms (which you can think of as point mutations) in databases that have been specially built to accommodate the most common polymorphisms. In a GWAS, you aren't doing deep sequencing of people's genes; you're looking for markers that cover a rather tiny percent of the genome. These markers occur commonly, which almost by definition prevents them from being very serious defects, because serious genetic defects are rendered rare by evolution (natural selection removes them from the gene pool over time). Therefore, in a technique that, by design, looks mainly at commonly occurring SNPs, which are overwhelmingly neutral (from an evolutionary standpoint), we should not be surprised to find that the markers that get spotted in GWAS (as genome-wide significant) are essentially neutral mutations without much phenotypic effect.
Unfortunately, we're at an awkward point in the history of science, because technology has given us some powerful DNA analysis tools (and the computers needed to crunch the data), but we're not quite yet at the point where detailed, deep genetic sequencing of whole human genomes is economically feasible for a study that includes, say, a thousand cases and a thousand controls. (Right now, it costs about $10,000 to do the necessary sequencing.) When a human genome can be reliably sequenced for under $2000, we'll have the tools to do really serious, deep genetic studies of hard-to-pin-down diseases. Whether we'll have the computer power needed to do the necessary statistical analysis on the scale of large (N > 1000) case-control studies is another matter; remember that in GWAS we're dealing with perhaps 500,000 data points per genome whereas in deep sequencing we're generating billions of data points per genome. It'll probably be doable using Amazon AWS/EC2 infrastructure (or Google's compute-cloud). Another reason to invest in Amazon, probably.
I'm in the process of writing a 100,000-word book on mental illness. It's evidence-based; part science, part memoir, with over 200 footnotes (so you can refer to the scientific literature yourself). To follow the progress of the book, check back here often, and also consider adding your name to our mailing list. Thanks!
Showing posts with label DNA. Show all posts
Showing posts with label DNA. Show all posts
Saturday, January 10, 2015
Tuesday, January 06, 2015
The Problem of "Missing Heritability"
When the human genome was sequenced in 2003, the expectation was that scientists, armed with a powerful array of new genetic techniques, would very quickly identify the genetic correlates of things like schizophrenia, depression, autism, and alcoholism, which we supposedly "know" have a large genetic component. Extremely large, well-funded genome-wide association studies (GWAS) were begun. (See this PLoS paper to learn how such studies are conducted.) Surprisingly, the studies failed, for the most part, to converge on genes, gene copy number variants, SNPs (single nucleotide polymorphisms), gene rearrangements, mutations, or other peculiarities of allelic architecture whose presence could predict important diseases. When genetic loci of interest were found, the findings often didn't replicate in followup studies; or if they did, the relative odds ratios (a measure of the ability of genetic features to predict the prevalence of a trait) fell far short of explaining the known magnitudes of various traits in target populations.
If things like schizophrenia and alcoholism truly do have a strong genetic basis (as we've been told), and if they were to involve tractable numbers of genes (say, dozens or scores of genes, rather than hundreds or thousands), DNA studies of the GWAS kind should produce immediate, strong, recognizable genetic signatures of disease. The results should leap off the page. But they don't. Time and again, the very few candidate alleles that are found are either found in low numbers, or have low "penetrance" (low capacity to predict disease), or both, and quite often the candidate alleles that are found are not confirmed by followup studies.
This has given rise to major crisis in genetics, summed up in a paper called "Finding the missing heritability of complex diseases" that appeared in Nature in 2009. Scientists are desperate to explain the "missing heratibility" of disorders we "know" are genetic.
It's assumed, by most scientists, that the failure of genome-wide association studies to find genetic explanations of complex diseases can be attributed to such studies' low power and resolution for genetic variations of modest effect. (So the quest has begun to increase the study size of future trials, in hopes of seeing more robust results.) In addition, there's a reasonable expectation that many complex diseases will doubtless be found to involve numerous genes, each contributing only a small effect to the total. It's also been suggested that some traits are dictated by extremely rare genetic features with high "penetrance." There's also an awareness that current laboratory techniques have low power to detect gene-gene interactions. What no scientist wants to say, however—and what none of the 24 co-authors of the Nature article (above) would say—is that maybe the genetic component(s) of schizophrenia, alcoholism, depression, etc. are simply vanishingly small to begin with. Yet that's exactly what the data are telling us. But we won't listen to the data because it doesn't fit our preconception of how the world should work.
We "know" that a person's height is largely controlled by genes. Studies going back almost a century have determined that body height is 80% to 90% heritable; no one seriously questions this fact. (Height is heritable—it "runs in families.") However, at least three large, modern genetic studies have been done to find "height genes"; the largest involved over 180,000 study subjects (and 291 co-authors). In all, some 180 genetic loci were identified that play a role in determining a person's height. But the 180 genomic features, put together, accounted for only 10% of observed variations in height. The rest appears to be environment.
Are we now supposed to enlarge our "study population" from 180,000 to several million, in order to find the genetic explanation for body height, just because we "know" one exists?
At some point, don't we have to just admit "the data's the data"?
Shouldn't we be willing (at least provisionally) to entertain the idea that maybe the twin studies and the data showing that certain things "run in families" (pre-GWAS-era data, mostly) are in need of reevaluation? Shouldn't we at least consider the idea that prior studies misjudged the importance of uncontrolled-for environmental variables? Now that we have powerful DNA-analytic techniques for investigating heritability, and the techniques aren't giving us the results we want, must we "increase the size of the microscopic" to make the results look bigger? Or shouldn't we just accept what the data are telling us (as painful as that may be)?
In a future post, I'll talk about some of the GWAS results for alcoholism. Stay tuned.
This post is, in part, derived from material for a forthcoming book on mental illness I'm writing. Please come back often to find out how to get free sample chapters.
If things like schizophrenia and alcoholism truly do have a strong genetic basis (as we've been told), and if they were to involve tractable numbers of genes (say, dozens or scores of genes, rather than hundreds or thousands), DNA studies of the GWAS kind should produce immediate, strong, recognizable genetic signatures of disease. The results should leap off the page. But they don't. Time and again, the very few candidate alleles that are found are either found in low numbers, or have low "penetrance" (low capacity to predict disease), or both, and quite often the candidate alleles that are found are not confirmed by followup studies.
This has given rise to major crisis in genetics, summed up in a paper called "Finding the missing heritability of complex diseases" that appeared in Nature in 2009. Scientists are desperate to explain the "missing heratibility" of disorders we "know" are genetic.
It's assumed, by most scientists, that the failure of genome-wide association studies to find genetic explanations of complex diseases can be attributed to such studies' low power and resolution for genetic variations of modest effect. (So the quest has begun to increase the study size of future trials, in hopes of seeing more robust results.) In addition, there's a reasonable expectation that many complex diseases will doubtless be found to involve numerous genes, each contributing only a small effect to the total. It's also been suggested that some traits are dictated by extremely rare genetic features with high "penetrance." There's also an awareness that current laboratory techniques have low power to detect gene-gene interactions. What no scientist wants to say, however—and what none of the 24 co-authors of the Nature article (above) would say—is that maybe the genetic component(s) of schizophrenia, alcoholism, depression, etc. are simply vanishingly small to begin with. Yet that's exactly what the data are telling us. But we won't listen to the data because it doesn't fit our preconception of how the world should work.
We "know" that a person's height is largely controlled by genes. Studies going back almost a century have determined that body height is 80% to 90% heritable; no one seriously questions this fact. (Height is heritable—it "runs in families.") However, at least three large, modern genetic studies have been done to find "height genes"; the largest involved over 180,000 study subjects (and 291 co-authors). In all, some 180 genetic loci were identified that play a role in determining a person's height. But the 180 genomic features, put together, accounted for only 10% of observed variations in height. The rest appears to be environment.
Are we now supposed to enlarge our "study population" from 180,000 to several million, in order to find the genetic explanation for body height, just because we "know" one exists?
At some point, don't we have to just admit "the data's the data"?
Shouldn't we be willing (at least provisionally) to entertain the idea that maybe the twin studies and the data showing that certain things "run in families" (pre-GWAS-era data, mostly) are in need of reevaluation? Shouldn't we at least consider the idea that prior studies misjudged the importance of uncontrolled-for environmental variables? Now that we have powerful DNA-analytic techniques for investigating heritability, and the techniques aren't giving us the results we want, must we "increase the size of the microscopic" to make the results look bigger? Or shouldn't we just accept what the data are telling us (as painful as that may be)?
In a future post, I'll talk about some of the GWAS results for alcoholism. Stay tuned.
This post is, in part, derived from material for a forthcoming book on mental illness I'm writing. Please come back often to find out how to get free sample chapters.
Saturday, November 01, 2014
A Virus Where It Shouldn't Be
Words like "bizarre" and "unexpected" hardly suffice when describing the results published this week in PNAS (one of the most respected journals in all of science) by Robert H. Yolken and 17 coworkers from Johns Hopkins, University of Nebraska, and Baltimore's Sheppard Pratt Health System, who wrote about finding a peculiar virus in the noses of a number of individuals. The virus they found (through DNA metagenome analysis) was something called ATCV-1, a large DNA virus normally associated with the freshwater alga Chlorella.
Just to bring you up-to-date quickly: The scientists in question stumbled onto the fact that DNA from the ATCV-1 virus appears to exist in the noses and throats of a surprising fraction (43%) of randomly chosen individuals. The study population of 92 people is not large, however, and we should not be too quick to jump to conclusions. Nevertheless, the mere finding of this virus's DNA in that many people's noses is shocking, because this is a virus that, until now, was thought to occur only in algae, not higher life forms (and certainly not in humans).
To understand how shocking this is, you have to realize that viruses are generally extremely highly adapted to specific hosts. A virus that attacks tobacco plants only attacks tobacco plants. A virus that infects bacteria only infects bacteria, and generally only a specific species of bacterium. Viruses coevolve very closely with their hosts, developing extremely intricate, fine-tuned adaptations to a specific organism. For a virus to cross species lines (let alone Kingdoms) is unusual, to say the least. And thank goodness! Otherwise, you might come down with pox from an insect bite, or get any number of deadly diseases from eating ordinary foods.
As it turns out, ATCV-1 (Acanthocystis turfacea Chlorella virus) has appeared in this blog once before, back in March, in a post I did called "A Virus, a Worm, and a Louse Walk into a Bar." The point of that post (ironically) was to show that while most virus genes tend to have a great deal of DNA homology with their host counterparts, the genes of certain algae viruses actually appear to have greater homology with genes in the human body louse. Which perhaps should have been a tipoff, of sorts, to the Yolken et al. findings. Certainly, it adds more color to the story, knowing (as we do) that phycodnavirus DNA may have have crossed species lines before. ("Phycodnavirus" is the name scientists have given to the family of large DNA viruses that infect algae.)
What do we know about ATCV-1? It's fairly large, as viruses go, with a genome of 288,047 base pairs, encompasssing as many as 860 genes. The fact that the genome has been fully sequenced doesn't mean we know how it works, though. In fact, we don't know what most of the 860 genes do. Some of the genes are (as in many king-size viruses) devoted to DNA-synthesis functions of a type commonly associated with nuclear-expressed genes in the host, meaning that the virus appears well adapted to take over the nucleus of a cell. This is not always the case; some viruses are adapted to thrive in the cytoplasm, away from the nucleus.
Intriguingly, the Yolken group conducted tests of cognitive function among enrollees in its study and found that there was a statistically significant reduction in cognitive ability (on the particular measures tested) in the group of people that tested positive for ATCV-1. To determine if this was just a fluke, researchers evaluated the effects of the virus on mice. As it happens, inoculated rodents showed deficits in recognition memory and attention while navigating mazes. In other words, exposure to the virus is associated with cognitive deficits in both humans and an animal model.
This is remarkable, since (if true) it would suggest that the virus traveled through the blood (of humans and mice), crossed the blood-brain barrier, gained entry to brain cells, and expressed its DNA inside brain cells (a type of cell the virus would presumably never encounter in its normal habitat of freshwater marshes and lakes). It all sounds (and is) very unlikely. But there you are: PNAS is one of the most esteemed science journals in the world. They published the result. It's as real as can be.
Of course, the work needs to be replicated and extended. And I'm sure it will be.
And I think what we'll find is that phycodnaviruses have been crossing species boundaries for some time. If you go back to my original blog about this, you'll see an example of a particular gene, encoding a ribonucleoside reductase, in another Chlorella virus (called PBCV-1), which shows greater homology to the ribonucleoside reductase of the human body louse than to the same gene in the virus's "normal" host, Chlorella. (The homology between PBCV-1 and louse reductases is 53%, versus 48% for virus and Chlorella. These are protein-sequence homology numbers.)
If we check the ATCV-1 reductase for homology with reductases in other organisms, we find that the highest scoring non-viral match (53%) occurs not with Chlorella (the virus's host), but with Lichtheimia corymbifera, a fungus found in soil and decaying plant matter. Interestingly, this fungus is known to cause pulmonary, CNS, and rhinocerebral infection in animals and humans. (But there is no known association whatsoever between ATCV-1 and the fungus.) One wonders whether the fungus has learned a few tricks from marine viruses; or perhaps vice versa?
If you're wondering how the people in the Yolken et al. study could have come in contact with ATCV-1, maybe the answer is as simple as taking a drink of water from a mountain stream, or drinking unfiltered water from a well, or swimming in a lake, or perhaps just walking in a marsh.
Suffice it to say, a good deal more remains to be learned regarding the ecology and life cycles of the phycodnaviruses, and cross-species infection/transfection generally. I feel certain the recent results of Yolken et al. will stimulate a great amount of much-needed followup research.
In the meantime, would you grab me a bottled water?
Please share this story on your favorite social channel, if you enjoyed it. Thanks!
![]() |
| ATCV-1 virions attached to a Chlorella cell. |
To understand how shocking this is, you have to realize that viruses are generally extremely highly adapted to specific hosts. A virus that attacks tobacco plants only attacks tobacco plants. A virus that infects bacteria only infects bacteria, and generally only a specific species of bacterium. Viruses coevolve very closely with their hosts, developing extremely intricate, fine-tuned adaptations to a specific organism. For a virus to cross species lines (let alone Kingdoms) is unusual, to say the least. And thank goodness! Otherwise, you might come down with pox from an insect bite, or get any number of deadly diseases from eating ordinary foods.
As it turns out, ATCV-1 (Acanthocystis turfacea Chlorella virus) has appeared in this blog once before, back in March, in a post I did called "A Virus, a Worm, and a Louse Walk into a Bar." The point of that post (ironically) was to show that while most virus genes tend to have a great deal of DNA homology with their host counterparts, the genes of certain algae viruses actually appear to have greater homology with genes in the human body louse. Which perhaps should have been a tipoff, of sorts, to the Yolken et al. findings. Certainly, it adds more color to the story, knowing (as we do) that phycodnavirus DNA may have have crossed species lines before. ("Phycodnavirus" is the name scientists have given to the family of large DNA viruses that infect algae.)
What do we know about ATCV-1? It's fairly large, as viruses go, with a genome of 288,047 base pairs, encompasssing as many as 860 genes. The fact that the genome has been fully sequenced doesn't mean we know how it works, though. In fact, we don't know what most of the 860 genes do. Some of the genes are (as in many king-size viruses) devoted to DNA-synthesis functions of a type commonly associated with nuclear-expressed genes in the host, meaning that the virus appears well adapted to take over the nucleus of a cell. This is not always the case; some viruses are adapted to thrive in the cytoplasm, away from the nucleus.
Intriguingly, the Yolken group conducted tests of cognitive function among enrollees in its study and found that there was a statistically significant reduction in cognitive ability (on the particular measures tested) in the group of people that tested positive for ATCV-1. To determine if this was just a fluke, researchers evaluated the effects of the virus on mice. As it happens, inoculated rodents showed deficits in recognition memory and attention while navigating mazes. In other words, exposure to the virus is associated with cognitive deficits in both humans and an animal model.
This is remarkable, since (if true) it would suggest that the virus traveled through the blood (of humans and mice), crossed the blood-brain barrier, gained entry to brain cells, and expressed its DNA inside brain cells (a type of cell the virus would presumably never encounter in its normal habitat of freshwater marshes and lakes). It all sounds (and is) very unlikely. But there you are: PNAS is one of the most esteemed science journals in the world. They published the result. It's as real as can be.
Of course, the work needs to be replicated and extended. And I'm sure it will be.
And I think what we'll find is that phycodnaviruses have been crossing species boundaries for some time. If you go back to my original blog about this, you'll see an example of a particular gene, encoding a ribonucleoside reductase, in another Chlorella virus (called PBCV-1), which shows greater homology to the ribonucleoside reductase of the human body louse than to the same gene in the virus's "normal" host, Chlorella. (The homology between PBCV-1 and louse reductases is 53%, versus 48% for virus and Chlorella. These are protein-sequence homology numbers.)
If we check the ATCV-1 reductase for homology with reductases in other organisms, we find that the highest scoring non-viral match (53%) occurs not with Chlorella (the virus's host), but with Lichtheimia corymbifera, a fungus found in soil and decaying plant matter. Interestingly, this fungus is known to cause pulmonary, CNS, and rhinocerebral infection in animals and humans. (But there is no known association whatsoever between ATCV-1 and the fungus.) One wonders whether the fungus has learned a few tricks from marine viruses; or perhaps vice versa?
If you're wondering how the people in the Yolken et al. study could have come in contact with ATCV-1, maybe the answer is as simple as taking a drink of water from a mountain stream, or drinking unfiltered water from a well, or swimming in a lake, or perhaps just walking in a marsh.
Suffice it to say, a good deal more remains to be learned regarding the ecology and life cycles of the phycodnaviruses, and cross-species infection/transfection generally. I feel certain the recent results of Yolken et al. will stimulate a great amount of much-needed followup research.
In the meantime, would you grab me a bottled water?
Please share this story on your favorite social channel, if you enjoyed it. Thanks!
Labels:
alga virus,
ATCV-1,
Chlorella,
DNA,
genomics,
Lichtheimia,
PBCV-1,
phycodnavirus,
reductase
Wednesday, June 25, 2014
Why So Many Helicases?
Many DNA-processing genes have an unusual amount of internal complementarity: regions of DNA in which the DNA can fold back on itself to form stable structures. A good example is the dinG gene of Mycobacterium tuberculosis, which encodes an ATP-dependent helicase. Using the DINAMelt server's Quikfold app, I obtained the following structure prediction for the first 1,000 bases of the (1,971-bases-long) dinG gene of M. tuberculosis Erdman strain.
Remarkably, almost the entire sequence can form a stable self-annealing structure. Only a few bases (see red arrow, above) lack the ability to form secondary structure. Bear in mind, what we're looking at is a stable conformation involving one strand of DNA only. (Each of the two strands of dinG can form this structure, independently of one another.) The structure shown above has a Tm (melting temperature) of 66.4°C in 1M saline, with mean-free-energy enthalpy of minus-2418.60 kcal/mol and a 37°C Gibbs free energy (ΔG) of minus-209.92 kcal, meaning that at 37°C, formation of the stable structure shown here (or one very much like it) is, energetically speaking, strongly favored.
Structures of this kind are often considered to occur in RNA, but if they also occur in single-stranded DNA, it raises interesting questions. If a gene has an energetically stable strands-apart configuration, getting the strands of duplex B-form DNA to separate might not be so hard. But more to the point, getting the self-annealing gene to come back together again as duplex DNA will require significant energy input. In molecular genetics, we're accustomed to the idea of duplex DNA requiring help from an ATP-dependent helicase to "open up" (unwind) the double helix in preparation for replication or transcription. The above diagram suggests that the problem isn't "opening up" a gene; the greater problem may be bringing the strands together again after they've assumed a stable strands-apart secondary structure. There's a substantial energy barrier to be overcome before the above structure can be relaxed into randomly coiling DNA.
This suggests that certain genes, like dinG, may be modal in terms of strand-separation state. Once the gene's strands are apart, they want to stay apart. There's an energy barrier to bringing the strands together again.
It's ironic that a helicase gene (dinG) has so much single-strand secondary structure. The gene product is a DNA-powered helicase; the gene needs its own protein product in order to zip up again. But maybe that's the point? Maybe it's a non-accidental feature.
In general, bacteria tend to have a remarkable number of helicase genes. M. tuberculosis (for example) has 16 different helicases. Other species have even more. (See table.)
One might ask why this is so; why would a bacterium need 15, 19, 25, or 42 different helicases? It's quite unusual for a bacterial genome to have significant redundancy of genes, because when there are two copies of a given gene, one copy usually eventually becomes disabled (pseudogenized) and lost through random mutations. The very few exceptions to this rule tend to involve highly transcribed, highly necessary genes (such as ribosomal-RNA genes). It would be extremely unlikely for M. tuberculosis to carry around 16 "flavors" of a gene if they weren't all absolutely necessary. Duplicates would almost certainly be lost over time, especially in M. tuberculosis, which lacks a mismatch repair system. (Bacteria that lack mismatch repair enzymes have been shown experimentally to lose DNA fifty times faster than other bacteria.) The most parsimonious view is that the 16 helicases of M. tuberculosis are, in fact, critically necessary and perform different jobs.
I would suggest that perhaps the reason bacteria have so many helicases is that these are actually the "nucleic acid chaperones" that manage secondary structure in the sizable minority of genes that exhibit pronounced self-annealing of separated strands. It could be that most helicases are tasked with separating individual strands of DNA (and/or RNA) from themselves. Different types of secondary structure require different types of helicase to unravel. This might be why bacteria need so many helicases.
References
Helpful articles on DNA secondary structure:
Remarkably, almost the entire sequence can form a stable self-annealing structure. Only a few bases (see red arrow, above) lack the ability to form secondary structure. Bear in mind, what we're looking at is a stable conformation involving one strand of DNA only. (Each of the two strands of dinG can form this structure, independently of one another.) The structure shown above has a Tm (melting temperature) of 66.4°C in 1M saline, with mean-free-energy enthalpy of minus-2418.60 kcal/mol and a 37°C Gibbs free energy (ΔG) of minus-209.92 kcal, meaning that at 37°C, formation of the stable structure shown here (or one very much like it) is, energetically speaking, strongly favored.
Structures of this kind are often considered to occur in RNA, but if they also occur in single-stranded DNA, it raises interesting questions. If a gene has an energetically stable strands-apart configuration, getting the strands of duplex B-form DNA to separate might not be so hard. But more to the point, getting the self-annealing gene to come back together again as duplex DNA will require significant energy input. In molecular genetics, we're accustomed to the idea of duplex DNA requiring help from an ATP-dependent helicase to "open up" (unwind) the double helix in preparation for replication or transcription. The above diagram suggests that the problem isn't "opening up" a gene; the greater problem may be bringing the strands together again after they've assumed a stable strands-apart secondary structure. There's a substantial energy barrier to be overcome before the above structure can be relaxed into randomly coiling DNA.
This suggests that certain genes, like dinG, may be modal in terms of strand-separation state. Once the gene's strands are apart, they want to stay apart. There's an energy barrier to bringing the strands together again.
It's ironic that a helicase gene (dinG) has so much single-strand secondary structure. The gene product is a DNA-powered helicase; the gene needs its own protein product in order to zip up again. But maybe that's the point? Maybe it's a non-accidental feature.
In general, bacteria tend to have a remarkable number of helicase genes. M. tuberculosis (for example) has 16 different helicases. Other species have even more. (See table.)
| Organism |
Helicase genes
|
| Myxococcus xanthus strain DK 1622 |
42
|
| Frankia sp. strain QA3 |
41
|
| Streptomyces cf. griseus strain XylebKG-1 |
39
|
| Clostridium botulinum Hall strain |
26
|
| Psychromonas ingrahamii strain 37 |
25
|
| Mesorhizobium sp. strain BNC1 |
21
|
| Bacillus cereus strain F837/76 |
21
|
| Escherichia coli B rel606 |
19
|
| Anabaena cylindrica strain PCC 7122 |
19
|
| Mycobacterium tuberculosis Erdman strain |
16
|
| Caulobacter crescentus strain NA1000 |
15
|
One might ask why this is so; why would a bacterium need 15, 19, 25, or 42 different helicases? It's quite unusual for a bacterial genome to have significant redundancy of genes, because when there are two copies of a given gene, one copy usually eventually becomes disabled (pseudogenized) and lost through random mutations. The very few exceptions to this rule tend to involve highly transcribed, highly necessary genes (such as ribosomal-RNA genes). It would be extremely unlikely for M. tuberculosis to carry around 16 "flavors" of a gene if they weren't all absolutely necessary. Duplicates would almost certainly be lost over time, especially in M. tuberculosis, which lacks a mismatch repair system. (Bacteria that lack mismatch repair enzymes have been shown experimentally to lose DNA fifty times faster than other bacteria.) The most parsimonious view is that the 16 helicases of M. tuberculosis are, in fact, critically necessary and perform different jobs.
I would suggest that perhaps the reason bacteria have so many helicases is that these are actually the "nucleic acid chaperones" that manage secondary structure in the sizable minority of genes that exhibit pronounced self-annealing of separated strands. It could be that most helicases are tasked with separating individual strands of DNA (and/or RNA) from themselves. Different types of secondary structure require different types of helicase to unravel. This might be why bacteria need so many helicases.
References
Helpful articles on DNA secondary structure:
- Dimitrov, R. A. & Zuker, M. (2004) Prediction of hybridization and melting for double-stranded nucleic acids. Biophys. J., 87, 215-226.
[Abstract] [Full Text] [PDF] - SantaLucia, Jr., J. (1998) A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proc. Natl. Acad. Sci. USA, 95, 1460-1465.
[Abstract] [Full Text] [PDF] - Walter, A. E., Turner, D. H., Kim, J., Lyttle, M. H., Müller, P.,
Mathews, D. H. & Zuker, M. (1994) Coaxial stacking of helixes
enhances binding of oligoribonucleotides and improves predictions of RNA
folding. Proc. Natl. Acad. Sci. USA, 91, 9218-9222.
[Abstract] [Full Text] [PDF]
Friday, June 13, 2014
Thermometer Genes
Heat shock proteins are an interesting class of proteins that provide "damage control" for enzymes when temperatures rise to the point where proteins start to unfold and refold improperly. Protein 3-dimensional structure is critical to proper enzyme function, and it doesn't take much thermal jostling to mess up a protein's structure. Therefore it's not surprising cells have their own miniature repair factories for refolding heat-misfolded proteins.
Collectively, heat shock proteins are part of a group of proteins known as chaperones, some of which are bonafide refoldases and others of which aid proteins in other ways. (For example, the ClpB protein rescues proteins from an aggregated state.)
GroEL is a so-called type I chaperonin involved in protein folding, assembly, and transport. Like many heat shock proteins, it's over-expressed at high temperatures and plays a critical role in growth and survival at non-permissive temperatures. Because of its importance in many cellular processes, GroEL is ubiquitous in bacteria, with most species having a single GroEL gene, but with about 30% of genomes having two or more GroEL copies.
The question of how cells up-regulate heat shock proteins during times of thermal stress is still largely open, although we know in some cases specific transcription regulator proteins are involved (but then the question becomes: how do the regulators know to up-regulate in times of heat stress?). The answer might not be that difficult. The messenger RNAs encoding GroEL and other heat shock proteins contain a great deal of secondary structure (that is to say, the RNA folds back on itself to form thermally sensitive structures). Recently, Wan et al. surveyed RNA thermal sensitivity in yeast and found thousands of so-called "mRNA thermometers": RNA molecules that unfold in response to heat. About three quarters of yeast RNA is thermo-stable at 37 degrees C, while around 55% of RNAs are unfolded at 55 degrees. In the folded state, RNA probably requires the help of helicases or other "helpers" to unfold, but at a high enough temperature, the molecules unfold by themselves and become eligible for translation by ribosomes.
To investigate the possible role of mRNA secondary structure in GroEL regulation, I wrote scripts that check a gene for all occurrences of length-9 (so-called "9-mer") nucleotide sequences that have a corresponding reverse-complement sequence in the same gene. When I checked the GroEL gene of Mycobacterium tuberculosis (Erdman strain), I found 14 pairs of complementarity 9-mers, representing regions of the gene that could, in theory, cause secondary structure to form in mRNA. A check of the sister organism M. intracellulare (whose GroEL gene is 80% identical to the M. tuberculosis version) showed 22 such complementary pairs.
Interestingly, the mutational differences between GroEL in M. tuberculosis and M. intracellulare do not appear to be randomly distributed along the gene. In M. tuberculosis, mutations occur at a rate of 0.12698 substitutions per site inside 9-mer regions (putative stems) versus a rate of 0.19722 for non-9-mer regions, indicating that (perhaps) selection pressure is different for self-complementing regions than for other regions. I found much the same thing in M. intracellulare, where the mutation rate was 0.16414 inside 9-mers and 0.19347 elsewhere.
Tending to confirm that selection pressure is different for the "secondary structure" regions versus other regions is the (surprising) finding that in M. tuberculosis complementary regions, the ratio of non-synonymous to synonymous mutations (Kn/Ks) is 0.526, versus 0.950 for other regions. In M. intracellulare, likewise, Kn/Ks is less in complementing regions (0.635) than in non-complementing regions (0.975).
To check whether these observations apply only to Mycobacterium or might be more widely applicable, I took a look at GroEL genes in Clostridium acetobutylicum strain ATCC 824 and Clostridium lentocellum strain DSM 5427. The Clostridia are phylogenetically quite distant from Mycobacteria (as confirmed by the fact that their GroEL genes share only 50% nucleotide sequence identity). A total of 462 mutations separated the two Clostridial genes. But again, the mutations segregated non-randomly according to whether they occurred in putative regions of complementarity (secondary structure) as opposed to non-complementing regions. In C. acetobutylicum the mutation rate in complementing regions was 0.23015 substitutions per site (29/126 bases) versus 0.28618 (433/1513 bases) for non-complementing regions, while in C. lentocellum the rates were
0.19047 (24/126) substitutions per site vs. 0.28949 (438/1513). For C. acetobutylicum the Kn/Ks ratios were 0.277 in 9-mers and 1.466 otherwise. For C. lentocellum, Kn/Ks was 0.538 in 9-mers and 1.415 outside 9-mers, tending to confirm that selection pressures are different in stems than in loops.
Bottom line, the data are consistent with a scenario in which secondary structure of GroEL mRNA (and/or ssDNA) plays a role in heat-activation of the gene, such that when temperatures exceed the melting point of secondary structures, the gene is eligible for transcription and/or translation. The gene is, in effect, its own thermometer.
Collectively, heat shock proteins are part of a group of proteins known as chaperones, some of which are bonafide refoldases and others of which aid proteins in other ways. (For example, the ClpB protein rescues proteins from an aggregated state.)
GroEL is a so-called type I chaperonin involved in protein folding, assembly, and transport. Like many heat shock proteins, it's over-expressed at high temperatures and plays a critical role in growth and survival at non-permissive temperatures. Because of its importance in many cellular processes, GroEL is ubiquitous in bacteria, with most species having a single GroEL gene, but with about 30% of genomes having two or more GroEL copies.
![]() |
| GroEL mRNA of Mycobacterium intracellulare can fold into the advanced low-energy secondary structure shown here. |
To investigate the possible role of mRNA secondary structure in GroEL regulation, I wrote scripts that check a gene for all occurrences of length-9 (so-called "9-mer") nucleotide sequences that have a corresponding reverse-complement sequence in the same gene. When I checked the GroEL gene of Mycobacterium tuberculosis (Erdman strain), I found 14 pairs of complementarity 9-mers, representing regions of the gene that could, in theory, cause secondary structure to form in mRNA. A check of the sister organism M. intracellulare (whose GroEL gene is 80% identical to the M. tuberculosis version) showed 22 such complementary pairs.
Interestingly, the mutational differences between GroEL in M. tuberculosis and M. intracellulare do not appear to be randomly distributed along the gene. In M. tuberculosis, mutations occur at a rate of 0.12698 substitutions per site inside 9-mer regions (putative stems) versus a rate of 0.19722 for non-9-mer regions, indicating that (perhaps) selection pressure is different for self-complementing regions than for other regions. I found much the same thing in M. intracellulare, where the mutation rate was 0.16414 inside 9-mers and 0.19347 elsewhere.
Tending to confirm that selection pressure is different for the "secondary structure" regions versus other regions is the (surprising) finding that in M. tuberculosis complementary regions, the ratio of non-synonymous to synonymous mutations (Kn/Ks) is 0.526, versus 0.950 for other regions. In M. intracellulare, likewise, Kn/Ks is less in complementing regions (0.635) than in non-complementing regions (0.975).
To check whether these observations apply only to Mycobacterium or might be more widely applicable, I took a look at GroEL genes in Clostridium acetobutylicum strain ATCC 824 and Clostridium lentocellum strain DSM 5427. The Clostridia are phylogenetically quite distant from Mycobacteria (as confirmed by the fact that their GroEL genes share only 50% nucleotide sequence identity). A total of 462 mutations separated the two Clostridial genes. But again, the mutations segregated non-randomly according to whether they occurred in putative regions of complementarity (secondary structure) as opposed to non-complementing regions. In C. acetobutylicum the mutation rate in complementing regions was 0.23015 substitutions per site (29/126 bases) versus 0.28618 (433/1513 bases) for non-complementing regions, while in C. lentocellum the rates were
0.19047 (24/126) substitutions per site vs. 0.28949 (438/1513). For C. acetobutylicum the Kn/Ks ratios were 0.277 in 9-mers and 1.466 otherwise. For C. lentocellum, Kn/Ks was 0.538 in 9-mers and 1.415 outside 9-mers, tending to confirm that selection pressures are different in stems than in loops.
Bottom line, the data are consistent with a scenario in which secondary structure of GroEL mRNA (and/or ssDNA) plays a role in heat-activation of the gene, such that when temperatures exceed the melting point of secondary structures, the gene is eligible for transcription and/or translation. The gene is, in effect, its own thermometer.
Monday, June 09, 2014
How Do Bacteria Survive Radiation Damage?
![]() |
| Secondary structure of the Trad_1400 gene (encoding a MutT hydrolase) in Truepera radiovictrix. |
Members of the Deinococcus group probably learned their DNA repair tricks very early in the history of terrestrial life, before there was sufficient oxygen in the atmosphere to support an ozone layer. In those times (before about one billion years ago), ultaviolet light from the sun would have been strong enough to sterilize almost any exposed surface. UV radiation, when it's strong enough, causes single- and double-strand breaks in DNA (just as ionizing radiation from radioisotopes or cosmic rays will). How Deinococcus manages to survive such radiation is still something of a mystery, although several repair modalities have been elucidated. We know that these bacteria have high copy numbers of their genetic material, and this no doubt facilitates repair. Still, double-stranded DNA breaks, in most organisms, are quickly fatal if they accumulate.
It turns out, DNA from Deinococcus-group bacteria is unusually rich in internal (intra-strand) complementarity, which means single strands of DNA are capable (in theory) of folding back on themselves to form elaborate secondary structures of high thermal stability. One such structure, for the Trad_1400 gene of Truepera radiovictrix (a radiation-tolerant member of the Deinococcus group), is shown above. This particular structure has a 37°C Gibbs free energy of minus-71.47 kcal/mol (meaning the structure is more likely to form than randomly coiled ssDNA) and a Tm (melting temperature) of 63.1°C, meaning it should be thermo-stable to around 145°F. Almost the entire gene folds back on itself; the only portion that doesn't self-anneal is the flat line on the bottom containing 29 bases.
If the (separated) strands of Truepera DNA can assume stable self-annealed structures of this type, it would go a long way toward explaining how the organism could survive double-stranded breaks. Fire a random bullet at the DNA and you're bound to hit secondary structure, not canonical (B-form) duplex DNA. A double-strand break in a stem structure might liberate a stem/loop from one strand, but the other strand could unfold to form a template for immediate repair of the damaged strand. Something like this is probably going on in radiation-resistant Deinococcus members, which have evolved to allow more than the usual secondary structure in their DNA.
Tuesday, June 03, 2014
Odd Structures in a Prophage Genome
![]() |
| Contrary to what it looks like, this is not an arrival-gate diagram for an eastern European airport. It's a diagram of secondary structure of a portion of the mRNA for the yoqJ gene in B. subtilis phage SPBc2. Click to enlarge. |
The SPBc2 genome is strange, first of all, in its base composition. The following stats come from analysis of the coding regions of the genome:
| Base | Abundance |
| A | 0.3670 |
| T | 0.2825 |
| G | 0.1983 |
| C | 0.1506 |
Notice how the purines, A and G (adenine and guanine) are about 30% more abundant than the pyrimidines, C and T (cyotsine and thymine). This is a stark violation of Chargaff's Second Parity Rule (which says A=T and G=C not only for double-stranded DNA but for a single strand, as well).
When purine abundance is tallied on a per-gene basis, you get the following histogram:
![]() |
| Purine abundance for N=185 protein-coding genes in B. subtilis phage SPBc2. Only 8 genes contain less than 50% purine bases. |
And if you look at the gene sequence data, the genes even look funny to the naked eye, containing, as they do, long runs of bases, with funny repeats, like TTGAAAGGAAAAAAAGACGGCCTAAATAAA. One gets the feeling, when examining the sequence data, that one is looking at tRNA or rRNA data rather than protein-coding sequences. And yet, they really are protein-coding sequences.
But it turns out there's a lot of secondary-structure info in the sequences. (See the picture at the top of this post.) Many of the genes contain self-complementing intragenic regions. I decided to verify this with some custom scripts. First, I had scripts look inside each gene for length-10 nucleotide sequences with a matching reverse-complement sequence further downstream (in the same gene). I expected to find one or two (or ten) hits, based on the fact that any given length-10 sequence of the bases A, G, C, and T can turn up at random in about one in every million base pairs. (Each base is worth about two bits, so a length-10 DNA sequence can encode as much info as a length-20 binary sequence; 2-to-the-20th is a little over a million.) What I found was an astounding 419 complementary length-10 sequences within 89 genes.
When I upped the sequence length to 12, I still found 64 intragenic complement sequences in 38 genes. A length-12 match should happen randomly about once every 16 million base pairs. The SPBc2 genome (as I said) is 134,416 base pairs long.
The high degree of internal complementarity means that when the phage's DNA strands are separated during transcription or replication, they probably assume very particular 3-dimensional structures, and also, the messenger RNA made from the phage's genes are probably rich in secondary structure. Exactly why those structures are needed is anyone's guess. They may protect the virus from restriction nucleases (although frankly this chore is already taken care of by the prophage's own methylases). Or they may attract ribosomes in some special way (many of the genes form tRNA-lookalike structures). They may help with virion packaging. Or they may do nothing special at all. I kind of doubt the latter possibility, though.
Sunday, June 01, 2014
A Gene Heatmap
Lately I've been using the great tools at genomevolution.org plus custom Canvas API scripts to render colorful heatmaps of aligned genes from phylogeneticaly diverse microorganisms. The following graphic is one such.
What are we looking at? This is actually a composite rendering of the glyceraldehyde-3-phosphate dehydrogenase genes (DNA sequence info) from 134 bacterial species. Each gene is painted left-to-right (5' to 3') in a strip 4 pixels tall, with hot colors assigned to DNA bases G and C (guanine and cytosine), and cool colors assigned to bases A and T (adenine and thymine). Wherever there's a G or C, it gets painted red or red-orange. Wherever there's A or T, blue or blue-green. Same gene, 134 versions, varying significantly in G+C content. (The gene GC content ranges from a maximum of 69.2% at the top to 29.4% at the bottom.)
Why glyceraldehyde-3-phosphate dehydrogenase (GAPDH)? No real reason, except that it's a fairly universal (indeed, quite ancient) metabolic enzyme, reasonably compact (making possible a rendering that's not super-wide, as it would be for a larger gene), well-delineated genetically (not a fusion protein or an enzyme with multiple isoforms), and probably representative of a good many core metabolic enzymes. This is the enzyme that catalyzes the sixth step of glycolysis (sugar-breakdown). You may recall from Biochem 101 that the breakdown of glucose proceeds by splitting the twice phosphorylated molecule into two 3-carbon pieces. The triose phosphates in turn get phosphorylated by GAPDH before they transfer a phosphate to ADP to yield ATP, the 5-hour energy drink of all cells everywhere.
Alignment of genes was done via ClustalW in MEGA6 freeware. Rendering of the alignment FASTA file took about two seconds, in the browser, using 133 lines of custom JavaScript.
![]() |
| Glyceraldehyde-3-phosphate dehydrogenase genes of N=134 bacterial species, arranged in order of gene G+C content (high GC at the top). Hot colors are G and C. Cool colors are A and T. |
What are we looking at? This is actually a composite rendering of the glyceraldehyde-3-phosphate dehydrogenase genes (DNA sequence info) from 134 bacterial species. Each gene is painted left-to-right (5' to 3') in a strip 4 pixels tall, with hot colors assigned to DNA bases G and C (guanine and cytosine), and cool colors assigned to bases A and T (adenine and thymine). Wherever there's a G or C, it gets painted red or red-orange. Wherever there's A or T, blue or blue-green. Same gene, 134 versions, varying significantly in G+C content. (The gene GC content ranges from a maximum of 69.2% at the top to 29.4% at the bottom.)
Why glyceraldehyde-3-phosphate dehydrogenase (GAPDH)? No real reason, except that it's a fairly universal (indeed, quite ancient) metabolic enzyme, reasonably compact (making possible a rendering that's not super-wide, as it would be for a larger gene), well-delineated genetically (not a fusion protein or an enzyme with multiple isoforms), and probably representative of a good many core metabolic enzymes. This is the enzyme that catalyzes the sixth step of glycolysis (sugar-breakdown). You may recall from Biochem 101 that the breakdown of glucose proceeds by splitting the twice phosphorylated molecule into two 3-carbon pieces. The triose phosphates in turn get phosphorylated by GAPDH before they transfer a phosphate to ADP to yield ATP, the 5-hour energy drink of all cells everywhere.
Alignment of genes was done via ClustalW in MEGA6 freeware. Rendering of the alignment FASTA file took about two seconds, in the browser, using 133 lines of custom JavaScript.
Saturday, May 31, 2014
Not All Organisms Use Amino Acids the Same Ways
Organisms vary greatly in the GC (guanine plus cytosine) content of their DNA, and yet all organisms can still make ribosomal proteins, DNA and RNA polymerases, and the various other essential proteins of life, no matter what their DNA vocabulary limitations might be. A high-GC organism like Streptomyces can make a given enzyme (DNA polymerase, say) using mostly G and C bases in its DNA, but a low-GC organism like Clostridium botulinum can also make the same kind of enzyme, even though it uses mostly A and T in its DNA. How is this possible?
It's possible in part because of the many synonyms for amino acids available in the genetic code. But it's a mistake to think the same amino acids are used in equal numbers by high-GC organisms and low-GC organisms. Organisms at opposite ends of the GC spectrum use different amino acids.
I was curious to see which amino acids correlate most strongly with genomic GC, so I gathered codon usage statistics for 109 organisms of widely varying genomic GC content and used JavaScript to calculate Pearson correlation coefficients for all 20 amino acids with respect to GC content. The results are shown in the following table.
TABLE 1. Correlation (r) between amino acid usage and genome GC content (N=109 organisms).
Ten amino acids correlate positively with GC and ten correlate negatively. Alanine and arginine have the strongest positive correlation with GC, while isoleucine and asparagine have the strongest negative correlation with genomic GC content. (But note that these data apply only to the 109 organisms studied. For the complete list of 109 organisms, see this post.)
If you were to extract all the amino acids out of Clostridium botulinum (28% GC), you would get far more lysine than alanine. Conversely, if you were to hydrolyze all the proteins in Streptomyces griseus (GC 72%), you would find far more alanine than lysine.
Interestingly, serine has six synonymous codons (AGT, AGC, CTA, CTG, CTC, CTT) and can just as easily be specified with G and C as with A and T; so overall, you'd expect little correlation with genomic GC. And yet serine use correlates strongly with low GC. In a sense, this is not surprising. Certain low-GC organisms (like Streptococcus) are known to produce serine-rich cell-coat proteins, some of which are important determinants of pathogenicity. But it may simply be that the high utilization of serine in low-GC organisms is related to one-carbon chemistry. Serine, after all, is the source of the methyl group that, by way of methylenetetrahydrofolate, converts dUMP to TMP (thymidine monophosphate, a DNA precursor). Any organism whose DNA is unusually rich in thymine (low in GC) will almost certainly be processing large quantities of serine. Serine is also a carbon source in the biosynthetic pathways for cysteine and methionine, both of which, like serine itself, are negatively correlated with genomic GC content.
It's possible in part because of the many synonyms for amino acids available in the genetic code. But it's a mistake to think the same amino acids are used in equal numbers by high-GC organisms and low-GC organisms. Organisms at opposite ends of the GC spectrum use different amino acids.
I was curious to see which amino acids correlate most strongly with genomic GC, so I gathered codon usage statistics for 109 organisms of widely varying genomic GC content and used JavaScript to calculate Pearson correlation coefficients for all 20 amino acids with respect to GC content. The results are shown in the following table.
TABLE 1. Correlation (r) between amino acid usage and genome GC content (N=109 organisms).
Code
|
Amino Acid |
r
|
A
|
Alanine (Ala) |
0.9634
|
R
|
Arginine (Arg) |
0.9495
|
G
|
Glycine (Gly) |
0.9472
|
P
|
Proline (Pro) |
0.9436
|
V
|
Valine (Val) |
0.7725
|
W
|
Tryptophan (Trp) |
0.7497
|
H
|
Histidine (His) |
0.4660
|
L
|
Leucine (Leu) |
0.3364
|
D
|
Aspartic Acid (Asp) |
0.3347
|
T
|
Threonine (Thr) |
0.3099
|
C
|
Cysteine (Cys) |
-0.1280
|
Q
|
Glutamine (Gln) |
-0.1668
|
M
|
Methionine (Met) |
-0.2863
|
E
|
Glutamic Acid (Glu) |
-0.4621
|
S
|
Serine (Ser) |
-0.6831
|
F
|
Phenylalanine (Phe) |
-0.8550
|
Y
|
Tyrosine (Tyr) |
-0.8983
|
K
|
Lysine (Lys) |
-0.9389
|
N
|
Asparagine (Asn) |
-0.9391
|
I
|
Isoleucine (Ile) |
-0.9558
|
Ten amino acids correlate positively with GC and ten correlate negatively. Alanine and arginine have the strongest positive correlation with GC, while isoleucine and asparagine have the strongest negative correlation with genomic GC content. (But note that these data apply only to the 109 organisms studied. For the complete list of 109 organisms, see this post.)
If you were to extract all the amino acids out of Clostridium botulinum (28% GC), you would get far more lysine than alanine. Conversely, if you were to hydrolyze all the proteins in Streptomyces griseus (GC 72%), you would find far more alanine than lysine.
Interestingly, serine has six synonymous codons (AGT, AGC, CTA, CTG, CTC, CTT) and can just as easily be specified with G and C as with A and T; so overall, you'd expect little correlation with genomic GC. And yet serine use correlates strongly with low GC. In a sense, this is not surprising. Certain low-GC organisms (like Streptococcus) are known to produce serine-rich cell-coat proteins, some of which are important determinants of pathogenicity. But it may simply be that the high utilization of serine in low-GC organisms is related to one-carbon chemistry. Serine, after all, is the source of the methyl group that, by way of methylenetetrahydrofolate, converts dUMP to TMP (thymidine monophosphate, a DNA precursor). Any organism whose DNA is unusually rich in thymine (low in GC) will almost certainly be processing large quantities of serine. Serine is also a carbon source in the biosynthetic pathways for cysteine and methionine, both of which, like serine itself, are negatively correlated with genomic GC content.
Friday, May 30, 2014
Comparative Genomics Using the Canvas API
An interesting and somewhat mysterious aspect of biodiversity is that the relative proportions of the bases in DNA can take on wildly different values in different organisms even though they're making many of the same proteins. I'm referring to the fact that the G+C (guanine plus cytosine) content of DNA can vary from more than 70% (e.g., Streptomyces species) to less than 30% for certain bacterial endosymbionts (and even for some free-living bacteria, such as Clostridium botulinum). Remarkably, the DNA of the tiny bacterium Buchnera aphidicola (which is distantly related to E. coli but entered into a symbiotic partnership with the aphid around 200 million years ago) has a GC content of only 26%, making its DNA look almost like a two-letter code (A and T, with the occasional G or C).
Many genome-reduced endosymbionts have lost most of their DNA repair enzymes (this is true of mitochondria, incidentally, which are thought to have arisen from a symbiosis with an ancient ancestor of today's Alphabproteobacteria), and this loss of repair capability could well explain much of the GC-to-AT shift seen in endosymbiont genomes. (Left on its own, DNA tends to accumulate 8-oxo-guanine residues, which incorrectly pair with thymine and result in GC-to-TA transversions at DNA replication time.) Whatever the cause(s), GC reduction has left many organisms with severely A- and T-enriched DNA. But the organisms in question are still able to encode perfectly functional proteins, even with a limited nucleic-acid vocabulary.
I decided it might be interesting to try to visualize "GC diversity" by creating a heat map of G and C (in hot colors) and A and T (in cool colors) for certain genes that occur across all bacteria. For example, the following graphic is a heat map of GC and AT usage in the gene for thymidine kinase as it occurs in 61 different organisms with genomic G+C ranging from just over 70% to just under 30%.
To create this map, I obtained DNA sequences for thymidine kinase from 61 organisms (using the excellent online tools at UniProt.org and genomevolution.org), then aligned them in MEGA6 and drew colors (red or red-orange for G or C, blue or blue-green for A or T) corresponding to the bases, using the Canvas API. What you're looking at are 61 rows of data (one row per gene; which is also to say, per organism). Gaps created during alignment are shown in grey.
Several things are apparent from this graphic, aside from the obvious fact that GC usage (indicated by red and orange) tends to be high in the genes for organisms like Brachybacterium faecium (top line, GC 72%) and low for bottom-of-the-chart organisms like Clostridium perfringens B strain ATCC 3626 (28.7% GC) and Ureaplasma urealyticum (25.9%). First, high-GC genes tend to be somewhat longer. (Indeed, in the upper right you can see that the longest genes opened up a sizable alignment gap that extends down the whole graphic.) Also, the genes differ substantially in their leader and trailer sequences, although I think what we're really seeing here is (at least in part) inaccurate annotation of start and stop codons. What's interesting to me is the way certain GC regions trail all the way down to the bottom of the graph while others fade to blue. I think it could be argued that the nucleotides represented by the red blips in the final few lines of the graph, at the bottom, are positions in the gene that are under strong functional constraints. It would be interesting to test those positions for evidence of selection pressure. It could be argued that all areas of low selection pressure have turned blue by the time you get to the bottom of the graph.
I think it's also interesting that several red-orange clusters just to the right of the midpoint, about three-quarters of the way down the graph, all but disappear just as a new red-orange zone begins to appear on the left around 80% of the way down. It's as if certain GC-to-AT mutations in one protein domain led to AT-to-GC mutations in another domain upstream. (The 5' side of the graph is on the left, 3' on the right.)
Getting this many genes (from such divergent sources) to align is not easy. You can do it, though, by setting the gap-open penalty in ClustalW to a low value and aligning genes in small batches, row by row, if need be.
As a technical aside: I first tried creating this graph using SVG (Scalable Vector Graphics), but the burden of creating a separate DOM node for every pixel (which is what it amounts to) was way too much for the browser to handle (Firefox choked, as did Chrome), so I quickly switched to the Canvas API, which puts no heavy DOM burdens on the browser and can convert a FASTA-formatted alignment file to a nice picture in about two seconds.
For what it's worth, here are the names of the organisms whose genes appeared in the above graphic, arranged in order of GC content of the kinase gene only (not whole-genome GC). High-GC organisms are listed first:
Blastococcus saxobsidens strain DD2
Geodermatophilus obscurus strain DSM 43160
Streptomyces cf. griseus strain XylebKG-1
Streptosporangium roseum strain DSM 43021
Kribbella flavida strain DSM 17836
Brachybacterium faecium strain DSM 4810
Deinococcus radiodurans strain R1
Gordonia bronchialis strain DSM 43247
Mesorhizobium australicum strain WSM2073
Turneriella parva strain DSM 21527
Mesorhizobium ciceri biovar biserrulae strain WSM1271
Propionibacterium acnes TypeIA2 strain P.acn33
Novosphingobium aromaticivorans strain DSM 12444
Geobacillus kaustophilus strain HTA426
Geobacillus thermoleovorans strain CCB_US3_UF5
Halogeometricum borinquense DSM 11551
Alistipes finegoldii strain DSM 17242
Agrobacterium radiobacter strain K84
Rhizobium tropici strain CIAT 899
Aeromonas hydrophila strain ML09-119
Rhodopirellula baltica SH strain 1
Porphyromonas gingivalis strain ATCC 33277
Dyadobacter fermentans strain DSM 18053
Enterobacter cloacae strain SCF1
Paenibacillus polymyxa strain M1
Parabacteroides distasonis strain ATCC 8503
Bacillus amyloliquefaciens strain Y2
Vibrio cholerae strain BX 330286
Erwinia amylovora strain ATCC 49946
Bacillus subtilis BEST7613 strain PCC 6803
Gramella forsetii strain KT0803
Prevotella copri strain DSM 18205
Anaerolinea thermophila strain UNI-1
Klebsiella oxytoca strain 10-5243
Bacteroides ovatus strain 3_8_47FAA
Aggregatibacter phage S1249
Escherichia coli B strain REL606
Shigella boydii strain Sb227
Psychroflexus torquis strain ATCC 700755
Bacteroides dorei strain 5_1_36/D4
Leuconostoc gasicomitatum LMG 18811 strain type LMG 18811
Mycoplasma gallisepticum strain F
Myroides odoratimimus strain CIP 101113
Oceanobacillus kimchii strain X50
Chryseobacterium gleum strain ATCC 35910
Carboxydothermus hydrogenoformans strain Z-2901
Yersinia pestis D106004
Bacillus thuringiensis serovar andalousiensis strain BGSC 4AW1
Bacillus anthracis strain CDC 684
Coprobacillus sp. strain 8_2_54BFAA
Proteus mirabilis strain HI4320
Staphylococcus aureus strain 04-02981
Aerococcus urinae strain ACS-120-V-Col10a
Lactobacillus acidophilus strain 30SC
Lactobacillus reuteri strain MM4-1A
Lactococcus lactis subsp. cremoris strain A76
Haemophilus ducreyi strain 35000HP
Ureaplasma urealyticum serovar 5 strain ATCC 27817
Clostridium botulinum A strain Hall
Streptococcus agalactiae strain 2603V/R
Clostridium perfringens B strain ATCC 3626
Many genome-reduced endosymbionts have lost most of their DNA repair enzymes (this is true of mitochondria, incidentally, which are thought to have arisen from a symbiosis with an ancient ancestor of today's Alphabproteobacteria), and this loss of repair capability could well explain much of the GC-to-AT shift seen in endosymbiont genomes. (Left on its own, DNA tends to accumulate 8-oxo-guanine residues, which incorrectly pair with thymine and result in GC-to-TA transversions at DNA replication time.) Whatever the cause(s), GC reduction has left many organisms with severely A- and T-enriched DNA. But the organisms in question are still able to encode perfectly functional proteins, even with a limited nucleic-acid vocabulary.
I decided it might be interesting to try to visualize "GC diversity" by creating a heat map of G and C (in hot colors) and A and T (in cool colors) for certain genes that occur across all bacteria. For example, the following graphic is a heat map of GC and AT usage in the gene for thymidine kinase as it occurs in 61 different organisms with genomic G+C ranging from just over 70% to just under 30%.
To create this map, I obtained DNA sequences for thymidine kinase from 61 organisms (using the excellent online tools at UniProt.org and genomevolution.org), then aligned them in MEGA6 and drew colors (red or red-orange for G or C, blue or blue-green for A or T) corresponding to the bases, using the Canvas API. What you're looking at are 61 rows of data (one row per gene; which is also to say, per organism). Gaps created during alignment are shown in grey.
Several things are apparent from this graphic, aside from the obvious fact that GC usage (indicated by red and orange) tends to be high in the genes for organisms like Brachybacterium faecium (top line, GC 72%) and low for bottom-of-the-chart organisms like Clostridium perfringens B strain ATCC 3626 (28.7% GC) and Ureaplasma urealyticum (25.9%). First, high-GC genes tend to be somewhat longer. (Indeed, in the upper right you can see that the longest genes opened up a sizable alignment gap that extends down the whole graphic.) Also, the genes differ substantially in their leader and trailer sequences, although I think what we're really seeing here is (at least in part) inaccurate annotation of start and stop codons. What's interesting to me is the way certain GC regions trail all the way down to the bottom of the graph while others fade to blue. I think it could be argued that the nucleotides represented by the red blips in the final few lines of the graph, at the bottom, are positions in the gene that are under strong functional constraints. It would be interesting to test those positions for evidence of selection pressure. It could be argued that all areas of low selection pressure have turned blue by the time you get to the bottom of the graph.
I think it's also interesting that several red-orange clusters just to the right of the midpoint, about three-quarters of the way down the graph, all but disappear just as a new red-orange zone begins to appear on the left around 80% of the way down. It's as if certain GC-to-AT mutations in one protein domain led to AT-to-GC mutations in another domain upstream. (The 5' side of the graph is on the left, 3' on the right.)
Getting this many genes (from such divergent sources) to align is not easy. You can do it, though, by setting the gap-open penalty in ClustalW to a low value and aligning genes in small batches, row by row, if need be.
As a technical aside: I first tried creating this graph using SVG (Scalable Vector Graphics), but the burden of creating a separate DOM node for every pixel (which is what it amounts to) was way too much for the browser to handle (Firefox choked, as did Chrome), so I quickly switched to the Canvas API, which puts no heavy DOM burdens on the browser and can convert a FASTA-formatted alignment file to a nice picture in about two seconds.
For what it's worth, here are the names of the organisms whose genes appeared in the above graphic, arranged in order of GC content of the kinase gene only (not whole-genome GC). High-GC organisms are listed first:
Blastococcus saxobsidens strain DD2
Geodermatophilus obscurus strain DSM 43160
Streptomyces cf. griseus strain XylebKG-1
Streptosporangium roseum strain DSM 43021
Kribbella flavida strain DSM 17836
Brachybacterium faecium strain DSM 4810
Deinococcus radiodurans strain R1
Gordonia bronchialis strain DSM 43247
Mesorhizobium australicum strain WSM2073
Turneriella parva strain DSM 21527
Mesorhizobium ciceri biovar biserrulae strain WSM1271
Propionibacterium acnes TypeIA2 strain P.acn33
Novosphingobium aromaticivorans strain DSM 12444
Geobacillus kaustophilus strain HTA426
Geobacillus thermoleovorans strain CCB_US3_UF5
Halogeometricum borinquense DSM 11551
Alistipes finegoldii strain DSM 17242
Agrobacterium radiobacter strain K84
Rhizobium tropici strain CIAT 899
Aeromonas hydrophila strain ML09-119
Rhodopirellula baltica SH strain 1
Porphyromonas gingivalis strain ATCC 33277
Dyadobacter fermentans strain DSM 18053
Enterobacter cloacae strain SCF1
Paenibacillus polymyxa strain M1
Parabacteroides distasonis strain ATCC 8503
Bacillus amyloliquefaciens strain Y2
Vibrio cholerae strain BX 330286
Erwinia amylovora strain ATCC 49946
Bacillus subtilis BEST7613 strain PCC 6803
Gramella forsetii strain KT0803
Prevotella copri strain DSM 18205
Anaerolinea thermophila strain UNI-1
Klebsiella oxytoca strain 10-5243
Bacteroides ovatus strain 3_8_47FAA
Aggregatibacter phage S1249
Escherichia coli B strain REL606
Shigella boydii strain Sb227
Psychroflexus torquis strain ATCC 700755
Bacteroides dorei strain 5_1_36/D4
Leuconostoc gasicomitatum LMG 18811 strain type LMG 18811
Mycoplasma gallisepticum strain F
Myroides odoratimimus strain CIP 101113
Oceanobacillus kimchii strain X50
Chryseobacterium gleum strain ATCC 35910
Carboxydothermus hydrogenoformans strain Z-2901
Yersinia pestis D106004
Bacillus thuringiensis serovar andalousiensis strain BGSC 4AW1
Bacillus anthracis strain CDC 684
Coprobacillus sp. strain 8_2_54BFAA
Proteus mirabilis strain HI4320
Staphylococcus aureus strain 04-02981
Aerococcus urinae strain ACS-120-V-Col10a
Lactobacillus acidophilus strain 30SC
Lactobacillus reuteri strain MM4-1A
Lactococcus lactis subsp. cremoris strain A76
Haemophilus ducreyi strain 35000HP
Ureaplasma urealyticum serovar 5 strain ATCC 27817
Clostridium botulinum A strain Hall
Streptococcus agalactiae strain 2603V/R
Clostridium perfringens B strain ATCC 3626
Tuesday, May 27, 2014
An Evolutionary Debate that Misses the Point
One of many hotly debated topics in evolutionary biology is how codon usage bias (preference of an organism for certain codons, when other, synonymous variants are available) relates to transfer-RNA abundance. It's clear the two are related; no one disagrees on that. The question is whether codon usage bias is an outcome of tRNA abundance ratios, or the reverse.
What I think most people are missing in this argument is that the whole discussion might very well be mooted by a huge factor in tRNA evolution that no one seems to be taking into account. I'm talking about the fact that tRNAs are insertion targets for various kinds of mobile genetic elements, from phages to plasmid-borne genomic islands to transposons. As Jörg Hacker and Elisabeth Carniel point out in "Ecological fitness, genomic islands and bacterial pathogenicity" (EMBO Reports, 2001):
Many examples can be found of ancient tRNA signatures inside the tail ends of protein genes, no doubt leftovers from millions of years of insertion events.
What I think most people are missing in this argument is that the whole discussion might very well be mooted by a huge factor in tRNA evolution that no one seems to be taking into account. I'm talking about the fact that tRNAs are insertion targets for various kinds of mobile genetic elements, from phages to plasmid-borne genomic islands to transposons. As Jörg Hacker and Elisabeth Carniel point out in "Ecological fitness, genomic islands and bacterial pathogenicity" (EMBO Reports, 2001):
Genomic islands are part of the flexible bacterial gene pool and are somewhere between 10 and 100 kilobases (kb) in length. They frequently harbor phage- and/or plasmid-derived sequences, including transfer genes or integrases and IS elements. These particular blocks of DNA are most often inserted into tRNA genes and may be unstable.(Emphasis added.) Transfer RNAs are constantly being "inserted into" (and next to, not always into) by mobile elements, a phenomenon that's been well studied not only in bacteria but in yeast and elsewhere. Over evolutionary timespans, tRNA genes are duplicated, then disrupted, over and over again, by mobile DNA elements. These elements (whether from phages, viruses, transposons, or what have you) are known to have played (and continue to play) a significant role in shaping genome diversity, across all taxa. This is not a trivial factor, in other words. Transfer RNA genes are insertion hotspots. Surely the patterns of tRNA disruption caused by gene-hopping, over the eons, cannot be unimportant in the determination of codon usage patterns.
Many examples can be found of ancient tRNA signatures inside the tail ends of protein genes, no doubt leftovers from millions of years of insertion events.
Monday, May 26, 2014
Shannon Entropy and DNA
Claude Shannon made an important finding when he realized that the information contribution of a symbol could be estimated very simply as -f(x) log(f(x)), where f(x) is the frequency of occurrence of the symbol x. For example, a series of a coin tosses can be considered a binary information stream with symbol values H and T (heads and tails). If the frequency of H is 0.5 and f(T) is 0.5, the entropy E, in bits per toss, is -0.5 times log (base 2) 0.5 for heads, and a similar value for tails. The values add up (in this case) to 1.0. The intuitive meaning of 1.0 (the Shannon entropy) is that a single coin toss conveys 1.0 bit of information. Contrast this with the situation that prevails when using a "weighted" or unfair penny that lands heads-up 70% of the time. We know intuitively that tossing such a coin will produce less information because we can predict the outcome (heads), to a degree. Something that's predictable is uninformative. Shannon's equation gives -0.7 times log(0.7) = 0.3602 for heads and -0.3 * log (0.3) = 0.5211 for tails, for an entropy of 0.8813 bits per toss. In this case we can say that a toss is 11.87% redundant.
DNA is an information stream resembling a series of four-sided-coin tosses, where the "coin" can land with values of A, T, G, or C. In some organisms, the four bases occur with equal rates (25% each), in which case the DNA has a Shannon entropy of 2.0 bits per base (which makes sense, in that a base can encode one of 22 possible values). But what about organisms in which the bases occur with unequal frequencies? For example, we know that many organisms have DNA with G+C content quite a bit less than (or in some cases more than) 50%. The information content of the DNA will be less than 2 bits per base in such cases.
As an example, let's take Clostridium botulinum (the source of "Botox" serum), a soil bacterium with unusually low G+C content, at 28%. If we go through the organism's 3,404 protein-coding genes, we find actual base contents of:
A 0.40189
T 0.30603
G 0.18255
C 0.10840
These numbers are for a single strand (the coding strand or "message" strand) of DNA, which is why A and T aren't equal. For whole DNA, of course, A =T and G = C, but that's not the case here. We're just interested in the message strand.
If we put the above base frequencies into the Shannon equation, we come up with a value of 1.8467 for the information content (in bits) of one base of C. botulinum DNA. The DNA is about 7.67% redundant on a zero-order entropy basis. The DNA may be over 70% A and T, but it's a long way from being a two-base (one bit) information stream. Each base encodes an average of 1.8467 bits of information, which is a surprising amount (surprisingly close to 2.0) for such a skewed alphabet.
![]() |
| Claude Shannon |
As an example, let's take Clostridium botulinum (the source of "Botox" serum), a soil bacterium with unusually low G+C content, at 28%. If we go through the organism's 3,404 protein-coding genes, we find actual base contents of:
A 0.40189
T 0.30603
G 0.18255
C 0.10840
These numbers are for a single strand (the coding strand or "message" strand) of DNA, which is why A and T aren't equal. For whole DNA, of course, A =T and G = C, but that's not the case here. We're just interested in the message strand.
If we put the above base frequencies into the Shannon equation, we come up with a value of 1.8467 for the information content (in bits) of one base of C. botulinum DNA. The DNA is about 7.67% redundant on a zero-order entropy basis. The DNA may be over 70% A and T, but it's a long way from being a two-base (one bit) information stream. Each base encodes an average of 1.8467 bits of information, which is a surprising amount (surprisingly close to 2.0) for such a skewed alphabet.
Friday, May 23, 2014
Looking for LUCA
In 1964, Emile Zuckerkandl and Linus Pauling wrote a paper (published the following year) for the Journal of Theoretical Biology suggesting the use of amino-acid and nucleic-acid sequences for deducing phylogenetic relationships. Ever since then, biologists have been trying to use sequence data to get to the root of the tree of life. Darwinian logic says that at some point, all cells had to have diverged from a Last Universal Common Ancestor (LUCA). Unfortunately, as pointed out by Doolittle and others, the quest for LUCA is greatly complicated by mutational saturation effects, reductive genome loss in important members of the most ancient taxa, convergent evolution, and non-negligible (yet difficult to estimate) amounts of horizontal gene transfer, among other serious problems.
The difficulty (I won't say folly) of trying to construct a well-rooted tree of life is made evident in various failed attempts to trace common descent via protein sequences. In January 2010, a few months after the sequencing of the 1000th bacterial genome, Karin Lagesen, Dave W. Ussery, and Trudy M. Wassenaar published a paper in which they expressed surprise over the fact that when they looked all 1,000 then-existing genomes (the number is now more than 10 times that), they could not find a single protein that was conserved across all bacteria. (Here, "conserved" means >50% amino-acid sequence identity.) Harris et al. took a slightly different approach, using the Clusters of Orthologous Groups (COG) database to search for universally conserved genes that follow the same phylogenetic patterns as ribosomal RNA (and therefore might constitute the ancestral genetic core of today's cells). The upshot:
These disappointing results are understandable and perhaps expected, given the huge amount of deck-reshuffling that's happened in three billion years. It might well be that genome sequence data, with its constant churn, represents the wrong level of granularity for deep-phylogenetic studies. What matters for organisms, after all, is function, and function is an outcome of protein tertiary structure, not just primary structure.
With that in mind, Kyung Mo Kim and Gustavo Caetano-Anollés in 2011 published a brilliant study in BMC Evolutionary Biology called "The proteomic complexity and rise of the primordial ancestor of diversified life," relying on major structural motifs as the unit of phylogenetic discrimination. Defining protein domains at the highly conserved fold superfamily (FSF) level of structure, Kim and Caetano-Anollés used an iterative, parsimony-based phylogenomic approach to reconstructing FSF repertoires as upper and lower bounds of a presumed urancestral proteome ("ur" here meaning universal). Their conclusion:
While parasitic microorganisms were found to occupy some of the most ancient branches of the superkingdom tree, Kim and Caetano-Anollés nevertheless decided to omit such organisms from their study since reductive evolution (wholesale loss of entire families of enzymes and their control systems) might otherwise queer the results. The final set of free-living organisms included 48 archaeal, 239 bacterial, and 133 eukaryotic members. To avoid potential problems with long-branch attraction, the researchers wisely sampled (at random) equal numbers of proteomes per superkingdom and replicated trees of proteomes, so that bacterial data (which of course predominated) wouldn't swamp archaea or eukaryota.
Among the many fascinating findings in the study:
The authors have many interesting things to say about the evolution of archaeal and bacterial membrane-lipid chemistry (and much else). If you're a biologist and you haven't yet read the Kim and Caetano-Anollés paper, do yourself favor and take a look at it now. It's a fascinating read, no matter what side of the LUCA fence you're on.
![]() |
| An evolutionary tree of life based on analysis of N=420 genomes of free-living organisms. Proteomes are taxa and protein fold superfamilies are character data. Adapted from Kim and Caetano-Anollés, BMC Evolutionary Biology (2011), 11:140. Click to enlarge. See text for discussion. |
The difficulty (I won't say folly) of trying to construct a well-rooted tree of life is made evident in various failed attempts to trace common descent via protein sequences. In January 2010, a few months after the sequencing of the 1000th bacterial genome, Karin Lagesen, Dave W. Ussery, and Trudy M. Wassenaar published a paper in which they expressed surprise over the fact that when they looked all 1,000 then-existing genomes (the number is now more than 10 times that), they could not find a single protein that was conserved across all bacteria. (Here, "conserved" means >50% amino-acid sequence identity.) Harris et al. took a slightly different approach, using the Clusters of Orthologous Groups (COG) database to search for universally conserved genes that follow the same phylogenetic patterns as ribosomal RNA (and therefore might constitute the ancestral genetic core of today's cells). The upshot:
Of the roughly 3100 COGs analyzed, only 80 were found to occur in all organisms. Fifty of these universally present genes showed the same phylogenetic relationships as rRNA.Harris et al. found that the majority of universally conserved three-domain COG genes (37 of 50) are physically associated with the ribosome. Surprisingly, they found that "relatively few genes encoding proteins involved in DNA replication or transcription from DNA to RNA proved to be three-domain." In particular, RNA polymerases (except for certain subunits) did not follow rRNA distribution patterns and are not conserved across the three domains of life (archaea, bacteria, and eukaryotes). Moreover, the only component of the replicative DNA polymerase in modern cells that was found to be conserved across domains was DnaN (COG0592), the gene for the “sliding clamp.”
These disappointing results are understandable and perhaps expected, given the huge amount of deck-reshuffling that's happened in three billion years. It might well be that genome sequence data, with its constant churn, represents the wrong level of granularity for deep-phylogenetic studies. What matters for organisms, after all, is function, and function is an outcome of protein tertiary structure, not just primary structure.
With that in mind, Kyung Mo Kim and Gustavo Caetano-Anollés in 2011 published a brilliant study in BMC Evolutionary Biology called "The proteomic complexity and rise of the primordial ancestor of diversified life," relying on major structural motifs as the unit of phylogenetic discrimination. Defining protein domains at the highly conserved fold superfamily (FSF) level of structure, Kim and Caetano-Anollés used an iterative, parsimony-based phylogenomic approach to reconstructing FSF repertoires as upper and lower bounds of a presumed urancestral proteome ("ur" here meaning universal). Their conclusion:
The minimum urancestral FSF set reveals the urancestor had advanced metabolic capabilities, was especially rich in nucleotide metabolism enzymes, had pathways for the biosynthesis of membrane sn1,2 glycerol ester and ether lipids, and had crucial elements of translation, including a primordial ribosome with protein synthesis capabilities. It lacked however fundamental functions, including transcription, processes for extracellular communication, and enzymes for deoxyribonucleotide synthesis. Proteomic history reveals the urancestor is closer to a simple progenote organism but harbors a rather complex set of modern molecular functions.The paper is quite long (14,700 words) and often relentlessly technical, but convincingly restores the quest for LUCA to the firm empirical grounding that such a quest seemed (for a while) to have been robbed of after Doolittle's "Uprooting the Tree of Life" and Dagan and Martin's "The Tree of One Percent."
While parasitic microorganisms were found to occupy some of the most ancient branches of the superkingdom tree, Kim and Caetano-Anollés nevertheless decided to omit such organisms from their study since reductive evolution (wholesale loss of entire families of enzymes and their control systems) might otherwise queer the results. The final set of free-living organisms included 48 archaeal, 239 bacterial, and 133 eukaryotic members. To avoid potential problems with long-branch attraction, the researchers wisely sampled (at random) equal numbers of proteomes per superkingdom and replicated trees of proteomes, so that bacterial data (which of course predominated) wouldn't swamp archaea or eukaryota.
Among the many fascinating findings in the study:
- The earliest start of organismal diversification occurred sometime between 2.91 and 2.03 billion years ago.
- Translation had metabolic origins. It appeared only after the emergence "of a large number of metabolic functions, but before enzymes necessary for the synthesis of DNA."
- Proteomic analysis of extant fold superfamilies (FSFs) showed that "over 200 additional FSFs are necessary in urancestral FSF sets to account for the complexity of the simplest organism in existence today."
- None of the domains present in ribonucleotide reductase (RDR) enzymes was present in the min_set (representing the LUCA lower bound of complexity). Further, "We note that the reduction of ribonucleotides to deoxyribonucleotides involves the production of an active site thiyl radical that requires contacts with cysteines in all protein domains of the catalytic subunit of the oligomeric enzymatic complex, suggesting modern ribonucleotide reductase functions is [sic] indeed derived."
- Commenting on the known active-site domain homology between class III ribonucleotide reductase and pyruvate formate lyase (a link proposed to have mediated the RNA-to-DNA biological transition), Kim and Caetano-Anollés point out that phylogenomic analysis at the fold-family level suggests the pyruvate formate-lyase domain emerged later than its ribonucleotide reductase counterpart. Therefore it's likely that the urancestor stored genetic information as RNA and not DNA.
The authors have many interesting things to say about the evolution of archaeal and bacterial membrane-lipid chemistry (and much else). If you're a biologist and you haven't yet read the Kim and Caetano-Anollés paper, do yourself favor and take a look at it now. It's a fascinating read, no matter what side of the LUCA fence you're on.
Wednesday, May 21, 2014
The Pseudogene Hall of Fame
For "budget reasons" (supposedly), the Joint Genome Institute now requires all users to be registered and approved before they can use the (taxpayer-funded) https://img.jgi.doe.gov site. Fortunately, my registration was approved and I can use the excellent online genomics tools there, one of which produced the following table of Top Ten Organisms by Pseudogene Count.
Again: Don't expect the above links to work if you're not a registered JGI user. (I don't know if they will work for you or not.) This list was automagically generated by the Department of Energy's Joint Genomes Institute and I thought you might get the same kick out of it that I got. It's eye-opening to see the ratio of pseudogenes to "normal" genes in these organisms. Isn't it?
Taking the top three spots are mouse, rat, and humans. (Note: These counts should be taken with a bit of caution, as some have estimated the number of pseudogenes in the human genome to be much higher than the 7,186 shown here.) All the other spots in the chart except Arabidopsis (which is a leafy plant) are bacteria. The leprosy bacterium, which I've written about before, comes in tenth place.
If you're not familiar with the concept of pseudogenes, you might want to look at this post. Basically we're talking about genes that are thought to be disabled and no longer functional in the normal sense, although they may well be functional in some as-yet-unappreciated sense. (Otherwise, evolutionary theory says they should have been eliminated from most genomes eons ago.)
Personally, I believe pseudogenes are as much a feature of DNA as regular genes; certainly in higher life forms, they occur in great numbers. The vast majority of bacterial genomes in public databases are shown as having no pseudogenes. I find that (how shall I say?) not at all credible. Some day it will be obvious that almost every genome harbors pseudogenes; we simply lack smart enough software to detect them all right now.
Organism
|
Genes
|
Pseudogenes
|
|---|---|---|
60745
|
11398
| |
38115
|
9178
| |
38612
|
7186
| |
6325
|
4011
| |
31392
|
3818
| |
5379
|
1670
| |
4760
|
1589
| |
5775
|
1523
| |
5016
|
1507
| |
2750
|
1086
|
Again: Don't expect the above links to work if you're not a registered JGI user. (I don't know if they will work for you or not.) This list was automagically generated by the Department of Energy's Joint Genomes Institute and I thought you might get the same kick out of it that I got. It's eye-opening to see the ratio of pseudogenes to "normal" genes in these organisms. Isn't it?
Taking the top three spots are mouse, rat, and humans. (Note: These counts should be taken with a bit of caution, as some have estimated the number of pseudogenes in the human genome to be much higher than the 7,186 shown here.) All the other spots in the chart except Arabidopsis (which is a leafy plant) are bacteria. The leprosy bacterium, which I've written about before, comes in tenth place.
If you're not familiar with the concept of pseudogenes, you might want to look at this post. Basically we're talking about genes that are thought to be disabled and no longer functional in the normal sense, although they may well be functional in some as-yet-unappreciated sense. (Otherwise, evolutionary theory says they should have been eliminated from most genomes eons ago.)
Personally, I believe pseudogenes are as much a feature of DNA as regular genes; certainly in higher life forms, they occur in great numbers. The vast majority of bacterial genomes in public databases are shown as having no pseudogenes. I find that (how shall I say?) not at all credible. Some day it will be obvious that almost every genome harbors pseudogenes; we simply lack smart enough software to detect them all right now.
Subscribe to:
Posts (Atom)









