Showing posts with label phage. Show all posts
Showing posts with label phage. Show all posts

Friday, May 16, 2014

Evolution of Viral Genes vs. Host Genes

Some viral genes show evidence of having an ancient origin, predating the divergence of host organisms into different species. The enzyme thymidine kinase (TK), for example, has followed a certain set of evolutionary paths in enteric bacteria: the E. coli version differs slightly from the Salmonella version, which differs from the Yersinia version or the Shigella version, but they all show show evidence of common descent from ancestors that included Vibrio and Proteus. By contrast, the bacteriophages (viruses that attack enteric bacteria) have their own thymidine kinases that followed different evolutionary trajectories, resulting in genes with quite different G+C contents. The thymdine kinase in today's T4 phage shows a closer phylogenetic relation to the TK of Rhizobium and Agrobacterium than to E. coli or Salmonella. I presented a phylogenetic tree to this effect in an earlier post.

But sometimes, viral genes follow host-gene evolution more closely. The thymidine kinases of eukaryotic organisms and their viruses provided a good example of this. Consider the following tree developed from thymidine kinase genes of Guinea pig (Cavia porcellus), bull (Bos taurus), Cowpox and Swinepox viruses, amoeba species (Entamoeba), amoeba Mimivirus, two algal viruses, and finally, two strains of the alga Micromonas.

Thymidine kinase genes for cowpox and swinepox viruses occur in the same sub-branch with bull (Bos taurus) and Guinea pig (Cavia porcellus) genes; see the top four lines. Likewise, the Mimivirus TK gene is not far from Entamoeba. and the TK genes for algal viruses cluster near the TK genes of the alga Micromonas. Branch confidence is high (after 500 bootstraps, no nodes separate with less than 67% confidence). The 0.1 scale marker (lower left) represents substitutions per site, as represented by leaf-node depth. Tree generated using Mega6 freeware.
If you're not familiar with interpreting these trees, node depth (horizontal line length) is proportional to the fraction of amino acid substitutions per site, with a line length of about a centimeter representing 10 substitutions per 100 amino acids (note the 0.1 marker, lower left). The numbers at the branch points (100, 91, 86, etc.) represent the confidence that nodes to the right were located in the proper branches. These numbers are perecntages, representing the number of bootstrap trials (out of 100) that resulted in no "node-jumping." (Bootstrap testing is a way of introducing systematic noise into sequences to try to "trick them" into jumping to a new spot in the tree. If a little noise makes a node switch locations, the original location is suspect.) Every node in this tree was tested with 500 bootstrap tests. Overall, we can be fairly certain the node locations are correct.

What does the tree mean? Unlike the situation with bacterial and bacteriophage kinases (see earlier post), the various thymidine kinases of the DNA viruses shown here do tend to evolve in parallel with host equivalents. However, this doesn't automatically mean the viral genes originated with host genes, because (for example) the amoeba version of thymidine kinase shares only 39% amino-acid sequence identity with the Mimivirus version. That's a lot of divergence. It means the viral version of the gene (no doubt highly optimized for viral needs) could have come from a very-long-ago ancestor of the present-day host. Likewise, the Bos taurus version of the TK gene has only 69% similarity with the cowpox version. By comparison, the bovine gene is 88% similar to the human gene. The cowpox thymidine kinase gene is further away (evolutionarily) from the host gene than human TK is from rainbow trout TK.

It's prudent to conclude that viral genes are not simply orthologues of host-gene counterparts; they certainly don't represent recent gene transfer events (because there's way too much sequence divergence, even accounting for faster evolution in DNA viruses than in host DNA). The viral genes come from "somewhere else," probably a primordial past. Any resemblance to present-day host genes is largely incidental. 

The great majority of so-called "virus hallmark genes" (involving things like capsid proteins) have no counterpart at all in modern host cells, with very rare exceptions. Large DNA viruses may have gotten their start a long, long time ago, possibly in the pre-cellular world, where communities of free-ranging genes (some "selfish," some less so) coexisted in a common broth, with the only "cells" being micro-compartments in ocean-bed minerals.

RNA viruses are an entirely different matter. It's well known that RNA viruses evolve thousands of times faster than other viruses and often evolve in close harmony with hosts. Even so, faster evolution doesn't mean RNA virus genes came from host cells. RNA-virus genes, too, are probably of ancient provenance, maybe predating cellular life. As Koonin et al. said in 2006:
The existence of several genes that are central to virus replication and structure, are shared by a broad variety of viruses but are missing from cellular genomes (virus hallmark genes) suggests the model of an ancient virus world, a flow of virus-specific genes that went uninterrupted from the precellular stage of life's evolution to this day. This concept is tightly linked to two key conjectures on evolution of cells: existence of a complex, precellular, compartmentalized but extensively mixing and recombining pool of genes, and origin of the eukaryotic cell by archaeo-bacterial fusion. The virus world concept and these models of major transitions in the evolution of cells provide complementary pieces of an emerging coherent picture of life's history.

Wednesday, May 14, 2014

Where Do Virus Genes Come From?

There's a memorable moment in War of the Worlds (Spielberg version) when the Tom Cruise character tries to explain to his son, while driving a minivan, that they have to leave town immediately because evil, marauding machines "from somewhere else" are on the rampage. Robbie (the son) says: "What, you mean, like, from Europe?"

That's the image that comes to mind when I try to explain where virus genes come from. They don't come from the mother ship (the host). Oh sure, in some cases they clearly do derive from host genes. But in most cases, they clearly don't. The overwhelming majority of viral genes have no counterparts in host cells, and even for those that do, the genes in question are rarely true host orthologues.

"No, Robbie, not like Europe."
Where do they come from, then?

Short answer: Somewhere else.

Using a program like the excellent (and free) Mega6 ("Molecular Evolutionary Genetic Analysis") you can easily create phylogenetic trees from genetic data (FASTA sequences, protein or DNA), and when you do this for genes that occur in both viruses and host cells, you can see how they separate in phylo-space.

Example: I decided to look at the gene for thymidine kinase, which occurs in most living things plus a certain number of undead, maybe not-quite-living (in the usual sense) things known as viruses. Thymine (T) is, of course, an essential ingredient of DNA. Thymidine kinase converts thymidine (thymine bonded to deoxyribose, on the left, below) to the phosphorylated form (TMP, right) so it can participate in DNA synthesis.  

It's important to note that thymidine (above, left) is not a synthesis product but a breakdown product. The biosynthetic pathway for TMP actually starts with dUMP, which is converted to TMP via thymidylate synthetase, an entirely different enzyme. Where does thymidine come from, then, and why does a cell need thymidine kinase? The answer is that thymidine is a breakdown product of DNA. When you eat plant or animals cells, the DNA in those cells gets broken down, and thymidine is one of the breakdown products. For cells, thymidine is a valuable product to have, so thymidine kinase recovers free thymidine via the reaction shown above. This is a scavenging pathway, in other words. Most organisms have it, but some don't (e.g., Pseudomonas lacks it).

When a virus attacks a cell, there's a lot of DNA turnover as the virus prepares to manufacture its own DNA, so thymidine kinase is a handy enzyme to have around if you're a virus. Many viruses, as a result, bring their own copy of the TK enzyme gene. But is the viral TK gene derived (in some kind of ancestral way) from the host cell's own TK gene? Not necessarily.

When you gather up the amino acid sequence data for a bunch of TK genes from bacteria and the phages (viruses) associated with them, you find that, phylogenetically speaking, the phage/viral TK genes are not very similar to the host genes.

Thymidine kinase phylo tree for enteric bacteria and their phages. (Click to enlarge.) Note that the phage genes (top cluster) segregate from the host genes in the lowermost branch. Interestingly, TK genes of various alphaproteobacteria of the Rhizobiales class cluster near the phage genes (much nearer than the enteric bacteria). This suggests a primordial origin of the phage genes rather than recent acquisition of TK from host cells.

In the above tree, phage (viral) genes cluster at the top. Host-cell kinases cluster at the bottom. The two clusters may be related to a distant ancestor (not shown), but one thing is certain: the phage versions of this gene are not simply a slight modification of the host gene. We know that's true because, remarkably, the thyK genes of certain alphaproteobacteria (Agrobacterium and its relatives) cluster with the phage genes, even though the bacteriophages are adapted to E. coli and Salmonella (and closely related enterics). In theory, the phage genes should cluster with the enteric bacteria, not with Agrobacterium and Rhizobium.

Where do the phage genes come from? Some have speculated that these phages originated with escaped bacterial secretion-system cassettes. That may well be, but the escape event had to have occurred many hundreds of millions of years ago. The T4 thymidine kinase gene has only 62% sequence homology with the E. coli gene. T4's version of the gene has a G+C content of 34%, with GC3 (third codon base) of just 19%. The E. coli version of the gene has overall G+C of 42% and GC3 of 36%. While it's generally conceded that evolution of viral genes occurs faster than host genes (particularly for RNA viruses and single-stranded DNA viruses), DNA viruses like T4 can't evolve outside of the host cell, as far as we know, and although DNA viruses may evolve faster than host DNA, they don't evolve thousands of times faster. (RNA viruses do evolve thousands of times faster, but that's an entirely different matter.)

To me, the above tree says that enteric phages diverged from the common ancestor of alphaproteobacteria and gammaproteobacteria (the latter include the enteric bacteria). We're talking hundreds of millions of years ago. (Note that E. coli and its closest relatives are thought to have diverged 140 million years ago.) A recent transfer of thymidine kinase genes from enteric bacteria to their phages is not credible. It's far more likely that an ancient precursor of today's T4 phage (and similar enteric phages) had the gene, and passed it down through the ages as the phages adapted to new hosts (first the alphaproteobacteria, then the phylogenetically newer enterics).

In tomorrow's post, I want to explain in detail how the above phylo tree was made and show you how you can make your own phylogenetic trees using the popular Mega6 program. If you've never used Mega6, you're in for a treat.  

Sunday, March 30, 2014

The primordial nature of phage genes

Recent ecological studies have shown that bacteriophage (viruses that attack bacteria) are numerically the most abundant biological entities on the planet. The estimated 1030 viruses (mostly phage) in the oceans, if stretched end to end, would span farther than the nearest 60 galaxies. These viruses are thought to cause the turnover, by virus-related death, of 20% of the ocean's biomass per day. In shotgun sequencing of marine samples, the majority of phage gene sequences are invariably found to be novel (not corresponding to any other known gene sequences). Hence, the bulk of genetic diversity on the planet may well be tied up in viral/phage "dark matter."

One of the most-studied bacteriophage classes is the so-called "T-even" (T2, T4, T6, etc.) class of phages, of which the poster child, arguably, is T4. These are phages that attack enteric bacteria (E. coli and its relatives), hence are commonly found in sewage.
T4 phage morphology.

T4 is interesting from a number of standpoints, not least of which is its distinctive head/tail morphology (see diagram). In phylogenetic studies, T4 typically shows up in basal positions on trees, meaning it is presumably ancient. Increasingly, viruses and phages are considered to be of primordial origin, possibly predating cellular life. Certainly, any theory on the origin of life has to come to grips with the fact that the major biomolecules (proteins, nucleic acids, lipids) had to exist, in some form, prior to the appearance of the first cell. Some experts suggest nucleic acids and proteins may have interacted with each other in a so-called Virus World scenario (see the excellent paper by Koonin) wherein microscopic hydrothermal pore systems (in mineral formations at the ocean floor) provided for sequestration of prebiotic processes in physical compartments that could be invaded by "selfish replicators."

Primordial interaction of proteins with nucleic acids (and their precursors) presumably gave rise to a number of artifacts that survive today, such as ribosomes (which contain over 50 small proteins in tight association with RNA), tRNA (RNA covalently bound to an amino acid),  adenine-containing cofactors (e.g. SAMe, NADPH), and viruses (capsid and other proteins bound to RNA or DNA). Conceptually, one can think of protein/nucleic-acid complexes as having diverged, at the Darwinian Threshold, along two lines: toward ribosomal life, or toward the viral world.
Bacterial cell covered with T4 virions.

The genes for certain viral capsid proteins (with colorful names like Jelly Roll Capsid) are among a number of "viral hallmark genes" that show no homology to any genes from the cellular world. Presumably, some of these genes are of truly primordial ancestry. We have a valuable clue to the origin of at least some of these genes in the case of phage T4. A number of tantalizing reports from the 1970s (see here, here, and here) suggest that the enzymes dihydrofolate reductase and thymidylate synthase (both encoded by T4 DNA) are, in fact, components of the virion baseplate and/or tail structure of T4. Hence, at least in some cases, it's conceivable that virion structural proteins began as enzymes.

What's particularly intriguing about the T4 enzymes is that T4's thymidylate synthase (ThyA), which is phylogenetically ancient, is encoded in the phage DNA immediately downstream of the gene for dihydrofolate reductase, with no intervening "junk DNA." Why is this significant? In many organisms (as I explained in an earlier post), these two enzymes occur in a single large bifunctional enzyme that's proposed to be the result of a gene fusion event. In organisms that have the double enzyme, the reductase occurs at the beginning (the N-proximal end) of the protein.

Just for fun, I took the protein sequence for T4 dihydrofolate reductase and fused it (in Notepad) with the sequence for T4 thymidylate synthase, then did a BLAST search of the fusion sequence against all the protein sequences at UniProt.org. The naturally occurring bifunctional ThyA/dihydrofolate reductase enzymes from peach, balsam, rice, castor bean, and clementine (Citrus clementina) all showed up as hits, with E-values of 10-67 or better.

This doesn't prove that the bifunctional enzymes of the peach, etc. came from T4 phage, of course, but it is consistent with the general idea that the bifunctional ThyA/dihydrofolate reductases of algae and protists could (at least in theory) have started out as phage gene fusion products.

Let's put it this way: Weirder things have been known to happen.

Friday, April 26, 2013

Science on the Desktop

For decades, I've been hoping I'd live long enough to see a day when serious science could be done on the desktop by dedicated amateurs. Amateur astronomers know what I'm talking about. You can't do much particle physics on the desktop, and there are no affordable desktop electron microscopes (yet), but if comparative genomics is your thing? Get ready to rock and roll, my friend.

Over the weekend I discovered http://genomevolution.org and promptly went nuts. Let me take you on a tour of what's possible.

First I should explain that my background is in microbiology, and I've always had a soft spot in my heart (not literally) for organisms with ultra-tiny genomes: things like Chlamydia trachomatis, the sexually transmitted parasite. It's technically a bacterium, but you can't grow it in a dish. It requires a host cell in which to live.

It turns out there are many of these itty-bitty obligate endosymbionts (at least a dozen major families are known), and because of their small size and obligate intracellular lifestyle, they have a lot in common with mitochondria. Which is to say, like mitochondria, they're about a micron in size, they divide on their own, they have circular DNA, and they provide services to the host in exchange for living quarters.

When you look at one of these little creatures under the microscope (whether it's Chlamydia or Ehrlichia or Anaplasma or what have you), you see pretty much the same thing. (See photo.) Namely, a tiny bacterium living in cytoplasm, mimicking a mitochondrion.

When Lynn Margulis wrote her classic 1967 paper suggesting that mitochondria were once tiny bacterial endosymbionts, it seemed laughable at the time, and her ideas were widely criticized (in fact her paper was "rejected by about fifteen journals," she once recalled). Now it's taught in school, of course. But we have a long way to go before we understand how mitochondria work. And we really, really need to know how they work, because for one thing, mitochondria seem to be deeply involved in orchestrating apoptosis (programmed cell death) and various kinds of signal transduction, and until we understand how all that works, we're going to be hindered in understanding cancer.

When I discovered the tools at http://genomevolution.org, one of the first things I did, on a what-the-hell basis, was compare the genomes of two small endosymbionts, Wolbachia pipientis and Neorickettsia sennetsu. The former lives in insects; the latter, in flatworms that infect fish, bats, birds, horses, and probably lots else. Note that for a horse to get Potomac horse fever, first the Neorickettsia has to infect a tiny flatworm; then the flatworm has to be ingested by a dragonfly, caddisfly, or mayfly; then the horse has to eat (or maybe be bitten by, although only infection-by-ingestion has been demonstrated) the worm-infected fly. The parasite-of-a-parasite chain of events is not only fascinating in its own right, it suggests (to me) that parasites enable each other through shared strategies at the biochemical level, and I might as well spoil some suspense here by revealing that there's even yet another layer of parasitism (and biochemical enablement) going on in this picture, involving viruses. But we're getting ahead of ourselves.

I mentioned Wolbachia a second ago. Wolbachia is a fascinating little critter, because it's found in the reproductive tract of anywhere from 20% to 70% of all insects (plus an undetermined number of spiders, mites, crustaceans, and nematodes), but they don't cause disease, and in fact it appears many insects are unable to survive without them. Wolbachia are unusual in that the extracellular phase of their lifecycle (the part where they spread from one host to another) isn't known; no one has observed it. What's more (and this part is incredible), Wolbachia have adapted to a stem-cell niche: They live in the cells that give rise to insect egg cells. Thus, all newborn female progeny of an infected mother are infected, and all eggs pass on the Wolbachia. In this sense, the genetics of Wolbachia obey mitochondrial genetics (whereby the mother passes on the organelle and its genome).

I quickly found, via Sunday afternoon desktop genomics, that Wolbachia and Neorickettsia (and other endosymbionts: Anaplasma, Ehrlichia, etc.) have many genes in common—hundreds, in fact. And when I say "genes in common," I mean that the genes often show better-than-50% similarity in DNA base-pair matching.

It's important to put some context on this. These little organisms have DNA that encodes only 1,000 genes. (By comparison, E. coli has around 4,400 genes.) Endosymbionts lack genes for common metabolic pathways. They cannot biosynthesize amino acids, for example; instead they rely on the host to provide such nutrients ready-made. If 400 to 500 of an endosymbiont's 1,000 genes are shared across major endosymbiont families, that's a huge percentage. It suggests there's a set of core genes, numbering in the low hundreds, that encapsulate the basic "strategy" of endosymbiosis.

A little more context: Mitochondria have their own DNA and look a lot like endosymbionts. But here's the thing: Mitochondrial DNA is tiny (only about 15,000 base pairs, versus a million for an endosymbiont). It turns out, 97% of the "stuff" that makes up a mitochondrion is encoded in the nucleus of the host. If you include these nuclear genes, mitochondria actually rely on about 1,000 genes total, of which only 3% are in the organelle's DNA. Lynn Margulis would say that what happened is, the endosymbiont ancestor of today's mitochondrion originally had DNA of about a million base-pairs (1,000 genes), but some time after taking up residency in the host cell, the invader's DNA mostly migrated to the host nucleus.

Why did symbiont-to-host DNA migration stop at 97%? Why not 100%? If we look at that 3%, we find genes coding for tRNA and bacterial ribosomes (specialized protein-making machinery) plus genes for enormous, complex transmembrane enzyme systems: cytochrome c oxidase and NADH dehydrogenase. (The former is the endpoint of oxidative respiration; the latter the entry-point.) Obviously it must be advantageous for these genes to be proximal to the organelle.

But why even have an organelle (a physical compartment)? One might ask why it's necessary to have a mitochondrial parasite swimming around in the cytoplasm at all, when most of the genes are part of the host's DNA? The answer is, the stuff that goes on inside the confines of the mitochondrion needs to be contained, because it's violently toxic stuff involving superoxide radicals, redox reactions, "proton pumps," and Fenton chemistry (transition-metal peroxide reactions). A containment structure is definitely called for, to segregate this toxic chemistry from the rest of the cell.

We might ask how it is that the DNA of the protobacterial ancestor of today's mitochondria wound up in the host nucleus in the first place. Let's consider the possibilities. Protobacterial (symbiont) DNA may have transferred to the host all at once, or it might have migrated piecemeal, over time. Or both. Is it realistic that huge amounts of endosymbiont DNA could have migrated to the host nucleus all at once? Yes. It's been suggested that vacuolar phagocytosis drove invader DNA to the nucleus in a big gulp. Evidence? Wolbachia inhabits the vacuolar space.

But export of genes and gene products to the host might have occurred piecemeal as well. A little desktop exploration provides some clues. If you use GenomeView or any number of other online tools to explore the DNA of Wolbachia, several things pop out at you. First is that many Wolbachia genes are mitochondria-like: They encode for things like cytochrome c oxidase, cytochrome b, NADH dehydrogenase, succinyl-CoA synthetase, Fenton-chemistry enzymes, and a slew of oxidases and reductases (including a nitroreductase). Wolbachia is clearly engaged in providing what might be called redox-detox services for the host—the same value proposition that mitochondria offer. This makes sense, because if Wolbachia cells were a net drag on the respiratory potential of host-cell mitochondria (if they couldn't at least hold their own with respect to mitochondria), the host would die.

The second thing that jumps out at you when you look at the Wolbachia genome is the abundance of genes devoted to export processes: membrane proteins, permeases, type I, II, and IV secretion systems, ABC transporters, etc., plus at least 60 ankyrin-repeat-domain genes—all powerful evidence of specializations aimed at export of genes and gene products to the host. But the most stunning "smoking gun" of all is the presence, in Wolbachia DNA, of five reverse-transcriptase genes, plus genes for resolvases, recombinases, transposases, DNA polymerases, RNA polymerases, and phage integrases. In essence, there's a complete suite of retroviral machinery, designed for export of foreign DNA into host DNA.

An example of one of 113 phage-derived genes in Wolbachia (lower gene array). In this case, the gene matches a phage gene found in Candidatus hamiltonella (upper gene array). The two isoforms exhibit 59% DNA sequence similarity, despite widely differing GC ratios. See text for discussion.

But wait. There's more. The third thing that jumps straight in your face when you start looking at the Wolbachia genome is the presence of (are you ready?) no less than 113 genes for phage-related proteins, including major and minor capsid and HK97-style prohead proteins, plus tail proteins, baseplate, tail tube, tail tape-measure, and sheath proteins; late control gene D; phage DNA methylases; and so on. (For non-biologists: phage is the term for viruses that attack bacteria.)

In the above screenshot, I'm comparing Wolbachia DNA (lower strip) to DNA from another insect-infecting endosymbiont, Candidatus hamiltonella, which is known to contain an intact virus (phage) in its DNA. Many phage proteins in Wolbachia have corresponding matches in the Candidatus genome. In this case, we're looking at a gene (the gold-colored stretch pointed at by red arrows) that is 1440 nucleotides long, with a 59% sequence match across genomes. The match percentage is remarkably high given that the Candidatus version of this gene has a 51.7% GC content while the Wolbachia version has a 40.6% GC. Also, note that Wolbachia itself has an overall GC of 34.2%. The fact that Wolbachia's putative phage genes are significantly higher in GC content than Wolbachia's non-phage genes is good confirmation that the genes really are from phage.

It's 100% clear that viral DNA has made its way into the DNA of Wolbachia (either recently or long ago), and it's reasonable to hypothesize that Wolbachia has repurposed the retrovirus-like phage genes for packaging and exporting Wolbachia DNA to the host nucleus.

Okay, so maybe you have to be a biologist for any of this stuff to make your hairs stand on end. To me, it's a dream come true to be able to do this kind of detective work on a Sunday afternoon while sitting on the living-room couch, using nothing more than a decrepit five-year-old Dell laptop with a wireless connection. The notion that you can do comparative genomics and proteomics while watching an Ancient Aliens rerun on TV is (for me) totally cerebrum-blowing. It makes me wonder what's just around the corner.