Showing posts with label Fels-2. Show all posts
Showing posts with label Fels-2. Show all posts

Saturday, May 17, 2014

Evolution of Prophage Genes

Viruses have two modes of reproductive existence. In the familiar lytic cycle, the virus infects a cell, replicates itself until the cell bursts, and hundreds (or thousands) of virions are produced. But there is also a lysogenic mode of viral existence, in which the virus inserts a copy of its DNA in the host's own DNA. The viral DNA thus inserted becomes known as a prophage, which can remain dormant for long periods of time. The prophage can often be induced to enter a lytic cycle by exposure of cells to hydrogen peroxide or Mitomycin C. (Induction of phages in this fashion is thought to occur when a phage repressor protein is cleaved by recA after the latter is upregulated in the SOS response.)

Prophage genes are seen in a wide variety of bacteria (a 2008 paper estimated that over 60% of bacterial genomes contain prophage genes), and in fact human DNA is thought to contain at least 8% retroviral gene remnants. There's reason to suspect that certain large DNA animal viruses (such as herpes and vaccinia) have a lysogenic cycle. Certainly, viruses like varicella zoster (which can produce shingles many years after a person's initial infection) can remain dormant for decades before suddenly undergoing induction to a lytic phase.

Viruses that live an exclusively lytic lifecycle have relatively few opportunities to co-evolve with the host, because they spend little time in the host. Such a virus might spend years "hanging around" in the environment before encountering a host cell; then the lytic reproductive cycle may last only minutes or hours, and it's back to "hanging around" in the environment.

The situation is much different for a temperate virus (i.e., one that has a lysogenic cycle). A lysogenic virus essentially becomes an integral, first-class component of the host DNA and undergoes the same replication and repair processes that apply to host DNA. Accordingly, we should expect to see a much different pattern of evolution in the genes of lysogenic viruses (or prophages). And indeed we do.

The phylogenetic tree below was prepared using viral (phage) and bacterial genes for DNA adenine methylase (dam), an enzyme involved in DNA repair and replication. What's interesting about this gene is that many bacteria have their own (native) copies of this gene plus a prophage copy. And they differ, but not as much as, say, lytic-phage thymidine kinase versus native bacterial TK. (I showed phylo-trees for viral and bacterial TK enzymes in a prior post. If you'll recall, these enzymes differ so drastically that it's not at all clear that one derives from the other, ancestrally.)

DNA adenine methylase genes from three enteric bacteria and two phages (marked with asterisks). The top branch shows very close homology between prophage genes and their bacterial paralogues. The bottom branch shows that the native bacterial isoform of the enzyme is not as closely related to the prophage version(s).
With the dam genes, we see an interesting segregation pattern. There are two main branches to the phylo-tree. In the upper branch are the phage dam genes along with bacterial paralogues of these genes. The bottom branch shows how the non-paralogous (non-prophage) dam genes segregate.

To make these relationships clearer, here's a chart showing the overall G+C content as well as the GC3 (G+C content at codon base 3) for the various genes. The entries shaded in grey represent prophage genes. Notice that the G+C percentages are significantly lower for the prophage genes, but are higher than in free-living lytic-cycle phages (where GC3, in particular, is often less than 20%).


DNA adenine methylase genes for enteric bacteria and their temperate phages. Base-composition stats for prophage isoforms are shown in grey.
Organism Gene G+C GC3
Shigella sp. strain D9 ZP_05434596.1 49.60% 55.90%
E. coli EHW52521.1 49.60% 55.90%
S. enterica AAL22346.1 49.10% 54.50%
E. coli EHW55384.1 47.20% 47.90%
Salmonella phage RE-2010 YP_007003503.1 46.50% 46.20%
Shigella sp. strain D9 EGJ07993.1 46.20% 46.30%
S. enterica ETB92379.1 46.20% 43.00%
Fels-2 phage YP_001718754.1 46.20% 43.00%

If you compare the phylo tree shown further above with the phylo tree in my earlier post about thymidine kinase genes, you'll note that the prophage dam genes cluster very tightly with bacterial versions of these genes. That's because, as a fully integrated part of the genome, the prophage genes benefit from the host's DNA repairosome. They evolve gradually over long periods of time by the usual mechanisms. The genes are notably host-like because they're continuously repaired and groomed in the same manner as host DNA.

The takeaway here is: If you create a phylo-tree for a set of genes from hosts and viruses, and the genes cluster tightly with host versions, you're probably looking at the result of longterm lysogeny. On the other hand, if the virus genes do not cluster with host genes (as they usually don't!), that means you're looking at viruses that have a predominantly lytic mode of existence; viruses that probably got their genes from a far-distant ancestor of the modern-day host, if not from a primordial precellular precursor of some kind.

Sunday, May 04, 2014

A Tale of Nematodes, Scarabs, Germs, and Genes

Genes get around. Sometimes they go from organism to organism to organism. Take gene ECB_00841 of E. coli B, for example. This baby's been around the world.

Pristionchus pacificus (the handsome guy
on the right) is a nematode, magnified
here about 50 times.
By outward appearances, ECB_00841 is just another "hypothetical protein" gene (which is Genomic for "we don't know what the heck this thing does"), one of 793 genes of unknown function in E. coli B. But here's the funny thing (and boy is it odd): If you copy the DNA sequence of that gene and translate it into an amino acid sequence using the online translation tool at http://web.expasy.org/translate, you'll get six different translations of the sequence, based on the six possible reading frames in which you can parse DNA. One of the six translations, namely the reverse translation (Expasy.org calls it 3'5' Frame 1) looks like this:

SYRASGCIAAFQDASMLMRINHDPVRTHRSAALGNATIDHHFAVNARAVSLQYKLTGFHG
DVISGLNDH-FDAPDIPAPGGGFVFKPATVRMFCHAGVRRRRRWCELIRIDSGQRKGSLQ
IAAQTQQHHLLTFRWSPPCAGITRAQRQPADPVGFKVARFHPTKPVFPVHFGDYTYADQI
GDKAHDFG-LCVHRYMMTAENLFITSCKLWYRWYKLIKG-KTTNRWINECYESNKF-YI-
C-IFKHPGSGTCRMDFHYWSCYNFGDDSYKSTIQ

This is the amino acid sequence you get by reading the "wrong strand" of the DNA. Notice there's pink highlighting every time a sequence begins with the letter 'M' (because methionine is associated with the start codon of a gene), but the highlighting ends whenever a hyphen (representing a stop codon) is encountered. If this gene were translated in vivo as shown, it would be in four rather pathetic chunks. If this is truly the correct DNA strand, we're looking at a pseudogene.

But now comes the Real Magic. (I promise you, this part is amazing.)

Go ahead and Select and Copy everything in the above sequence from PAPGGG to YKSTIQ, then go to http://www.uniprot.org/?tab=blast and Paste the sequence into the BLAST field, but before you click the Blast button, be sure to remove the hyphens from the amino acid sequence, as otherwise you'll get an immediate error.

After 30 seconds or so, the BLAST search will come back with a short list of hits. At the very top, with an Identity score of 100%, is an Uncharacterized Protein belonging to Pristionchus pacificus ("parasitic nematode").

Yes, the backwards-translated E. coli gene is actually a forward gene in a worm. It's not a fake hit or a ruse. This is a genuine gene, wrongly annotated as to DNA strand in E. coli, but definitely existing in both a bacterium and a worm. The fact that the amino-acid sequence identity is 100% (not 90%, not 98%, but 100%) is striking confirmation that the gene really exists and is conserved in both organisms.

You're probably wondering how an E. coli gene gets into a nematode's DNA in the first place. I'm glad you asked, because the answer is fascinating.

The gene, it turns out, originates neither with E. coli nor with the nematode. If you search online databases, you'll eventually find that the gene is actually a baseplate assembly protein from the enterobacterial phage Fels-2.

Fels-2 bacteriophage (virus).
The gene occurs in a part of the E. coli genome that happens to contain a large cluster of prophage genes. Recall that viruses can coexist with hosts in two ways: the familiar lytic cycle (where the virus takes over the host cell, eventually exploding it to release thousands of new virions), or the stealth-mode lysogenic cycle, wherein a viral genome integrates itself into a host genome, where it can remain dormant for anywhere from a few hours to all eternity, depending.

At some point in the past, Fels-2 integrated itself into the E. coli B genome, where parts of it have remained for probably millions of years, although (intriguingly) not a single amino acid has changed between the nematode version and the E. coli version.

"But how did it get into the worm?" you're asking. Well, in the wild, P. pacifica likes to hang out with scarabs. Nematodes are often mistakenly identified as parasites of beetles. The truth is, they like to feed on dead beetles, but they do not attack beetles directly. Rather, the so-called dauer larvae of the worm (a durable, environmentally hardened larval form) bring with them bacterial stealth payloads, some of which are toxic to the beetle and can kill it. The dauer larvae patiently wait for the beetle to die so they can begin feeding.
Beetles are in constant contact
with dung and enteric bacteria.

Of course, scarabs are dung-mongers, and as such, they're no strangers to the likes of E. coli, but the Xenorhabdus bacteria carried by nematodes can be deadly to the beetle. As it happens, E. coli and Xenorhabdus are both enteric bacteria, and both carry the ECB_00841 gene. In fact, some version of this gene exists in a wide variety of enteric bacteria, including members of Salmonella, Yersinia, Klebsiella, and other genera. It could be that each bacterium acquired the phage gene separately, through individual lysogeny events, but a more parsimonious view is that the common ancestor of these bacteria acquired the first copy, many millions of years ago, and passed it down through the ages.

Somehow, at some point, the gene for the phage baseplate protein made its way into a nematode's reproductive cells. Nematodes can feed on microorganisms, and it's possible a nematode engulfed an infected bacterium (infected with the Fels-2 phage), a bacterium that then underwent lysis inside the nematode host cell, releasing thousands of virions. Fels-2 brings with it its own recombinases and integrases, enzymes that would have facilitated transfer of the phage DNA to the nematode. By chance, the baseplate gene stuck.

Why a baseplate gene? Who knows. Phage proteins are often multifunctional, and what appears to be nothing more than a structural protein (a baseplate protein) can sometimes turn out to play other roles. No doubt, the so-called baseplate protein plays some kind of useful role for Pristionchus pacificus and for the various bacteria in which the gene exists today. Otherwise, according to evolutionary theory, the gene would have been lost eons ago.

One thing seems likely: The gene has probably been around a very long time, probably as long as dung-pushing beetles (and the nematodes that eat them when they die) have been pushing balls of dung. And that's a long, long time indeed.

If you enjoyed this post, please share the link with a friend. Thanks!