Showing posts with label Chlorella. Show all posts
Showing posts with label Chlorella. Show all posts

Saturday, November 01, 2014

A Virus Where It Shouldn't Be

Words like "bizarre" and "unexpected" hardly suffice when describing the results published this week in PNAS (one of the most respected journals in all of science) by Robert H. Yolken and 17 coworkers from Johns Hopkins, University of Nebraska, and Baltimore's Sheppard Pratt Health System, who wrote about finding a peculiar virus in the noses of a number of individuals. The virus they found (through DNA metagenome analysis) was something called ATCV-1, a large DNA virus normally associated with the freshwater alga Chlorella.

ATCV-1 virions attached to a Chlorella cell.
Just to bring you up-to-date quickly: The scientists in question stumbled onto the fact that DNA from the ATCV-1 virus appears to exist in the noses and throats of a surprising fraction (43%) of randomly chosen individuals. The study population of 92 people is not large, however, and we should not be too quick to jump to conclusions. Nevertheless, the mere finding of this virus's DNA in that many people's noses is shocking, because this is a virus that, until now, was thought to occur only in algae, not higher life forms (and certainly not in humans).

To understand how shocking this is, you have to realize that viruses are generally extremely highly adapted to specific hosts. A virus that attacks tobacco plants only attacks tobacco plants. A virus that infects bacteria only infects bacteria, and generally only a specific species of bacterium. Viruses coevolve very closely with their hosts, developing extremely intricate, fine-tuned adaptations to a specific organism. For a virus to cross species lines (let alone Kingdoms) is unusual, to say the least. And thank goodness! Otherwise, you might come down with pox from an insect bite, or get any number of deadly diseases from eating ordinary foods.

As it turns out, ATCV-1 (Acanthocystis turfacea Chlorella virus) has appeared in this blog once before, back in March, in a post I did called "A Virus, a Worm, and a Louse Walk into a Bar." The point of that post (ironically) was to show that while most virus genes tend to have a great deal of DNA homology with their host counterparts, the genes of certain algae viruses actually appear to have greater homology with genes in the human body louse. Which perhaps should have been a tipoff, of sorts, to the Yolken et al. findings. Certainly, it adds more color to the story, knowing (as we do) that phycodnavirus DNA may have have crossed species lines before. ("Phycodnavirus" is the name scientists have given to the family of large DNA viruses that infect algae.)

What do we know about ATCV-1? It's fairly large, as viruses go, with a genome of 288,047 base pairs, encompasssing as many as 860 genes. The fact that the genome has been fully sequenced doesn't mean we know how it works, though. In fact, we don't know what most of the 860 genes do. Some of the genes are (as in many king-size viruses) devoted to DNA-synthesis functions of a type commonly associated with nuclear-expressed genes in the host, meaning that the virus appears well adapted to take over the nucleus of a cell. This is not always the case; some viruses are adapted to thrive in the cytoplasm, away from the nucleus.

Intriguingly, the Yolken group conducted tests of cognitive function among enrollees in its study and found that there was a statistically significant reduction in cognitive ability (on the particular measures tested) in the group of people that tested positive for ATCV-1. To determine if this was just a fluke, researchers evaluated the effects of the virus on mice. As it happens, inoculated rodents showed deficits in recognition memory and attention while navigating mazes. In other words, exposure to the virus is associated with cognitive deficits in both humans and an animal model.

This is remarkable, since (if true) it would suggest that the virus traveled through the blood (of humans and mice), crossed the blood-brain barrier, gained entry to brain cells, and expressed its DNA inside brain cells (a type of cell the virus would presumably never encounter in its normal habitat of freshwater marshes and lakes). It all sounds (and is) very unlikely. But there you are: PNAS is one of the most esteemed science journals in the world. They published the result. It's as real as can be.

Of course, the work needs to be replicated and extended. And I'm sure it will be.

And I think what we'll find is that phycodnaviruses have been crossing species boundaries for some time. If you go back to my original blog about this, you'll see an example of a particular gene, encoding a ribonucleoside reductase, in another Chlorella virus (called PBCV-1), which shows greater homology to the ribonucleoside reductase of the human body louse than to the same gene in the virus's "normal" host, Chlorella. (The homology between PBCV-1 and louse reductases is 53%, versus 48% for virus and Chlorella. These are protein-sequence homology numbers.)

If we check the ATCV-1 reductase for homology with reductases in other organisms, we find that the highest scoring non-viral match (53%) occurs not with Chlorella (the virus's host), but with Lichtheimia corymbifera, a fungus found in soil and decaying plant matter. Interestingly, this fungus is known to cause pulmonary, CNS, and rhinocerebral infection in animals and humans. (But there is no known association whatsoever between ATCV-1 and the fungus.) One wonders whether the fungus has learned a few tricks from marine viruses; or perhaps vice versa?

If you're wondering how the people in the Yolken et al. study could have come in contact with ATCV-1, maybe the answer is as simple as taking a drink of water from a mountain stream, or drinking unfiltered water from a well, or swimming in a lake, or perhaps just walking in a marsh.

Suffice it to say, a good deal more remains to be learned regarding the ecology and life cycles of the phycodnaviruses, and cross-species infection/transfection generally. I feel certain the recent results of Yolken et al. will stimulate a great amount of much-needed followup research.

In the meantime, would you grab me a bottled water?

Please share this story on your favorite social channel, if you enjoyed it. Thanks!

Tuesday, May 20, 2014

More Secrets of the Virus World

It's generally conceded that viruses evolve more rapidly than host cells, but the rates vary tremendously depending on the type of virus. Generally, large DNA viruses that infect algae (the phycodnavirus family) are considered to have some of the slowest rates of change, whereas the fastest-to-change viruses tend to be small RNA viruses that infect animal cells (e.g., HIV). In terms of substitutions per nucleotide per cell infection (s/n/c), one recent study found rates of 10−8 to 10−6 s/n/c for DNA viruses and 10−6 to 10−4 s/n/c for RNA viruses, which means the fastest-mutating viruses change 10,000 times faster than the slowest-mutating viruses.

Given the ultra-rapid rate of change of RNA viruses and their generally impressive level of adaptation to host-cell environments, one might expect a virus like HIV-2 to show a codon usage bias similar to that of the host. And that's approximately true.

HIV-2 codon usage (left), in DNA format (T for U), versus overall human-cell codon usage (right).

The above graph shows codon usage for HIV-2 on the left and codon usage for human cells on the right. (HIV is an RNA virus, but codons are shown here in DNA format, with T in place of U.) R-squared/adjusted comes to 0.2204, so we can't very well say confidently that the codon values are highly correlated. But if you look at the smaller bars (not the "peaky" ones), they tend to taper down on the left, just as on the right.

It might be instructive to go from one of the fastest-changing viruses in the biosphere (HIV) to one of the slowest, and see how its codon usage compares to that of its host. This time, we're looking at the large DNA virus known as PBCV-1 (left) versus its Chlorella host (an alga, right):

Codon usage in Paramecium bursaria Chlorella virus 1 (PBCV-1), left, and Chlorella variabilis strain NC64A, right.
These two data sets are not only not correlated, they appear to be anticorrelated, which is quite unexpected. Bear in mind, PBCV-1 is relatively large, with a genome of 330,601 base pairs encoding hundreds of proteins (and ten tRNAs). Thus the pattern shown here isn't likely to be random noise. Note that PBCV-1 has a genomic G+C content of 40%, versus 61% for the host, which is a pretty sizable separation. It's almost as if PBCV-1 has spent part of its life coexisting with an entirely different host.

Which brings me to the final and most intriguing (I might even say shocking) graphic, which compares codon usage in PBCV-1 virus with codon usage in Chlorella's own host, Paramecium.

Codon usage in PBCV-1 virus (left) and Paramecium (right).
Recall that when it is not free-living on its own, the tiny unicellular Chlorella alga has an endosymbiotic relationship with the comparatively much larger unicellular ciliate protist, Paramecium. That is to say, Chlorella can live inside Paramecium. Chlorella allows Paramecium to thrive in high-sunlight/low-nutrient conditions, whereas Paramecium, in return, gives the non-motile Chlorella free transportation and protection against viruses. (PBCV-1 can infect free-living Chlorella, but does not infect Chlorella living inside Paramecium.) As far as I know, no one has ever reported that PBCV-1 virus can infect Paramecium. Supposedly, it infects only free-living Chlorella And yet, we find that the pattern of codon usage in PBCV-1 is very strongly correlated with the pattern of codon usage in Paramecium. (R-squared/adjusted: 0.527.)

Paramecium filled with Chlorella cells.
This chart is a real shocker from a couple of standpoints. First, as I say, PBCV-1 virus is not known to infect Paramecium. And yet codon usage patterns in the virus are much more closely aligned to Paramecium's patterns than to Chlorella's. Notice that AAA is the No. 1 most-used codon in PBCV-1 as well as Paramecium. Seven of Paramecium's top ten codons are in PBCV-1's top ten.

Secondly, Paramecium doesn't use the standard genetic code! It uses the Ciliate Code (Translation Table 6), in which TAA and TAG encode glutamine instead of serving as stop codons. (TGA is the one and only stop codon in Table 6.) If Paramecium used the standard genetic code, the alignment of the two organisms would be even stronger.

Also interesting is that PBCV-1 and Paramecium are quite far apart in G+C content (the former is 40%, the latter is 28%).

Perhaps at some point in its past, PBCV-1 had a wider host range, one that included Paramecium. It's possible that even today, it has hosts other than Chlorella that have yet to be observed experimentally. Certainly, the pattern of codon usage is consistent with such an idea.

Saturday, March 29, 2014

Virus genes don't always come from the host

Usually, when a virus contains a certain kind of gene, and the host contains the same kind of gene, it's assumed the virus got its copy from the host. This is not a terribly safe assumption, however. In some cases it's demonstrably wrong.

For a striking example of how wrong this assumption can be, you need look no further than the tiny Chlorella alga that can be found living symbiotically inside the fresh-water ciliate Paramecium. When it's not living inside Paramecium, Chlorella is subject to infection by PBCV-1 (the Paramecium bursaria Chlorella virus).

Tiny green Chlorella cells can be seen here growing
inside Paramecium. The full Paramecium cell is
shown in the inset at lower left. (Photo by Charles Krebs.)
Both Chlorella and PBCV-1 have a gene for an enzyme called thymidylate synthase, which is the enzyme that produces thymidine monophosphate (dTMP, or just TMP), a precursor molecule for making DNA. Ordinarily, one would assume that the virus picked up the gene for this enzyme from its host at some point in the past. But there's a problem.

The only thymidylate synthase gene in Chlorella's genome codes for a protein with 508 amino acids. The PBCV-1 virus version of this gene codes for a much shorter protein with only 216 amino acids. It turns out there's a perfectly good explanation for the size difference. Like other small algae (such as Micromonas, Ostreococcus, and Bathycoccus) and certain protozoans as well, Chlorella has evolved a bifunctional enzyme. In Chlorella, the same enzyme acts as both a thymidylate synthase and as a dihydrofolate reductase. In most higher organisms, two different enzymes carry out these functions. Organisms that have the dual-function enzyme are presumed to have developed this capability through a gene fusion event sometime in the (most likely distant) past.

It turns out the PBCV-1 virus synthase not only isn't bifunctional, it carries out its thymidylate reaction by an entirely different mechanism than that used in the host enzyme. The host enzyme employs folate (but no flavins) as a cofactor, whereas PBCV-1 is strictly dependent on flavin adenine dinucleotide (FAD), as verified experimentally by Graziani et al. in 2006. We now know that many bacteria use the FAD version of this enzyme (often called ThyX, as disintguished from ThyA, the folate-only enzyme). And the FAD users all have relatively small thymidylate synthases, of about 200 to 300 amino acids.

The above scenario isn't exclusive to Chlorella and PBCV-1. It turns out, certain other small algae (Micromonas, Ostreococcus, and Bathycoccus; all happen to be salt-water algaae) have a bifunctional thymidylate kinase, yet they are subject to infection by viruses that use the much smaller, mono-functional flavin-binding enzyme.

In all these cases, the virus uses an entirely different style of enzyme than the host to carry out TMP production. There is essentially zero chance that the virus derived its enzyme from the host (or vice versa), because the reaction mechanisms of ThyA and ThyX are radically different. (For more detail on this, see the excellent review article at http://www.ncbi.nlm.nih.gov/books/NBK6401/.) These aren't orthologues; these aren't paralogues; these are entirely different enzymes.

So where did the virus get its thymidylate synthase from, if not the host?

If you take the protein sequence for the PBCV-1 thymidylate synthase and run a BLAST search at UniProt.org, the best non-viral hits (in the range of 58% identities, 77% similarities, E-value 10-69) are for the thymidylate synthases of Prochlorococcus marinus and other cyanobacteria, with cyanophages also scoring high. This makes a great deal of sense, because the photosynthetic Prochlorococcus and its relatives are thought to be some of the most ancient bacteria on earth (possibly going back 3.8 billion years). They're thought to be the ancestors of chloroplasts. At one point, they were almost certainly the predominant life form in the oceans. Since phycodnaviruses (of which PBCV-1 is a member) are thought to be quite ancient, it's entirely possible they got their thymidylate synthase from cyanobacteria. That's certainly what the protein-sequence evidence suggests.

I'll go with the evidence.

Wednesday, March 26, 2014

A virus, a worm, and a louse walk into a bar

Since large DNA viruses are in the business of making large amounts of DNA, it shouldn't come as a surprise that many of them carry a gene for ribonucleoside diphosphate reductase, the enzyme that allows deoxy-bases (dADP, dCDP, etc.) to be created for use in deoxyribonucleic acid (DNA). The host organism, of course, has its own reductases for this purpose. But you have to imagine that when a giant DNA virus comes barging into a host cell and begins its crash program of digesting host nucleic acids into monomers (free nucleotides), the virus has a huge need to convert those monomers, quickly, into the deoxy form.

So I wasn't totally surprised to find that the genome for PBCV-1, the virus that infects Chlorella algae, contains genes for RNDR (ribonucleoside diphosphate reductase). What's surprising is that the virus brings not one, but two such genes. One gene encodes a short protein (about 370 amino acids); another gene encodes a protein with 771 amino acids. In the case of Paramecium bursaria Chlorella virus NY2A (PBCV-NY2A), which is essentially a variant of PBCV-1, there's actually a third gene, for a protein having 1,103 amino acids.

Why so many genes?

It turns out there are three major types of RNDR enzyme in living organisms, and a given organism can have more than one type. There's an aerobic enzyme (class I) that uses a tyrosine oxygen for radical generation. There's a larger (~1200 AA) class II enzyme that requires adenosylcobalamin (B12) as a coenzyme. And there's an anaerobic class III enzyme that relies on S-adenosylmethionine (SAMe) as a cofactor. Based on the relative sizes of these various enzymes, it appears the PBCV-NY2A virus may be harboring all three. However, most phyocodnaviruses infecting algae seem to have class I and class III reductases, but not the bigger class II.

Human body louse.
The more-or-less standard assumption, when a virus has an enzyme that the host also has, is that the virus obtained its copy of the gene from the host (at some point in the distant or not-so-distant past). That assumption may have to be revisited for PBCV-1's class III reductase. When you do a protein alignment of the viral reductase against the sequence for the host alga's reductase, you expect to see a lot of sequence similarity. What you find in the case of PBCV-1 vs. Chlorella is that the host enzyme shares only 48% amino-acid identities with the viral enzyme. "Well," you're saying, "but that's pretty good, right?" Not so fast. When you take the virus's enzyme sequence and run a search against the entire UniProt.org database, the most similar non-viral sequence turns out to be the reductase enzyme not of the virus's host (Chlorella) but of Haemonchus contortus, the barber-pole worm, with 53% sequence identities. Also very closely matched: the reductase from Pediculus humanas, the human body louse. Three other organisms also have a closer match of their reductases to the PBCV-1 reductase than Chlorella. (See table further below.)

So did this marine-virus reductase gene actually come from a louse, a worm, or a fungus, rather than from an algal host? Not likely. What's going on here, then? Frankly, it's a mystery. For one thing, we have no way of knowing how ancient the PBCV-1 reductase gene is or how fast it has evolved over the ages, relative to the host gene. Some scientists believe the three classes of ribonucleotide reductase originally stemmed from a common ancestor that was similar to the current class III (anaerobic) enzyme. This makes sense, in that the enzyme probably first came about in a highly anoxic ocean environment, billions of years ago, well before atmospheric oxygen began to accumulate, and maybe before sea water had accumulated much dissolved oxygen gas. The PBCV-1 virus reductase may derive from this ancient design. It's possible that Chlorella and its ancestors evolved extensively over the last few hundred million years, whereas the barber-pole worm and body louse (whose ancestors got the ancient class III proto-enzyme) may not have evolved as rapidly. Therefore, the worm enzyme, the louse enzyme, and the viral enzyme may all still share similarities with the progenitor enzyme that Chlorella no longer shares.

But there are also the forces of selection to consider. Modern ribonucleoside reductases incorporate allosteric control mechanisms that fine-tune the enzyme's capabilities with respect to deoxynucleotide (and small-peptide) concentrations. For example, a 50-amino-acid region at the beginning (N-terminal) end of the enzyme allows the enzyme to be feedback-inhibited by dATP. A virus interested in maximizing the production of deoxy-nucleotides might not want or need this sort of allosteric feedback mechanism. Also, the G+C content of the viral genome is significantly lower than that of the host  (40% vs. 60%), meaning that the viral enzyme might very well be optimized to produce deoxy-nucleotides in different ratios than the normal NTP-pool setpoints desired by the host. In short, it's possible to imagine that the virus's nucleotide requirements are, in fact, much more like a barber pole worm's than those of a healthy Chlorella.

Still, you have to admit: Nature comes up with strange bedfellows.

Here are a few protein matches between PBCV-1 (virus) reductase and other reductases:

Organism Length %ID Score E-value Gene identifier
Paramecium bursaria Chlorella virus 1 (PBCV-1) 771 100% 4727 0 A629R
Acanthocystis turfacea Chlorella virus Canal-1 763 76% 3746 0 Canal-1_104L ATCVCanal1_104L
Haemonchus contortus (Barber pole worm) 795 53% 2513 0 HCOI_01437900
Pediculus humanus subsp. corporis (Body louse) 795 53% 2483 0 Phum_PHUM350970
Salpingoeca rosetta (choanoflagellate) 779 51% 2479 0 PTSG_01558
Pneumocystis murina (fungus) 844 51% 2479 0 PNEG_03325
Schizosaccharomyces japonicus (yeast)83451%24780SJAG_04665
Chlorella variabilis (Green alga) 810

48% 2276 0 CHLNCDRAFT_32953
Cellulophaga phage phi13:1 789 47% 2039 0 Phi13:1_gp061
Cyprinid herpesvirus 3 806 45% 2092 0 CyHV3_ORF141 KHVJ151
Acanthamoeba polyphaga moumouvirus 849 43% 1947 0 Moumou_00516

Length refers to the total protein length in amino acids. Percent ID means the percent of target-protein amino acids that were an exact match against (aligned) query-sequence amino acids. Score is a figure of merit for the total matching; E-value represents the expectation that the matches could have occurred by chance (zero, here, in every case; meaning, these similarities probably could not have happened by chance). Finally, the Gene Identifier will let you look up these sequences at UniProt.org or other sequence database sites.

For more on the subject of ribonuceotide reductases in viruses, see the review of phage metagenome RNRs at http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3653736/.

Sunday, March 23, 2014

Nucleus-like viruses and their enzymes

Recent findings in virology have forced biologists to consider many notions that just a few years ago would have seemed heretical and/or science-fiction-like. For example, there is now serious discussion of the possibility that cellular life descended from viruses (the Virus World theory; see also this paper). A growing (but still minority) viewpoint is that viruses should be considered symbionts rather than simply parasites (see the review by Villereal). Some have dared to propose that the eukaryotic cell nucleus actually stemmed from a virus. Others have speculated the reverse: that the large DNA viruses are actually escaped, spore-like nucei. Meanwhile, some say that during an earlier RNA World, viruses became the original inventors of DNA.

There's no question that large viruses of the NCDLV class have nucleus-like properties. Within a short time of infection, these viruses set up a complex structure inside the cell known as the virus factory, and the factory looks a lot like a cell nucleus. The authors of a recent paper on Mimivirus (the famously huge virus that infects freshwater amoeba) admitted that in previous work, they did, in fact, mistake the virus factory for the nucleus. (See photo.)

Which is the nucleus and which is the virus factory? In this photo, VP is a virus particle (of the enormous mimivirus) developing inside Acanthamoeba. A smaller virus factory (S) is just beginning to form on the left.

Macroscopic aspects aside, the large "nucleocytoplasmic" viruses (some of which infect animals and marine life, not just amoeba) bring with them many genes for enzymes that are normally found in a cell nucleus. I'm not talking about genes for DNA polymerases, topoisomerases, etc., but genes that act on small molecules. In a previous post, I mentioned the example of PBCV-1 (a virus that infects the alga Chlorella) having its own gene for aspartate transcarbamylase (ATCase), which is an enzyme that catalyzes the first committed step in pyrimidine synthesis. This enzyme (common to most living things) is predominantly found in the cell nucleus of higher organisms.

There are other examples. Many NCLDV-group viruses have a gene for deoxy-UTP pyrophosphatase, an enzyme that breaks the high-energy phosphates off dUTP so that uracil isn't accidentally incorporated into DNA. One can imagine that after a virus invades a cell and unleashes its nucleases on the cell's own RNA, many ribonucleotides (breakdown products of RNA) will be liberated; and many of these will then be reduced to deoxy-nucleotides (by ribonucleoside-diphosphate reductase) in preparation for viral DNA synthesis. As it happens, dUTP is quite easily incorporated into DNA (and is promiscuous in its Watson-Crick pairing with other nucleobases); the resulting malformed DNA can trigger apoptosis in some cells. The virus takes no chances. It brings its own dUTPase to make sure uracil never gets into its DNA by mistake.

Some viruses bring their own gene for thymidylate synthase, to bring about the conversion of dUMP to dTMP (in other words, methylation of uracil, in its deoxy-ribonucleoside-monophosphate form, to give thymidine monophosphate). Some also have a gene for thymidylate kinase, which converts dTMP (often just called TMP) to dTDP (or TDP).

Yet another "small-molecule" enzyme encoded by large DNA viruses is ribonucleoside-diphosphate reductase (RDPR). This enzyme is fundamental to the whole DNA synthesis enterprise. Its job is to convert ordinary ribonucleotides to the deoxy form that DNA needs. Without this enzyme, you can make RNA but not DNA. So it's typically found in the cell nucleus (in higher organisms).

It turns out, a gene for RDPR is contained in a great many viral genomes. When I did a BLAST search of the protein sequence for Chlorella virus ribonucleoside reductase against the UniProt database of virus sequences, the search came back with 863 hits, spanning viruses belonging not only to the NCDLV class (pox, mimivirus, phycodnaviruses, etc.) but also the Herpesviridae, plus many bacteriophage groups as well. In terms of the sheer variety of virus groups involved, it's hard to think of another "small-molecule-processing" enzyme that spans as many viral taxa. We're talking about everything from relatively small bacteriophages to mimivirus, and lots in between.

The reductase gene is so widespread, it made me wonder what its phylogenetic distribution might look like. In other words: Are viral RDPRs related to each other? Are they related to the host's own RDPR? Does the enzyme's evolution follow the viral path, or the host path?

Just for fun, I obtained a number of ribonucleoside reductase (small subunit) protein sequences for viruses, plants, animals, bacteria, fungi, and various eukaryotic parasites (using the tools at UniProt.org), then fed the results to the tree-maker at http://www.phylogeny.fr. What I got was the following "maximum likelihood" phylogenetic tree. (See this paper for details on the tree algorithm. Also, be sure to check out this nifty paper to learn more about how to read this sort of tree.)

For convenience, names of viruses are depicted in blue. Notice how, except for the Vaccinia-Variola group, which is deeply nested, most of the viral nodes are ancestral to most of the higher-organism nodes; you have to go through many levels of viral ancestors to get from the original, universal ancestor (presuming there was one) to the reductase gene of the pig, say. From this diagram, it would appear that the Pox-family reductase gene is derived, in some way, from a highly evolved host. But that's the exception, not the rule. All of the other viral genes are outgroups and/or, more usually, ancestors of one another.

Mimivirus is fairly high up the chain and shows relatedness to two very common freshwater and soil bacteria (Pseudomonas and Burkholderia).

It would be fun to go back and remake the tree, adding more organisms. (If you end up trying this, let me know the results.) For now, I'm comfortable concluding that except for pox-family viruses, the ribonucleoside reductase produced by major DNA viruses and phages are not derived from current-day hosts. A parsimonious (but not necessarily correct!) explanation is that the phage reductases are ancestral to host orthologs; but it is also possible that the phage reductases derive from very ancient hosts (not depicted in the tree), with current-day hosts appearing to derive from phage genes when in fact the similarity is to a long-ago host ortholog. In any case, the tree shows that organismal RDPRs tend to be related to organismal RDPRs and viral versions are related to viral versions. What we don't see anywhere is a viral sub-tree growing out of a host sub-tree (as would be the case if the viral enzymes simply derived from modern host enzymes).

The UniProt identifiers of the protein sequences used in this study are given below in case you want to try to replicate these results (or perhaps extend them). To retrieve the protein sequences in question, go to http://www.uniprot.org/ and click the Retrieve tab, then Copy and Paste the following sequences (one to a line) exactly as shown:



O57175
P33799
M1I7H3
E5ERR7
Q6GZQ8
Q77MS0
P28847
M1I8A4
W0TWG5
Q7T6Y9
Q9HMU4
T0MT29
201403222BWOVN08AD
B3ERT4
F2II86
F2L908
U7RFH3
Q4KLN6
I3LUY0
B9RBH6
Q9LSD0
S8GD97
W4I9N3
Q4DFS6
A4HFY2
G3XP91
S8B144