Showing posts with label tree of life. Show all posts
Showing posts with label tree of life. Show all posts

Friday, May 23, 2014

Looking for LUCA

In 1964, Emile Zuckerkandl and Linus Pauling wrote a paper (published the following year) for the Journal of Theoretical Biology suggesting the use of amino-acid and nucleic-acid sequences for deducing phylogenetic relationships. Ever since then, biologists have been trying to use sequence data to get to the root of the tree of life. Darwinian logic says that at some point, all cells had to have diverged from a Last Universal Common Ancestor (LUCA). Unfortunately, as pointed out by Doolittle and others, the quest for LUCA is greatly complicated by mutational saturation effects, reductive genome loss in important members of the most ancient taxa, convergent evolution, and non-negligible (yet difficult to estimate) amounts of horizontal gene transfer, among other serious problems.

An evolutionary tree of life based on analysis of N=420 genomes of free-living organisms. Proteomes are taxa and protein fold superfamilies are character data. Adapted from Kim and Caetano-Anollés, BMC Evolutionary Biology (2011), 11:140. Click to enlarge. See text for discussion.

The difficulty (I won't say folly) of trying to construct a well-rooted tree of life is made evident in various failed attempts to trace common descent via protein sequences. In January 2010, a few months after the sequencing of the 1000th bacterial genome, Karin Lagesen, Dave W. Ussery, and Trudy M. Wassenaar published a paper in which they expressed surprise over the fact that when they looked all 1,000 then-existing genomes (the number is now more than 10 times that), they could not find a single protein that was conserved across all bacteria. (Here, "conserved" means >50% amino-acid sequence identity.) Harris et al. took a slightly different approach, using the Clusters of Orthologous Groups (COG) database to search for universally conserved genes that follow the same phylogenetic patterns as ribosomal RNA (and therefore might constitute the ancestral genetic core of today's cells). The upshot:
Of the roughly 3100 COGs analyzed, only 80 were found to occur in all organisms. Fifty of these universally present genes showed the same phylogenetic relationships as rRNA.
Harris et al. found that the majority of universally conserved three-domain COG genes (37 of 50) are physically associated with the ribosome. Surprisingly, they found that "relatively few genes encoding proteins involved in DNA replication or transcription from DNA to RNA proved to be three-domain." In particular, RNA polymerases (except for certain subunits) did not follow rRNA distribution patterns and are not conserved across the three domains of life (archaea, bacteria, and eukaryotes). Moreover, the only component of the replicative DNA polymerase in modern cells that was found to be conserved across domains was DnaN (COG0592), the gene for the “sliding clamp.”

These disappointing results are understandable and perhaps expected, given the huge amount of deck-reshuffling that's happened in three billion years. It might well be that genome sequence data, with its constant churn, represents the wrong level of granularity for deep-phylogenetic studies. What matters for organisms, after all, is function, and function is an outcome of protein tertiary structure, not just primary structure.

With that in mind, Kyung Mo Kim and Gustavo Caetano-Anollés in 2011 published a brilliant study in BMC Evolutionary Biology called "The proteomic complexity and rise of the primordial ancestor of diversified life," relying on major structural motifs as the unit of phylogenetic discrimination. Defining protein domains at the highly conserved fold superfamily (FSF) level of structure, Kim and Caetano-Anollés used an iterative, parsimony-based phylogenomic approach to reconstructing FSF repertoires as upper and lower bounds of a presumed urancestral proteome ("ur" here meaning universal). Their conclusion:
The minimum urancestral FSF set reveals the urancestor had advanced metabolic capabilities, was especially rich in nucleotide metabolism enzymes, had pathways for the biosynthesis of membrane sn1,2 glycerol ester and ether lipids, and had crucial elements of translation, including a primordial ribosome with protein synthesis capabilities. It lacked however fundamental functions, including transcription, processes for extracellular communication, and enzymes for deoxyribonucleotide synthesis. Proteomic history reveals the urancestor is closer to a simple progenote organism but harbors a rather complex set of modern molecular functions.
The paper is quite long (14,700 words) and often relentlessly technical, but convincingly restores the quest for LUCA to the firm empirical grounding that such a quest seemed (for a while) to have been robbed of after Doolittle's "Uprooting the Tree of Life" and Dagan and Martin's "The Tree of One Percent." 

While parasitic microorganisms were found to occupy some of the most ancient branches of the superkingdom tree, Kim and Caetano-Anollés nevertheless decided to omit such organisms from their study since reductive evolution (wholesale loss of entire families of enzymes and their control systems) might otherwise queer the results. The final set of free-living organisms included 48 archaeal, 239 bacterial, and 133 eukaryotic members. To avoid potential problems with long-branch attraction, the researchers wisely sampled (at random) equal numbers of proteomes per superkingdom and replicated trees of proteomes, so that bacterial data (which of course predominated) wouldn't swamp archaea or eukaryota.

Among the many fascinating findings in the study:
  • The earliest start of organismal diversification occurred sometime between 2.91 and 2.03 billion years ago.
  • Translation had metabolic origins. It appeared only after the emergence "of a large number of metabolic functions, but before enzymes necessary for the synthesis of DNA."
  • Proteomic analysis of extant fold superfamilies (FSFs) showed that "over 200 additional FSFs are necessary in urancestral FSF sets to account for the complexity of the simplest organism in existence today."
  • None of the domains present in ribonucleotide reductase (RDR) enzymes was present in the min_set (representing the LUCA lower bound of complexity). Further, "We note that the reduction of ribonucleotides to deoxyribonucleotides involves the production of an active site thiyl radical that requires contacts with cysteines in all protein domains of the catalytic subunit of the oligomeric enzymatic complex, suggesting modern ribonucleotide reductase functions is [sic] indeed derived."
  • Commenting on the known active-site domain homology between class III ribonucleotide reductase and pyruvate formate lyase (a link proposed to have mediated the RNA-to-DNA biological transition), Kim and Caetano-Anollés point out that phylogenomic analysis at the fold-family level suggests the pyruvate formate-lyase domain emerged later than its ribonucleotide reductase counterpart. Therefore it's likely that the urancestor stored genetic information as RNA and not DNA.
Kim and Caetano-Anollés note: "The urancestor had an advanced metabolic network, especially rich in nucleotide metabolism enzymes, had primordial pathways for the biosynthesis of membrane glycerol ether and ester lipids, crucial elements of translation, including amino-acyl tRNA synthases, regulatory factors, and a primordial ribosome with protein synthesis capabilities. It lacked however transcription and in advanced evolutionary stages stored genetic information in RNA (not DNA) molecules."

The authors have many interesting things to say about the evolution of archaeal and bacterial membrane-lipid chemistry (and much else). If you're a biologist and you haven't yet read the Kim and Caetano-Anollés paper, do yourself  favor and take a look at it now. It's a fascinating read, no matter what side of the LUCA fence you're on.

Monday, May 06, 2013

Hydrogen Peroxide Powers Evolution

I'm about to offer a conjecture that is a bit preposterous-sounding but could well hold true. I actually think it does.

I propose that evolution, at the level of bacteria (though probably not at higher levels), is driven by hydrogen peroxide.

This theory rests on three assumptions: One is that the creation of new bacterial species happens almost entirely via lateral gene transfer, not heritable point-mutations. Secondly, bacteria (marine and terrestrial) are regularly exposed to challenges by hydrogen peroxide in the environment. Thirdly, those challenges drive lateral gene transfer.

Evidence for the first assumption is embarrassingly abundant. If you're not up to speed on the subject, I suggest you read the excellent paper, "Lateral Gene Transfer," by Olga Zhaxybayeva and W. Ford Doolittle in Current Biology, April 2011, 21:7, pp. R242-246 (unlocked copy here). It's now common to find that any given bacterial species can trace a good percentage of its protein base to "ancestors" that are too far removed horizontally to be ancestors in the conventional sense.

Consider E. coli. There are hundreds of strains of E. coli, with genes ranging in number from 4,100 to about 5,300 per strain. The problem is, the various strains of E. coli have only about 900 genes in common (and that's far too few genes to render a fully functional E. coli). The E. coli pan-genome actually takes in more than 15,000 gene families, total. Certainly, you can draw a family tree of E. coli based on 16S ribosomal polymorphisms, but that doesn't explain where the 15,000 pan-genome genes came from. The "family tree" metaphor quickly breaks down if you start drawing trees based on proteins. You get many conflicting trees—all of them correct.

Trees like this are fiction where bacteria are concerned.
The tree of life is more like a net of life or web
of life than a directed acyclic graph.
Where are all of the genes coming from? Other species, of course. They arrive by way of mechanisms like transformation, transduction, and conjugation. all of which allow direct entry of foreign DNA into a bacterial cell. At one time it was thought that conjugation could only occur between bacteria of the same species, but it is now known that cross-species conjugation also occurs (as, for example, between E. coli and Streptomyces or Mycobacterium).

Transduction, which is where viruses package up an infected host's genes in virus capsules that are then taken up by another cell, occurs naturally in bacterial populations in response to environmental factors like ultraviolet light and hydrogen peroxide. Exposure of a virus-carrying (lysogenic) cell to UV light or peroxide can induce runaway production of virus, and in fact this mechanism is used by Streptococcus to kill competitive Staphylococcus cells, in a clever bit of chemical warfare. It's been known for years that hydrogen peroxide can cause many types of bacteria to shed DNA. Now we know why: Hydrogen peroxide is a signalling molecule. It signals (among other things) lysogenic bacteria to go into a lytic cycle. It also signals cells to mount what's known as the SOS response, which is a global response to oxidative challenge. Years ago, Bruce Ames and his colleagues showed that exposing Salmonella to very dilute (60 micromolar) hydrogen peroxide caused the cells to differentially express 30 "SOS" proteins, including heat-shock proteins and low-fidelity DNA-repair systems. We know that hydrogen peroxide as dilute as 0.1 micromolar can induce phage (virus) production in up to 11% of marine bacteria. This is significant, because rainwater contains hydrogen peroxide in concentrations of 2 to 40 micromolar, and ocean water has been known to reach millimolar levels of H2O2 after a rain storm.

If you're wondering why rain contains hydrogen peroxide, the peroxide gets there in two ways. One is UV-frequency photochemistry (where water is cleaved to H and OH, then reforms as H2 and H2O2); the other is via ionization reactions caused by lightning. (Lightning is energetic enough to bring airborne oxygen and water to a plasma state. The resulting ionization and rearrangement of free atoms yields a certain amount of hydrogen peroxide.) The presence of H2O2 in rainwater has been confirmed many times, and in fact there's a well-preserved "fossil record" of it in polar icepacks, going back centuries. (Polar snowpacks contain from 10 to 900 ppb of H2O2; it varies seasonally, the max coming in summer.)

Bottom line, every rain event (over land, over sea) constitutes a hydrogen peroxide challenge for microbes. Which induces viral transduction (and a release of whole-cell DNA through lysis, some of which will be inevitably be used in transformation). It also induces low-fidelity DNA repair (which is guaranteed to help evolution along). Every rain event, in other words, is a chance for evolution to do its thing. For bacteria, that means gene-sharing within and across species lines.
Darwin's theory of a tree-like ancestor basis
for all living things is dead wrong, at
least for bacteria.
W. Ford Doolittle (who wrote a classic book chapter about lateral gene transfer called "If the Tree of Life Fell, Would We Recognize the Sound?") estimates that if a horizontal gene transfer occurs once every ten billion vertical replications, "it would be enough to ensure that no gene in any modern genome has an unbroken history of vertical descent back to some hypothetical last universal common ancestor." (See this article.)

It's obvious (to me, at least) that every rain event carries with it the potential to cause far more gene transfers than are necessary (according to Doolittle) to make vertical inheritance fade into insignificance as an evolutionary bringer of change. The hydrogen peroxide in rain has been driving lateral gene transfer in bacteria for eons. In fact, it is arguably the dominant driver of evolution in bacteria.

Sorry, Mr. Darwin. Point mutations handed down to sons and daughters just isn't cutting it.