Showing posts with label tuberculosis. Show all posts
Showing posts with label tuberculosis. Show all posts

Monday, May 12, 2014

Problems with the Tuberculosis Genome

These days, most genes, in most sequenced genomes, are machine-annotated with a minimum of human intervention, and as a result around 30% of gene annotations are inaccurate, either as to functional assignment or as to reading frame.

It's not hard to find serious errors in published genomes. In the genome of Mycobacterium tuberculosis ATCC 35801, for example, there's a 17-kilobase-pair section of the genome that's full of errors (see below).

A view of two tuberculosis bacterial genomes showing a 17,000-base region (denoted by pink) of 100% sequence similarity. Notice that even though the genomes are 100% sequence-identical in the region, the number, strand orientation, and sizes of genes differ. The yellow gene in the top panel is identified by GeneMarkS+ as "secretion protein EspK," while the yellow gene in the lower panel is identified as DNA Ligase (ligA) and is much smaller (and occurs on the opposite strand).

In this graphic, the same region is shown for two strains of M. tuberculosis. On top is M. tuberculosis Strain EAI5 and on the bottom is M. tuberculosis ATCC 35801. To browse the genomes in your own browser, go to this link and click the pink Run GEvo Analysis! button. When the panels appear, you'll be able to click on individual genes to see what they (supposedly) are.

The yellow-colored gene in the top panel is annotated as "secretion protein EspK," while the corresponding region in the bottom genome has two genes (green and yellow) on opposing strands of DNA, with one gene (in green) given as a "hypothetical protein" and the other (in yellow) as "DNA ligase." Note that the entire region covered by the pink bands is 100% identical from organism to organism in terms of DNA sequence. And yet one annotation program found 13 genes while the other found 17.

Sadly, this is not an unusual situation.

How can you tell which annotations are correct? Unfortunately, it takes some investigation. Consider the two yellow genes. They can't both be correct as shown, because one is twice as long as the other and each is on a different strand of the DNA! And yet, if you obtain the DNA sequence of each, and use it as a query sequence in a BLAST search, you'll come up with "good" hits for each gene (because there are other incorrectly identified genes in public databases, matching each query).

It turns out the "secretion protein EspK" (the large yellow gene in the top panel) is correctly identified. To determine this, I downloaded the FASTA sequence for the large gene, plus the sequence for the same gene in the same location in the M. canetti CIPT genome (identified as "hypothetical alanine and proline rich protein"), plus the same gene in M. canetti CIPT 140070010, with the idea of identifying the differences in the (aligned) DNA sequences with respect to their codon positions.

When I had a script identify mutations by codon location, the number of differences at the three possible codon locations, were:

base 1: 4 changes
base 2: 2 changes
base 3: 19 changes

This is exactly the pattern you would expect if the reading frame is correct. Base 3 (the wobble base) tends to accumulate the majority of mutations, because changes in a codon's third base usually result in no amino acid change (due to the genetic code's degeneracy). Base 1 has a small amount of degeneracy and thus has the next-highest number of mutations. Base 2 has no degeneracy; all changes to this base result in a different amino acid. (All changes are non-synonymous, in other words.) 

I repeated this check using the so-called "DNA ligase" gene of Mycobacterium tuberculosis ATCC 35801, but using sequence data from the bottom strand of DNA, with 1329 bases' worth of additional 3'-end sequence data in order to cover the same amount of DNA as the other sequences. I got difference data of 3, 1, and 8 for bases one, two, and three. This verified that the active gene is actually on the bottom strand.  ERDMAN_4254 is incorrectly annotated as a top-strand DNA ligase. Instead of two genes, ERDMAN_4253 and ERDMAN_4254 (on two opposing strands), Mycobacterium tuberculosis ATCC 35801 should be shown as having a single large gene on the minus-one strand. The correct identification (based on more than thirty E=0, 100% identity hits in other strains) is, in fact, "secretion protein EspK."

Unfortunately, at least 3 other tuberculosis strains, namely M. tuberculosis GuangZ0019, plus Strain CCDC5180 and Strain CCDC5079, also harbor an incorrect ligA annotation, and Mycobacterium isn't the only organism with a "bogus ligA" problem. Micromonospora sp. M42 also has a fake ligA gene and I'm certain there are others.

Bottom line, don't believe everything you read in genome annotations. The annotation may say "DNA ligase," but you could actually be looking at something else entirely.




Monday, April 14, 2014

Pseudogenes Are Not Junk DNA

In 2007,  a PLoS ONE paper by Ahmed et al. proposed a phylogeny for Mycobacteria in which M. leprae (the leprosy organism) is shown as a relatively recent branch off a very long tree, with M. tuberculosis depicted (in a decidedly fanciful schematic) as being of relatively recent provenance (35,000 years), diverging from M. canettii (a recently discovered cousin of tuberculosis) 3 million years ago.

The rather fanciful phylogenetic picture of Mycobacterium evolution presented by Ahmed et al. (2007). Click to enlarge.

The only trouble with this picture is that we know it's wrong. More exacting work has shown that M. tuberculosis is at least 3 million years old, and one paper estimates that the common ancestor of TB and leprosy may go back 66 million years. If the latter figure sounds dubious, consider that until recently, M. leprae wasn't thought to have any sister strains that could aid with dating the organism phylogenetically. But in 2008, the situation changed dramatically when it was realized that in Mexico, a distinct form of leprosy known as "diffuse lepromatous leprosy" (DLL) was actually due to a genetically distinct variant of Mycobacterium known as M. lepromatosis. When the genome for the latter organism was analyzed, it was found to contain the same stupendous assortment of pseudogenes contained in M. leprae, but detailed analysis of polymorphisms in the genomes of the two strains led to a surprising finding: Divergence of the strains appears to have occurred around 10 million years ago.

Another team found that the massive "pseudogenization event" that caused M. leprae (and its cousin, M. lepromatosis) to become saddled with a record number (1,116) of pseudogenes probably occurred on the order of 20 million years ago.

The age and stability of the pseudogenes in M. leprae can only be described as stunning. Conventional evolutionary dogma says that pseudogenes will inevitably be degraded and lost over time. Surely M. leprae can't be conserving and repairing pseudogenes over 10-million-year-long timespans? Pseudogenes are discardable junk.

Or are they?

An analysis of Buchnera aphidicola (the tiny Enterobacterial endosymbiont of the pea aphid) put the half-life of pseudogenes in that organism at 23.9 million years.

Human DNA reportedly contains over 12,000 pseudogenes. Some of these pseudogenes are quite old. Parallel nonsense mutations caused a pseudogenization of the uricase gene in apes during the early Miocene era (17 million years ago). We still carry the pseudogene in question—and it gets transcribed. According to a report by James T. Kratzer and colleagues at the University of Texas, Austin:
Despite being nonfunctional, cDNA sequencing confirmed that uricase mRNA is present in human liver cells and that these transcripts have two premature stop codons.
The inevitable conclusion is that pseudogenes are not, and should not be considered by default, "junk DNA." To the contrary, the default assumption should be that pseudogenes are ancient and conserved—because in most cases, that's exactly what they are.

What causes genes to "go pseudo"? Why are they conserved? What are they really doing? I'll tackle some of those questions in a followup post. Stay tuned.


Saturday, April 12, 2014

The Most Deadly Pathogen of All Time

Few bacterial species have had as great an impact on humankind as the members of the Mycobacterium family, which encompass the causative agents of (among other ailments) leprosy, tuberculosis, and Crohn's Disease in humans, and Johne's Disease in farm animals. Leprosy is known from antiquity and continues to strike 200,000 or more people each year worldwide. Tuberculosis, which affects (subclinically) one in three persons worldwide, continues to kill well over a million people a year and has caused a billion deaths in the last two centuries, more than all the wars and genocides of history combined.

The association of M. avium subspecies paratuberculosis (MAP) with Crohn's Disease is still considered controversial by some, but if in fact Koch's criteria have already been met, MAP adds millions more to the toll of human misery caused by Mycobacterial infection.

Colonies of Mycobacterium have a
characteristically waxy consistency.
Shown here: colonies of M. tuberculosis.
What are these bacteria? Where did they come from? How have they managed to be so successful in causing death and disease?

The prefix "myco" means fungal, but these are not fungi we're talking about. Mycobacteria are soil- and water-borne bacteria that produce an extraordinarily complex cell wall containing not only the usual (for bacteria) peptidoglycans but also:
  • Arabinogalactan
  • Mycolic acids
  • Lipoarabinomannan
  • Extractable lipids including glycolipids, phenolic glycolipids (PGL), glycopeptidolipids (GPL), waxes, acylated trehaloses, and sulfolipids
In contrast to most cell-wall fatty acids (which contain carbon-carbon double bonds susceptible to oxidation), mycolic acids are cyclopropanated and resistant to oxidation, not to mention extremely hydrophobic. The Mycobacterial cell wall thus presents a formidable physical barrier to antibiotics, and it was with considerable dismay that physicians realized, early on, that penicillin would have no benefit in treating tuberculosis. When an antibiotic that could attack M. tuberculosis was finally discovered (streptomycin), it resulted in a 1952 Nobel Prize for Ukrainian American Selman Waksman (although in reality the discovery was made by a post-doc in Waksman's lab, Albert Schatz).

The Mycobacterial cell wall is famously complex, but it also has the curious habit of disappearing entirely, under nutrient-starvation conditions. Like many other bacteria, Mycobacteria can, under certain conditions, shed their cell walls and take on a so-called L-form morphology, in which cells (bounded only by a thin and osmotically vulnerable cell membrane containing just 7% of the usual amount of peptidoglycan) exist as protoplasts which are nonetheless able to reproduce and thrive, producing distinctive colonies on solid media and giving rise, in vivo, to tiny spherules that are often confused with Russell bodies in cancer biopsies. The medical significance of the mysterious L-forms is still debated, after more than 100 years.

The very small red filaments here are cells of  
Mycobacterium avium living inside lymph-node
macrophages in an immunocompromised individual.
One thing most Mycobacterial species have in common is slow growth. Cultures of M. tuberculosis and MAP often require weeks to develop, and M. leprae (which can't be grown in pure culture at all; it can be lab-grown only in the footpads of mice or armadillos) has the longest known generation time of any bacterium, at two weeks.

Ironically, pathogenic strains of Mycobacterium seem to have evolved slow growth as a survival strategy. (This certainly makes them hard to treat with antibiotics. Most antibiotics are effective only in disrupting the growth of actively growing cells.) The lack of DNA mismatch repair enzyme systems (MutS, MutL, and MutH) may be an outcome of the fact that slow DNA replication in these organisms, in and of itself, ensures reasonably high-fidelity replication. On the other hand, lack of a mismatch repair system could be why pseudogenes (genes inactivated due to frameshifts or other errors) abound in Mycobacterial species. M. leprae famously has over 1000 pseudogenes; M. smegmatis strain JS623 harbors over 200 pseudogenes; M. canettii (strain CIPT 140010059) and M. rhodesiae (strain NBB3) both have over 100. (For a good review of Mycobacterial DNA repair systems, see this 2011 paper.)

Unlike Yersinia pestis, the plague organism, which may be less than 20,000 years old (very young in bacterial species time), M. tuberculosis, as a species, appears to be at least 3 million years old, although this number should probably be considered a minimum age, subject to upward revision. (The species was thought to be only 35,000 years old as late as 2002, before a more detailed genetic analysis established the 3-million-year estimate of its age. The numbers should be viewed with caution, however, since they're based on mutation-rate assumptions derived from data for E. coli.)

The question of how M. tuberculosis has managed to achieve its distinctive pathogenic profile is a matter of active ongoing research, and likely will be for a long time. A recent review article reminds us: "The [complete genome] sequence of the pathogen Mycobacterium tuberculosis strain H37Rv has been available for over a decade, but the biology of the pathogen remains poorly understood."

Miscellaneous Links
List of famous T.B. victims—Brontë family, Balzac, Kafka, Thoreau, Kant, Chekhov, Orwell, Schrödinger, Vivien Leigh, Arline Feynman (wife of the famous physicist), the list goes on.
Tuberculosis in Literature and the Arts
The T.B. Blues (Jimmie Rodgers, 1931) This song, famously covered by Leon Redbone (among others), was written by Rodgers after he contracted the disease at age 27. He died eight years later.
World Health Organization TB Stats (landing page)
The Tuberculosis Systems Biology Program