Showing posts with label M. leprae. Show all posts
Showing posts with label M. leprae. Show all posts

Wednesday, May 21, 2014

The Pseudogene Hall of Fame

For "budget reasons" (supposedly), the Joint Genome Institute now requires all users to be registered and approved before they can use the (taxpayer-funded) https://img.jgi.doe.gov site. Fortunately, my registration was approved and I can use the excellent online genomics tools there, one of which produced the following table of Top Ten Organisms by Pseudogene Count.

Organism
Genes
Pseudogenes
60745
11398
38115
9178
38612
7186
6325
4011
31392
3818
5379
1670
4760
1589
5775
1523
5016
1507
2750
1086

Again: Don't expect the above links to work if you're not a registered JGI user. (I don't know if they will work for you or not.) This list was automagically generated by the Department of Energy's Joint Genomes Institute and I thought you might get the same kick out of it that I got. It's eye-opening to see the ratio of pseudogenes to "normal" genes in these organisms. Isn't it?

Taking the top three spots are mouse, rat, and humans. (Note: These counts should be taken with a bit of caution, as some have estimated the number of pseudogenes in the human genome to be much higher than the 7,186 shown here.) All the other spots in the chart except Arabidopsis (which is a leafy plant) are bacteria. The leprosy bacterium, which I've written about before, comes in tenth place.

If you're not familiar with the concept of pseudogenes, you might want to look at this post. Basically we're talking about genes that are thought to be disabled and no longer functional in the normal sense, although they may well be functional in some as-yet-unappreciated sense. (Otherwise, evolutionary theory says they should have been eliminated from most genomes eons ago.)

Personally, I believe pseudogenes are as much a feature of DNA as regular genes; certainly in higher life forms, they occur in great numbers. The vast majority of bacterial genomes in public databases are shown as having no pseudogenes. I find that (how shall I say?) not at all credible. Some day it will be obvious that almost every genome harbors pseudogenes; we simply lack smart enough software to detect them all right now.

Thursday, May 08, 2014

Ancient Drug Resistance Genes

Antibiotic-resistant bacteria have been in the news lately, with the usual scary talk about drug-resistant bacteria taking over the world, with a return to the Dark Ages if we don't Do Something.

Certain facts about drug resistance tend to get lost in these sorts of news stories. There's a common misconception that somehow the introduction of antibiotics (first in medicine, then in agriculture) caused bacteria to invent new drug-resistance genes out of nowhere. That's nonsense, of course. The Lederbergs showed in 1951 that drug resistance genes were/are preexisting, and did not come into being because of antibiotics. Rather, bacteria already equipped with such genes were able to survive treatment with antibiotics. Around 1968 it also became clear that bacteria can share drug resistance genes by exchange of extrachromosomal DNA (plasmids). Bacteria that don't have the magic genes can get them from their buddies.

But, so. How is it that the genes for antibiotic resistance already existed prior to the introduction of antibiotics into the food chain and into medicine? The answer is simple: Antibiotics are natural products produced by common soil and water microorganisms. Penicillin is produced by a mold; streptomycin is produced by the soil bacterium Streptomyces. Certainly hundreds (maybe thousands) of different kinds of antibiotics exist in the natural environment. They've been there for millions of years.

Mycobacterium tuberculosis ATCC 35801 (Erdman strain) has two "multidrug  resistance" genes, and they reside on the main chromosome (not on a plasmid). They're shared by other members of the Mycobacterium genus, including the leprosy bacterium, M. leprae. In the latter, one of the genes in question (MLBr_2224) is a pseudogene. It's interesting to compare this pseudogene with its counterpart, Erdman_0866, in Mycobacterium tuberculosis ATCC 35801. The two genes are nearly the same size: 1543 base pairs versus 1638. This is quite interesting in itself, inasmuch as most pseudogenes in M. leprae are truncated (averaging just 795 base pairs). But the comparison gets even more interesting when one does a side-by-side analysis of single nucleotide polymorphisms (changes to individual bases).

First I did a SNP comparison of the tuberculosis drug-resistance gene to its counterpart in M. indicus pranii, where there were 325 total base-pair differences, distributed amongst the 1st, 2nd, and 3rd base pairs of codons as:

98, 72, 155

This is the expected pattern: The largest number of mutations tends to accumulate in the third codon base (the so-called "wobble base"), because mutations in this base give a large percentage of synonymous codon changes owing to codon degeneracy. The next-highest number of mutations is expected in the first base, because there is some (but not much) degeneracy in that position. All mutations in the second base are non-synonymous and likely to affect enzyme function. Hence, the lowest number of mutations occurs in the second base.

When I ran the same check between the M. tuberculosis gene and its counterpart (the pseudogene) in M. leprae, I found that the mutational differences totaled 418 changes and segregated by base as:

119, 103, 196

Unexpectedly, the same pattern emerges. The reason this is unexpected is that one expects a pseudogene to contain frameshifts that would destroy the reading frame, mooting 1st/2nd/3rd-base comparisons. More generally, one expects massive random mutations all along a pseudogene's length (of the kind that would tend to make all three of the above numbers equal, even without frameshifts). So the fact that we still see the familiar pattern of 1st/2nd/3rd-base mutations is interesting.

The genes in question were unequal in length, so to do these comparisons I had to disregard a 112-base leader portion of the aligned genes. (I aligned the genes via ClustalW using Mega6.) I also ignored the unequal-length trailer portions of each gene, analyzing only the interior (aligned) portions from base 112 to base 1418. Also, to facilitate the analysis, I removed one adenine (representing a one-base insertion) in the M. leprae gene, at position 407, to restore the reading frame. But it's interesting that at that position there's a run of six adenines, indicative of a slippery site, of the kind associated with frameshift signalling.

The M. leprae pseudogene MLBr_2224 behaves as if it is still a normal gene, accumulating mutations in the "right places." An alternative explanation is that the M. leprae and M. tuberculosis genes were already in pretty much this orientation relative to each other before the two species diverged, and M. leprae simply accumulated very few additional mutations in the pseudogene after it became a pseudogene. (M. leprae is assumed to have experienced a massive pseudogenization event between 9 and 20 million years ago.)

For more about M. leprae's strange assortment of pseudogenes, see this post and also this post.

Monday, April 28, 2014

Are Some Leprosy Pseudogenes Turned On?

Most genomes, whether human or bacterial, contain significant numbers of pseudogenes (that is, genes that are presumed to be inactive due to internal stop codons, frameshift errors, or other serious defects). The usual presumption is that such genes are dormant or dead and thus are not expressed as proteins, since the proteins would be severely truncated or contain nonsense regions, etc.

However, we know that in Mycobacterium leprae (the leprosy bacterium), some 43% of the organism's 1,116 pseudogenes are transcribed. While most of the transcripts are no doubt used in some kind of regulatory capacity, it would be surprising if not a single transcript got translated into protein.

One sign that a protein gene is highly expressed is the presence, upstream of the start codon, of a strong Shine Dalgarno sequence. This is a special sequence of bases that serves as a binding area for 16S ribosomal RNA. The Shine Dalgarno sequence serves to increase the translational efficiency of the genes that have such signatures. (Not all do.) Generally, they are found ahead of high-priority/highly-expressed genes (such as genes for ribosomal proteins). Absence of a SD sequence doesn't mean the gene doesn't get translated. Many genes carry no SD signal.

I created scripts that looked at all 1,116 pseudogenes in M. leprae, to detect the occurrence of Shine Dalgarno sequences in the 20-base-pair region upstream of what would normally be the start codons of said genes. Intriguingly, 31% of pseudogenes carry a length-4 SD sequence (nearly 7 times the number of such sequences expected to occur by chance). By comparison, 50.6% of normal genes in M. leprae carry a length-4 SD sequence.

When I looked for length-5 SD signals, I found that 8.6% of pseudogenes carry such a signal, compared to 26.4% for regular genes. Length-6 signals were found for 33 pseudogenes (3% of the total of 1,116 pseudogenes) versus 176 normal genes (representing 10.9% of 1,604 normal genes). These numbers are about eight times higher than expected to occur by chance.

These numbers are summarized in the table below, where I also show similar findings for genes and pseudogenes of Bordetella pertussis.

Organism Data Set
Motif Length
Expected
Found
% of Genes
M. leprae

genes
(n = 1604)
SD6
5
176
10.9%
SD5
25
423
26.4%
SD4
119
812
50.6%
pseudogenes
(n = 1116)
SD6
4
33
3.0%
SD5
18
96
8.6%
SD4
83
346
31.0%
B. pertussis

genes
(n = 3377)
SD6
9
493
15.0%
SD5
49
1032
30.6%
SD4
259
1939
57.4%
pseudogenes
(n = 370)
SD6
0
38
10.20%
SD5
1
89
24.10%
SD4
11
194
52.40%

B. pertussis preserves a higher proportion of SD signals in pseudogenes than does M. leprae. This is expected, since most B. pertussis pseudogenes are still "in frame," whereas most (but not all) M. leprae pseudogenes harbor frameshifts.

Length-5-or-longer SD signals occur at about one third the rate in psuedogenes of M. leprae that they do in normal genes of M. leprae, but still much higher than expected by chance. If 43% of pseudogenes are transcribed (as we know they are), and a third of those transcripts have strong enough SD sequences to facilitate translation, it means about 159 pseudogenes in M. leprae could be expected to have expressed protein products. Those products would, of course, be truncated and/or contain nonsense regions. Presumably, many would be marked for proteolysis (either by the tmRNA system or through other mechanisms).


Thursday, April 17, 2014

The Pathogen's Playbook

When comparing pathogenic bacteria with non-pathogenic species of the same genus or family, we often find a common pattern. In the pathogen:
  • The genome is often reduced in size (particularly in endosymbionts, but also in others).
  • The genome is often shifted in the direction of higher A+T content (lower G+C content).
  • Many pseudogenes are present.
  • Often, the pathogen is a slow-grower in pure culture (if it can be cultured at all).
  • The pathogen has special nutritional needs.
An extreme case that illustrates all of these points is Mycobacterium leprae, the leprosy bacterium. It has fewer genes than its cousin, M. tuberculosis (which in turn has fewer genes than non-pathogenic Mycobacteria); its genomic G+C content is 8% lower than most other Mycobacteria; it contains over 1100 pseudogenes; it has a doubling time of two weeks; and it cannot be grown in pure culture (presumably because of fastidious nutritional requirements).

M. tuberculosis can be grown in the laboratory, but it, and its M. avium-group cousins, are very slow growers, taking anywhere from four days to two weeks to develop colonies on solid media.

It seems likely that some pathogens (certainly members of the Mycobacteria, but also the tiny Tenericutes, e.g. Mycoplasma, among many others) have evolved slow growth as a survival strategy. Certainly, organisms that have evolved an intracellular parasitic lifestyle need to be careful not to out-grow the host, if the relationship is to be a long one.

All of the factors listed above suggest a certain scenario, a "pathogen's playbook," if you will, which can be summarized as follows:
  1. The organism invades a warm-booded host.
  2. Phagocytes (white blood cells) ingest the organism.
  3. The phagocytes undergo a respiratory burst, flooding the microbe(s) with peroxides, hypochlorites, nitrous oxide, and other noxious oxidants.
  4. The flood of reactive oxygenated species triggers an SOS response in the microbe.
  5. The microbe's DNA undergoes massive damage. 
  6. Any surviving microbial cells are now pathogenic.
The SOS response is known to trigger mutagenicity. In Mycobacterium, for example, peroxides (as well as UV light) can induce up-regulation of dnaE, an error-prone polymerase. Since Mycobacteria are known to lack a MutS mismatch repair system, SOS-induced errors in DNA replication will almost certainly include uncorrected frameshift errors leading to the creation of pseudogenes. But that's a good thing, if you're a Mycobacterium interested in forming a longterm relationship with a host cell. The loss of certain genes (as long as they're not essential!) will likely slow your metabolism and make you dependent on host nutrients. Truly non-essential pseudogenes will simply be jettisoned over time, reducing the footprint of the remaining genome. Any pseudogenes that survive will likely have done so because they're now playing an essential gene-silencing role.

Let's expand on that last part. Take the dnaE gene, for example. M leprae has two copies of this gene, only one of which is functional. Suppose both copies were functional at the time of the massive pseudogenization event that converted so many of M. leprae's genes to pseudogenes 9 to 20 million years ago. After the pseudogenization event (probably a phagocytic respiratory burst), one copy of dnaE became a pseudogene. But continued transcription of the pseudogene in the forward direction means the pseudo-mRNA competes with the "normal" dnaE transcript for ribosomal attention. Transcription of the antisense strand of the disabled gene would, of course, create a messenger RNA product that could silence the normal transcript by doublestranded interaction. Either way, once the pseudogenization event is over, dnaE expression is attenuated—as it should be, once pathogenicity has been established.

Is it realistic to think M. leprae transcribes antisense strands of its pseudogenes? Given that E. coli has been found to contain ~1000 antisense transcripts, and given that we know M. leprae transcribes many of its pseudogenes, I think the answer has to be yes.

So the pattern is: infection, respiratory burst, massive mutation, silencing of many genes, and (oh by the way) creation of many brand-new gene products, some of them no doubt quite toxic to the host, as the result of gene truncation and pseudogene expression.

Tuesday, April 15, 2014

Coming to Grips with Pseudogenes

The term pseudogene was coined in 1977, when Jacq et al. discovered a version of the gene coding for 5S rRNA in the African clawed frog (Xenopus laevis) that was truncated yet retained homology with the active gene. Subsequent work has shown that in higher life forms, pseudogenes (genes that have been inactivated through one event or another) are almost as numerous as coding genes, with (for example) the human genome containing 10,000 or more pseudogenes. (A more recent estimate puts the number at 20,000.) Many of these pseudogenes are highly conserved. Looking at pseudogenes in the mouse and human, Svensson et al. found that of a group of 74 such genes that occur in both species, 30 appear to have been conserved since before the evolutionary divergence of mice and humans.

In higher organisms, pseudogenes are sometimes transcribed into RNA, with the RNA filling a regulatory function. For example, Korneev et al. found that simultaneous transcription of neural nitric oxide synthase (nNOS) and the antisense strand of a homologous pseudogene in the same neurons of Lymnaea stagnalis (a snail) leads to the formation of a duplex between the two strands and a reduction in nNOS translation. Further examples can be found in Pink et al. (2007), "Pseudogenes: Pseudo-functional or key regulators in health and disease?"

In bacteria, pseudogenes are somewhat rarer than in eukaryotes, but exist in significant numbers in many pathogens (including many species of Mycobacterium, Shigella, Brucella, Bordetella, and others). A study by Kuo and Ochman (2010) found that pseudogenes are swiftly eliminated from Salmonella. They describe "evidence of a strong deletional bias in Salmonella, such that genes that are not maintained by selection are rapidly inactivated and eliminated by mutational events." In fact, Kuo and Ochman found that pseudogenes are eliminated more rapidly than could be explained by the so-called neutral theory of evolution, indicating that the continued presence of pseudogenes exacts a high cost to the cell.

And yet, many bacteria with slow-evolving genomes (such as Mycobacterium species) retain their pseudogenes with high fidelity across evolutionary timespans. The most celebrated "pseudogene hoarder" of all time, M. leprae (the leprosy bacterium) appears to have acquired its 1000+ pseudogenes 9 to 20 million years ago. Meanwhile, the half-life of pseudogenes in Buchnera aphidicola was measured at 23.9 million years—a staggering number.

So on the one hand, we have work by Kuo and Ochman showing that pseudogenes in bacteria are rapidly eliminated, and on the other hand we have some bacterial lineages in which it seems pseudogenes are not only conserved but actively repaired over periods of tens of millions of years!

In Chapter 5 of Brucella: Molecular Microbiology and Genomics (2012, Caister Academic Press), Garcia-Lobo et al. describe their work with RNA sequence data from the bacterium Brucella abortus:
Twenty-four of the genes selected from the RNAseq data were annotated as pseudogenes in the B. abortus 2308 genome, which was considered a rather unexpected finding. By comparison with other Brucella genomes we can reduce the list of highly expressed pseudogenes to 16 (often, truncated parts of a gene are annotated as different pseudogenes especially in B. abortus 2308). This seems contradictory since high transcription of these genes, which should be not able to translate into functional proteins, will be contrary to biological economy. The high levels of transcription observed for these genes strongly suggest that they could be active genes and their products may perform functions unreported in metabolic reconstructions. High pseudogene expression may also indicate that these are very recently produced pseudogenes that did not turned down transcription yet by accumulation of mutations in their promoter or control regions. It is also possible that these pseudogenes may contain sequencing errors and they are indeed active genes.
It's almost comically obvious from this passage that the authors are troubled by their own finding that some pseudogenes in Brucella are highly transcribed. They try explaining it away by saying it could all be "sequencing errors."

A more parsimonious view is that pseudogenes that haven't been eliminated from a genome are, in fact permanent, legitimate fixtures of the landscape, in microbes just as in higher life forms. And as in higher life forms, pseudogenes in microbes are probably serving perfectly understandable regulatory functions (when they're not actually translated into protein products).

Kuo and Ochman have convincingly shown that useless pseudogenes are quickly eliminated. It follows that any pseudogenes that aren't swiftly eliminated are, in fact, serving a biological purpose, or else they wouldn't be there. This line of reasoning is already well accepted by researchers who study eukaryotic life forms. Those who study bacteria need to take a hint from their up-the-food-chain colleagues.

What could the hundreds of pseudogenes in Bordetella pertussis (or the 1000+ pseudogenes in M. leprae) be doing? First we need to get used to the idea that in bacteria, virtually all genes are transcribed, in both directions. It's been four years since Dornernberg et al. reported finding ~1000 antisense transcripts in E. coli, but no one seems to have gotten the memo.

A section of Rothia mucilaginosa genome (top) and a corresponding portion of Mycobacterium leprae (bottom); click to enlarge. The yellow gene, in each case, is DnaE (error-prone polymerase). Pink bands indicate areas of 65% or more homology between the two organisms. The small-diameter silver genes in the lower panel are M. leprae pseudogenes. "Normal genes" are shown in green. Notice that R. mucilaginosa has open reading frames on both strands of DNA, with many bidirectionally overlapping genes.

A look at the genome of the bacterium Rothia mucilaginosa DY18 shows that a very large proportion of "normal genes" have open reading frames on the opposite strand (see illustration). Bidirectional overlapping genes run throughout the Rothia genome. A massive annotation error? Maybe. Or maybe both strands are transcribed.

If massive wholesale transcription of antisense strands occurs in E. coli, as we know it does, certainly it's no stretch to imagine it occurring in Rothia mucilaginosa. And if it is occurring in Rothia, which is (incidentally) an opportunistic pathogen, how much harder can it be to imagine it occurring in another well-known pathogenic member of the Actinomycetales family, Mycobacterium leprae? We know already that upwards of 40% of M. leprae pseudogenes are transcribed. Antisense transcripts could well be playing a role in silencing certain gene essential genes when attempts are made to grow the organism in defined media. Forward transcripts could be producing nonsense or partial-nonsense/truncated proteins that are excreted as toxins or find their way to the cell wall as surface antigens. Any number of scenarios might be possible.

Some very low-hanging fruit is available to micobiologists who are willing to accept the obvious. Instead of wishing away pseudogenes or imagining them to be useless baggage, we should be looking at them as potential determinants of pathogenicity. We should consider their possible roles in modulating protein expression patterns. We should attempt to learn why they're conserved; what role(s) they're playing in cell physiology. The last thing in the world we should be doing is calling them "junk DNA."

Monday, April 14, 2014

Pseudogenes Are Not Junk DNA

In 2007,  a PLoS ONE paper by Ahmed et al. proposed a phylogeny for Mycobacteria in which M. leprae (the leprosy organism) is shown as a relatively recent branch off a very long tree, with M. tuberculosis depicted (in a decidedly fanciful schematic) as being of relatively recent provenance (35,000 years), diverging from M. canettii (a recently discovered cousin of tuberculosis) 3 million years ago.

The rather fanciful phylogenetic picture of Mycobacterium evolution presented by Ahmed et al. (2007). Click to enlarge.

The only trouble with this picture is that we know it's wrong. More exacting work has shown that M. tuberculosis is at least 3 million years old, and one paper estimates that the common ancestor of TB and leprosy may go back 66 million years. If the latter figure sounds dubious, consider that until recently, M. leprae wasn't thought to have any sister strains that could aid with dating the organism phylogenetically. But in 2008, the situation changed dramatically when it was realized that in Mexico, a distinct form of leprosy known as "diffuse lepromatous leprosy" (DLL) was actually due to a genetically distinct variant of Mycobacterium known as M. lepromatosis. When the genome for the latter organism was analyzed, it was found to contain the same stupendous assortment of pseudogenes contained in M. leprae, but detailed analysis of polymorphisms in the genomes of the two strains led to a surprising finding: Divergence of the strains appears to have occurred around 10 million years ago.

Another team found that the massive "pseudogenization event" that caused M. leprae (and its cousin, M. lepromatosis) to become saddled with a record number (1,116) of pseudogenes probably occurred on the order of 20 million years ago.

The age and stability of the pseudogenes in M. leprae can only be described as stunning. Conventional evolutionary dogma says that pseudogenes will inevitably be degraded and lost over time. Surely M. leprae can't be conserving and repairing pseudogenes over 10-million-year-long timespans? Pseudogenes are discardable junk.

Or are they?

An analysis of Buchnera aphidicola (the tiny Enterobacterial endosymbiont of the pea aphid) put the half-life of pseudogenes in that organism at 23.9 million years.

Human DNA reportedly contains over 12,000 pseudogenes. Some of these pseudogenes are quite old. Parallel nonsense mutations caused a pseudogenization of the uricase gene in apes during the early Miocene era (17 million years ago). We still carry the pseudogene in question—and it gets transcribed. According to a report by James T. Kratzer and colleagues at the University of Texas, Austin:
Despite being nonfunctional, cDNA sequencing confirmed that uricase mRNA is present in human liver cells and that these transcripts have two premature stop codons.
The inevitable conclusion is that pseudogenes are not, and should not be considered by default, "junk DNA." To the contrary, the default assumption should be that pseudogenes are ancient and conserved—because in most cases, that's exactly what they are.

What causes genes to "go pseudo"? Why are they conserved? What are they really doing? I'll tackle some of those questions in a followup post. Stay tuned.