Friday, June 27, 2014

A Tiny Genome with Lots of Structure

Mycoplasma genitalium is a super-tiny bacterial parasite of the human urinary system, causing about 15% of urinary infections in men. Its genome, encoding just 476 protein-coding genes, is often cited as the smallest genome of any organism that can be grown in pure culture. The genome is so small that scientists have for years used M. genitalium as a kind of litmus test for the essentiality of genes: If a particular kind of gene exists in M. genitalium (the reasoning goes), it must be truly essential for life.
Mycoplasma genitalium

Recently, I looked at the genome of M. genitalium from the point of view of latent nucleic-acid secondary structure. (Secondary structure refers to the ability of a single strand of DNA or RNA to base-pair with itself.) I probed each gene with a script designed to detect intragenic self-complementing regions of length eleven (so-called 11-mers), the idea being that complementary runs of bases shorter than that might occur at a high rate by chance. Given M. genitalium's base composition stats (adenine being most prevalent, at 36.24% of all bases in protein-coding message regions, thymine being next-most-abundant, at 32.21% of protein-coding bases), the most likely 11-mer, 'AAAAAAAAAAA', could be expected to occur by chance once every 70,672 bases. In actuality, that particular 11-mer doesn't occur in the M. genitalium genome, but if it did we could expect around 8 occurrences of it in a genome of 528,500 base pairs. Instead, what we actually find are 749 occurrences of complementary 11-mers in 283 genes.

The procedure used to check for 11-mers is as follows (in pseudocode):

for ( i = 0; i < genes.length; i++ )  // for each gene
    for ( k = 0; k < genes[i].length - 11; k++ ) // for all bases
        sequence = genes[i].substring( k, k + 11 ) // get 11 bases
        complement = getReverseComplement( sequence ) 
        if ( genes[i].match( complement ) ) // a match exists
        if ( matchNotWithinMask( mask ) ) // not inside a previous match
        if ( matchNotInsideSequence( sequence ) ) // not inside sequence 
             matches++    // count it as a hit
             updateMask( )  // update mask

Note that with even-length sequences it would be important to guard against self-matching sequences, since a sequence like AGCT is its own reverse-complement. With length-11 sequences this is not an issue. Nevertheless, it's possible for successive matches to overlap, and I wrote a check to guard against that. That's what  matchNotWithinMask( mask ) and updateMask() are all about.

It should go without saying, but I'll say it anyway: 749 complementing 11-mers in 283 genes is a very substantial (and quite unexpected) amount of internal complementarity. The question is: What's the purpose of all that (putative) secondary structure?

One possibility (which I've talked about in a previous post) is that the secondary structure allows for thermostatic control of gene expression. RNA thermometers are a well established phenomenon and it could be that M. genitalium needs sensitive control over thermal expression of certain genes.

Another possibility is that many genes incorporate magnesium-, manganese-, or calcium-sensitive riboswitches. RNA is a potent chelator of metal ions, particularly doubly-charged ions like Mg+2, which is smaller than sodium or potassium (with twice the charge) and thus has unusually high charge density. It's possible that under conditions of osmotic stress (with high influx of water and low ion concentrations) certain mRNAs relax or uncoil and thereby become translatable. If this is true (if certain genes are under osmotic control), we might expect to find that many M. genitalium genes with high secondary structure potential are membrane-targeted genes tasked with managing the import and export of various things. And that's indeed what we find.

Further below, I present a table with the top 100 M. genitalium genes containing the most putative secondary structure based on the 11-mer probe outlined above. The table shows the gene name, gene product or function (where known), gene size in base pairs, and the number of complementing 11-mers in the gene. The genes tend to fall into only a few categories. About a third of the genes (34 out of 100) are DNA- or RNA-binding genes. A dozen genes are transporters or permeases; another 12 fall in the category of "other membrane-associated genes" (these are marked with asterisks); and ten encode lipoproteins. Sixteen are "hypothetical proteins" (which should probably be reannotated as proteins of unknown function, since we can be fairly sure the genes are expressed).

It's interesting that so many genes with (putative) secondary-structure potential are nucleic acid processing genes. Among genes in this category are ten genes that either acylate or modify transfer RNAs. What's interesting about the latter group is that most involve tRNAs for non-polar amino acids (alanine, leucine, isoleucine, valine, phenylalanine, and methionine). The reason this is interesting is that non-polar amino acids are extensively used in membrane-associated proteins. Thus we have a situation, possibly, in which osmo-switches (osmotically sensitive mRNAs) control the expression of tRNA synthetases for amino acids used in membrane proteins.

It is quite possible that some of the (few) metabolic genes listed in the table are membrane-associated. This is likely true for phosphomannomutase, UDP-galactopyranose mutase, and glycerophosphoryl diester phosphodiesterase, for example. Also interesting is that 2,3-bisphosphoglycerate-independent phosphoglycerate mutase from Thermoplasma has been shown to be manganese-stimulated. Riboswitch modulation of such an enzyme would not be unexpected.

Altogether, 34 to 37 out of 100 genes listed in the table are membrane-associated, and another 7 are tRNA synthetases involving non-polar amino acids (heavily used in membrane proteins), tending to support the hypothesis that M. genitalium uses secondary structure of mRNA (and/or ssDNA) to modulate gene expression in osmotically sensitive manner.

Table 1. Genes with high secondary structure potential in M. genitalium. Genes marked with asterisks are membrane-associated genes other than transporters or permeases. The final column shows the number of complementary 11-mer pairs found in the gene.
Gene Product
Size (bp)
11-mers
MG_468 ABC transporter, permease protein
5353
22
MG_064 ABC transporter, permease protein, putative
3997
17
MG_218 HMW2 cytadherence accessory protein *
5419
16
MG_386 P200 protein
4852
15
MG_414 conserved hypothetical protein
3112
13
MG_075 116 kDa surface antigen *
3076
12
MG_422 conserved hypothetical protein
2509
12
MG_244 UvrD/REP helicase
2113
11
MG_191 MgPa adhesin *
4336
11
MG_292 alanyl-tRNA synthetase
2704
10
MG_018 helicase SNF2 family, putative
3097
10
MG_298 chromosome segregation protein SMC
2950
10
MG_345 isoleucyl-tRNA synthetase
2689
9
MG_338 lipoprotein, putative
3814
9
MG_307 lipoprotein, putative
3535
8
MG_031 DNA polymerase III, alpha subunit, Gram-positive type
4357
8
MG_080 oligopeptide ABC transporter, ATP-binding protein
2548
8
MG_321 lipoprotein, putative
2806
7
MG_340 DNA-directed RNA polymerase, beta' subunit
3880
7
MG_069 PTS system, glucose-specific IIABC component *
2728
7
MG_390 ABC transporter, ATP-binding/permease protein
1984
7
MG_525 conserved hypothetical protein
1996
7
MG_226 amino acid-polyamine-organocation (APC) permease family protein
1480
6
MG_378 arginyl-tRNA synthetase
1615
6
MG_192 P110 protein
3163
6
MG_312 HMW1 cytadherence accessory protein *
3421
6
MG_328 conserved hypothetical protein
2272
6
MG_341 DNA-directed RNA polymerase, beta subunit
4174
6
MG_291 phosphonate ABC transporter, permease protein (P69), putative
1633
5
MG_001 DNA polymerase III, beta subunit
1144
5
MG_411 phosphate ABC transporter, permease protein PstA
1966
5
MG_136 lysyl-tRNA synthetase
1474
5
MG_336 aminotransferase, class V
1228
5
MG_261 DNA polymerase III, alpha subunit
2626
5
MG_430 2,3-bisphosphoglycerate-independent phosphoglycerate mutase
1525
5
MG_053 phosphoglucomutase/phosphomannomutase, putative
1654
5
MG_195 phenylalanyl-tRNA synthetase, beta subunit
2422
5
MG_260 lipoprotein, putative
2299
5
MG_277 membrane protein, putative *
2914
5
MG_366 conserved hypothetical protein
2005
5
MG_250 DNA primase
1825
4
MG_447 membrane protein, putative *
1645
4
MG_254 DNA ligase, NAD-dependent
1981
4
MG_003 DNA gyrase, B subunit
1954
4
MG_397 conserved hypothetical protein
1702
4
MG_096 conserved hypothetical protein
1954
4
MG_334 valyl-tRNA synthetase
2515
4
MG_141 transcription termination factor NusA
1597
4
MG_241 conserved hypothetical protein
1864
4
MG_278 GTP pyrophosphokinase
2164
4
MG_012 alpha-L-glutamate ligases, RimK family, putative
865
4
MG_123 conserved hypothetical protein
1417
4
MG_266 leucyl-tRNA synthetase
2380
4
MG_119 ABC transporter, ATP-binding protein
1696
4
MG_068 lipoprotein, putative
1426
3
MG_047 S-adenosylmethionine synthetase
1153
3
MG_456 conserved hypothetical protein
1006
3
MG_122 DNA topoisomerase I
2131
3
MG_421 excinuclease ABC, A subunit
2866
3
MG_223 conserved hypothetical protein
1237
3
MG_375 threonyl-tRNA synthetase
1696
3
MG_364 expressed protein of unknown function
676
3
MG_185 lipoprotein, putative
2107
3
MG_423 conserved hypothetical protein
1687
3
MG_067 lipoprotein, putative
1552
3
MG_008 tRNA modification GTPase TrmE
1330
3
MG_306 membrane protein, putative *
1183
3
MG_242 expressed protein of unknown function
1894
3
MG_040 lipoprotein, putative
1777
3
MG_281 conserved hypothetical protein
1672
3
MG_259 modification methylase, HemK family
1372
3
MG_204 DNA topoisomerase IV, A subunit
2347
3
MG_216 pyruvate kinase
1528
3
MG_229 ribonucleoside-diphosphate reductase, beta chain
1024
3
MG_187 ABC transporter, ATP-binding protein
1759
3
MG_094 replicative DNA helicase
1408
3
MG_184 adenine-specific DNA modification methylase
955
3
MG_303 metal ion ABC transporter, ATP-binding protein, putative
1075
3
MG_045 spermidine/putrescine ABC transporter, spermidine/putrescine binding protein, putative
1453
3
MG_051 pyrimidine-nucleoside phosphorylase
1267
3
MG_004 DNA gyrase, A subunit
2512
3
MG_385 glycerophosphoryl diester phosphodiesterase family protein *
712
3
MG_360 ImpB/MucB/SamB family protein
1237
3
MG_457 ATP-dependent metalloprotease FtsH *
2110
3
MG_065 ABC transporter, ATP-binding protein
1402
3
MG_419 DNA polymerase III, subunit gamma and tau
1795
3
MG_295 tRNA (5-methylaminomethyl-2-thiouridylate)-methyltransferase
1105
3
MG_072 preprotein translocase, SecA subunit *
2422
3
MG_464 membrane protein, putative *
1159
3
MG_089 translation elongation factor G
2068
2
MG_439 lipoprotein, putative
820
2
MG_309 lipoprotein, putative
3679
2
MG_032 conserved hypothetical protein
2002
2
MG_194 phenylalanyl-tRNA synthetase, alpha subunit
1027
2
MG_137 UDP-galactopyranose mutase
1216
2
MG_203 DNA topoisomerase IV, B subunit
1903
2
MG_021 methionyl-tRNA synthetase
1540
2
MG_029 DJ-1/PfpI family protein
562
2
MG_255 conserved hypothetical protein
1099
2
MG_314 conserved hypothetical protein
1333
2

Thursday, June 26, 2014

Books Everyone Says Everyone Should Read

I found the graphic below in David McCandless's book Information Is Beautiful and felt it was worth sharing; I couldn't stop looking at it. Click to enlarge.


Sad to say, I've read shockingly few of these titles (maybe twenty percent of them?). Some, I read at too early an age and remember dimly. Most of the ones I did read do, IMHO, deserve to be on this list. I take issue, however, with The Da Vinci Code (an execrable dung-pile between two covers) and have to wonder why The Chronicles of Narnia made the list but Gravity's Rainbow didn't. Certainly Catch-22 deserves its prominent position in the word-cloud, but does the turgid Dune (look to the right of Lolita) deserve to displace The Road or [pick your favorite dystopian epic]? Does Stranger in a Strange Land (a wooden, laughably kitsch fable about a man from Mars) even hold a candle to a book like Flowers for Algernon? Does Ayn Rand rate not one, but two placements (for The Fountainhead and Atlas Shrugged)? Really??

Of course, David McCandless did not choose the items on this list. No single person did. It was drawn from Oprah's Book Club List, Goodreads.com, and other public sources. (Read the fine print at the bottom.) Garbage in, garbage out. Still, as infographics go, and considering how many word-clouds all of us (by now) have seen, it's surprisingly compelling.

Wednesday, June 25, 2014

Why So Many Helicases?

Many DNA-processing genes have an unusual amount of internal complementarity: regions of DNA in which the DNA can fold back on itself to form stable structures. A good example is the dinG gene of Mycobacterium tuberculosis, which encodes an ATP-dependent helicase. Using the DINAMelt server's Quikfold app, I obtained the following structure prediction for the first 1,000 bases of the (1,971-bases-long) dinG gene of M. tuberculosis Erdman strain.

Structure prediction for one strand of the M. tuberculosis dinG gene (first 1,000 bases). Almost the entire sequence folds back on itself. The only part of the original sequence that doesn't self-anneal is the tiny straight line at the lower right (red arrow). Click to enlarge.

Remarkably, almost the entire sequence can form a stable self-annealing structure. Only a few bases (see red arrow, above) lack the ability to form secondary structure. Bear in mind, what we're looking at is a stable conformation involving one strand of DNA only. (Each of the two strands of dinG can form this structure, independently of one another.) The structure shown above has a Tm (melting temperature) of 66.4°C in 1M saline, with mean-free-energy enthalpy of minus-2418.60 kcal/mol and a 37°C Gibbs free energy (ΔG) of minus-209.92 kcal, meaning that at 37°C, formation of the stable structure shown here (or one very much like it) is, energetically speaking, strongly favored.

Structures of this kind are often considered to occur in RNA, but if they also occur in single-stranded DNA, it raises interesting questions. If a gene has an energetically stable strands-apart configuration, getting the strands of duplex B-form DNA to separate might not be so hard. But more to the point, getting the self-annealing gene to come back together again as duplex DNA will require significant energy input. In molecular genetics, we're accustomed to the idea of duplex DNA requiring help from an ATP-dependent helicase to "open up" (unwind) the double helix in preparation for replication or transcription. The above diagram suggests that the problem isn't "opening up" a gene; the greater problem may be bringing the strands together again after they've assumed a stable strands-apart secondary structure. There's a substantial energy barrier to be overcome before the above structure can be relaxed into randomly coiling DNA.

This suggests that certain genes, like dinG, may be modal in terms of strand-separation state. Once the gene's strands are apart, they want to stay apart. There's an energy barrier to bringing the strands together again.

It's ironic that a helicase gene (dinG) has so much single-strand secondary structure. The gene product is a DNA-powered helicase; the gene needs its own protein product in order to zip up again. But maybe that's the point? Maybe it's a non-accidental feature.

In general, bacteria tend to have a remarkable number of helicase genes. M. tuberculosis (for example) has 16 different helicases. Other species have even more. (See table.)

Organism
Helicase genes
Myxococcus xanthus strain DK 1622
42
Frankia sp. strain QA3
41
Streptomyces cf. griseus strain XylebKG-1
39
Clostridium botulinum Hall strain
26
Psychromonas ingrahamii strain 37
25
Mesorhizobium sp. strain BNC1
21
Bacillus cereus strain F837/76
21
Escherichia coli B rel606
19
Anabaena cylindrica strain PCC 7122
19
Mycobacterium tuberculosis Erdman strain
16
Caulobacter crescentus strain NA1000
15

One might ask why this is so; why would a bacterium need 15, 19, 25, or 42 different helicases? It's quite unusual for a bacterial genome to have significant redundancy of genes, because when there are two copies of a given gene, one copy usually eventually becomes disabled (pseudogenized) and lost through random mutations. The very few exceptions to this rule tend to involve highly transcribed, highly necessary genes (such as ribosomal-RNA genes). It would be extremely unlikely for M. tuberculosis to carry around 16 "flavors" of a gene if they weren't all absolutely necessary. Duplicates would almost certainly be lost over time, especially in M. tuberculosis, which lacks a mismatch repair system. (Bacteria that lack mismatch repair enzymes have been shown experimentally to lose DNA fifty times faster than other bacteria.) The most parsimonious view is that the 16 helicases of M. tuberculosis are, in fact, critically necessary and perform different jobs.

I would suggest that perhaps the reason bacteria have so many helicases is that these are actually the "nucleic acid chaperones" that manage secondary structure in the sizable minority of genes that exhibit pronounced self-annealing of separated strands. It could be that most helicases are tasked with separating individual strands of DNA (and/or RNA) from themselves. Different types of secondary structure require different types of helicase to unravel. This might be why bacteria need so many helicases.

References
Helpful articles on DNA secondary structure:
  • Dimitrov, R. A. & Zuker, M. (2004) Prediction of hybridization and melting for double-stranded nucleic acids. Biophys. J., 87, 215-226.
    [Abstract] [Full Text] [PDF]
  • SantaLucia, Jr., J. (1998) A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proc. Natl. Acad. Sci. USA, 95, 1460-1465.
    [Abstract] [Full Text] [PDF]
  • Walter, A. E., Turner, D. H., Kim, J., Lyttle, M. H., Müller, P., Mathews, D. H. & Zuker, M. (1994) Coaxial stacking of helixes enhances binding of oligoribonucleotides and improves predictions of RNA folding. Proc. Natl. Acad. Sci. USA, 91, 9218-9222.
    [Abstract] [Full Text] [PDF]

Saturday, June 21, 2014

A Different Approach to Anti-Aging


In this fast-paced talk by Dr. Aubrey de Grey, we hear about a breathtakingly radical approach to forestalling aging, which (in a nutshell) involves moving all remaining mitochondrial genes into the nuclear DNA, so that mutations in mitochondrial DNA, per se, are rendered irrelevant. This technique of "obviation" of mitochondrial-DNA breakdown isn't a new idea (it's been around for at least 30 years). What's new is that, technologically, we're in a position to make it happen.

Most mitochondrial genes are, of course, already in the nucleus. The majority of scientists accept that mitochondria got their start as bacterial endosymbionts; a long-ago ancestor of today's alphaproteobacteria took up residency in an anaerobe. The anaerobe provided the invading bacterium with a nutrient-rich environment in which to live, while the bacterium provided oxygen-detoxification services (and a lot of adenosine triphosphate) to the host. Most likely, the invading bacterium had around 1,500 genes. Over time, ~500 redundant genes were lost and the remaining 1,000 or so migrated to the host cell's nuclear DNA (a much safer environment for DNA than the mitochondrion), leading to the present-day situation where (human) mitochondria have an extremely small circular chromosome encoding just 13 proteins. But we know mitochondria actually contain around 1,000 different proteins, most of which are encoded in nuclear genes.

The majority of mitochondrial genes (in the nucleus) encode proteins that are made in cyotplasm and imported into the mitochondrion. For import, proteins must be in an unfolded state. Some proteins (the most hydrophobic ones) are actually made on the surface of the mitochondrion and slurped into the interior of the mitochondrion as they're being made. Folding of the proteins then takes place inside the organelle.

Can the remaining 13 mitochondrial protein genes be moved to the nucleus? If we succeed in doing that, will cells live longer? What technical obstacles remain? What progress has been made? These and other questions are addressed in Aubrey de Grey's talk, which is well worth a listen.

Friday, June 13, 2014

Thermometer Genes

Heat shock proteins are an interesting class of proteins that provide "damage control" for enzymes when temperatures rise to the point where proteins start to unfold and refold improperly. Protein 3-dimensional structure is critical to proper enzyme function, and it doesn't take much thermal jostling to mess up a protein's structure. Therefore it's not surprising cells have their own miniature repair factories for refolding heat-misfolded proteins.

Collectively, heat shock proteins are part of a group of proteins known as chaperones, some of which are bonafide refoldases and others of which aid proteins in other ways. (For example, the ClpB protein rescues proteins from an aggregated state.)

GroEL is a so-called type I chaperonin involved in protein folding, assembly, and transport. Like many heat shock proteins, it's over-expressed at high temperatures and plays a critical role in growth and survival at non-permissive temperatures. Because of its importance in many cellular processes, GroEL is ubiquitous in bacteria, with most species having a single GroEL gene, but with about 30% of  genomes having two or more GroEL copies.
GroEL mRNA of Mycobacterium intracellulare can fold into the advanced low-energy secondary structure shown here.
The question of how cells up-regulate heat shock proteins during times of thermal stress is still largely open, although we know in some cases specific transcription regulator proteins are involved (but then the question becomes: how do the regulators know to up-regulate in times of heat stress?). The answer might not be that difficult. The messenger RNAs encoding GroEL and other heat shock proteins contain a great deal of secondary structure (that is to say, the RNA folds back on itself to form thermally sensitive structures). Recently, Wan et al. surveyed RNA thermal sensitivity in yeast and found thousands of so-called "mRNA thermometers": RNA molecules that unfold in response to heat. About three quarters of yeast RNA is thermo-stable at 37 degrees C, while around 55% of RNAs are unfolded at 55 degrees. In the folded state, RNA probably requires the help of helicases or other "helpers" to unfold, but at a high enough temperature, the molecules unfold by themselves and become eligible for translation by ribosomes.

To investigate the possible role of mRNA secondary structure in GroEL regulation, I wrote scripts that check a gene for all occurrences of length-9 (so-called "9-mer") nucleotide sequences that have a corresponding reverse-complement sequence in the same gene. When I checked the GroEL gene of Mycobacterium tuberculosis (Erdman strain), I found 14 pairs of complementarity 9-mers, representing regions of the gene that could, in theory, cause secondary structure to form in mRNA. A check of the sister organism M. intracellulare (whose GroEL gene is 80% identical to the M. tuberculosis version) showed 22 such complementary pairs.

Interestingly, the mutational differences between GroEL in M. tuberculosis and M. intracellulare do not appear to be randomly distributed along the gene. In M. tuberculosis, mutations occur at a rate of  0.12698 substitutions per site inside 9-mer regions (putative stems) versus a rate of 0.19722 for non-9-mer regions, indicating that (perhaps) selection pressure is different for self-complementing regions than for other regions. I found much the same thing in M. intracellulare, where the mutation rate was 0.16414 inside 9-mers and 0.19347 elsewhere.

Tending to confirm that selection pressure is different for the "secondary structure" regions versus other regions is the (surprising) finding that in M. tuberculosis complementary regions, the ratio of non-synonymous to synonymous mutations (Kn/Ks) is 0.526, versus 0.950 for other regions. In M. intracellulare, likewise, Kn/Ks is less in complementing regions (0.635) than in non-complementing regions (0.975).

To check whether these observations apply only to Mycobacterium or might be more widely applicable, I took a look at GroEL genes in Clostridium acetobutylicum strain ATCC 824 and Clostridium lentocellum strain DSM 5427. The Clostridia are phylogenetically quite distant from Mycobacteria (as confirmed by the fact that their GroEL genes share only 50% nucleotide sequence identity). A total of 462 mutations separated the two Clostridial genes. But again, the mutations segregated non-randomly according to whether they occurred in putative regions of complementarity (secondary structure) as opposed to non-complementing regions. In C. acetobutylicum the mutation rate in complementing regions was 0.23015 substitutions per site (29/126 bases) versus 0.28618 (433/1513 bases) for non-complementing regions, while in C. lentocellum the rates were
0.19047 (24/126) substitutions per site vs. 0.28949 (438/1513). For C. acetobutylicum the Kn/Ks ratios were 0.277 in 9-mers and 1.466 otherwise. For C. lentocellum, Kn/Ks was 0.538 in 9-mers and 1.415 outside 9-mers, tending to confirm that selection pressures are different in stems than in loops.

Bottom line, the data are consistent with a scenario in which secondary structure of GroEL mRNA (and/or ssDNA) plays a role in heat-activation of the gene, such that when temperatures exceed the melting point of secondary structures, the gene is eligible for transcription and/or translation. The gene is, in effect, its own thermometer.


Monday, June 09, 2014

How Do Bacteria Survive Radiation Damage?

Secondary structure of the Trad_1400 gene (encoding a MutT hydrolase) in Truepera radiovictrix.
In the 1950s, a tin of meat was exposed to a dose of radiation that was thought to be capable of killing all known forms of life, but the meat subsequently spoiled, and Deinococcus radiodurans (dubbed Conan the Bacterium by some) was isolated from it. Various members of the Deinococcus-Thermus group have shown themselves to be incredibly hardy, able to survive extremes of temperature and doses of radiation that, frankly, they shouldn't be able to survive. (If any group of bacteria were able to survive space travel, this group surely could.)

Members of the Deinococcus group probably learned their DNA repair tricks very early in the history of terrestrial life, before there was sufficient oxygen in the atmosphere to support an ozone layer. In those times (before about one billion years ago), ultaviolet light from the sun would have been strong enough to sterilize almost any exposed surface. UV radiation, when it's strong enough, causes single- and double-strand breaks in DNA (just as ionizing radiation from radioisotopes or cosmic rays will). How Deinococcus manages to survive such radiation is still something of a mystery, although several repair modalities have been elucidated. We know that these bacteria have high copy numbers of their genetic material, and this no doubt facilitates repair. Still, double-stranded DNA breaks, in most organisms, are quickly fatal if they accumulate.

It turns out, DNA from Deinococcus-group bacteria is unusually rich in internal (intra-strand) complementarity, which means single strands of DNA are capable (in theory) of folding back on themselves to form elaborate secondary structures of high thermal stability. One such structure, for the Trad_1400 gene of Truepera radiovictrix (a radiation-tolerant member of the Deinococcus group), is shown above. This particular structure has a 37°C Gibbs free energy of minus-71.47 kcal/mol (meaning the structure is more likely to form than randomly coiled ssDNA) and a Tm (melting temperature) of 63.1°C, meaning it should be thermo-stable to around 145°F. Almost the entire gene folds back on itself; the only portion that doesn't self-anneal is the flat line on the bottom containing 29 bases.

If the (separated) strands of Truepera DNA can assume stable self-annealed structures of this type, it would go a long way toward explaining how the organism could survive double-stranded breaks. Fire a random bullet at the DNA and you're bound to hit secondary structure, not canonical (B-form) duplex DNA. A double-strand break in a stem structure might liberate a stem/loop from one strand, but the other strand could unfold to form a template for immediate repair of the damaged strand. Something like this is probably going on in radiation-resistant Deinococcus members, which have evolved to allow more than the usual secondary structure in their DNA.

Friday, June 06, 2014

March of the Microbes

It can be hard to convey to non-biologists, who may think of germs as pests (an unseen, unclean residue to be neutralized with disinfectants), the grandeur and miraculous intricacy of microbial life. After more than three billion years of evolution, microbes are still, numerically speaking, the principal inhabitants of this planet; most of life on earth is invisible to the naked eye. But sightings of microbial life are all around us if we know where to look. In March of the Microbes: Sighting the Unseen (2010, Harvard University Press), Professor John L. Ingraham shows those of us with eyes to see just how and where to look.

It's impossible for me to come at this book with anything approaching objectivity, because I know the author well: I was Dr. Ingraham's student at U.C. Davis in the 1970s and did my graduate studies under his kind, insightful direction. And it was thanks largely to him that I developed what has turned out to be a deep, abiding, lifelong appreciation for the stunning ingeniousness (for there's no other way to put it) of microbial life.

The Ingraham lab, in those days, specialized in the study of pyrimidine metabolic pathways in Salmonella, but Dr. Ingraham's interests were broad and his understanding of microbial ecology impressively deep. Before undertaking some of the first studies of cold-sensitive mutants of bacteria, Ingraham had made important contributions to the understanding of fusel oil production by enologically important yeasts. But Ingraham was and is also an avid naturalist, as likely to identify a particular oak species as to diagnose the fungal disease causing its slow death. Serve him rainbow trout and he'll likely comment on the multiple layers of guanine crystals in the fish's shimmering scales before reminding you (as he does on page 24 of March of the Microbes) that the lack of a "fishy" smell in rainbow trout is due to the absence of trimethyl amine oxide, an osmolyte that would otherwise be converted by bacteria to oh-so-smelly trimethyl amine (TMA).

Go with Ingraham on a walk in the woods and he's sure to point out (as on p. 76 of MoM) that the age of a lichen on a boulder can be estimated from its size, based on its half-millimeter-a-year growth rate; and don't be surprised if he can name the particular cyanobacterium that gives the lichen its greenish hue. You'll forgive him, of course, when he uproots a small plant to show you the Rhizobium-fostered nodules on its roots, nodules that (when sliced open with a pocket knife) are red inside, crimson with the plant hemoglobin (p. 118) that only a symbiotically infected legume can produce. If you're worried about the safety of drinking water from that nearby mountain stream, he'll explain that you'd have to drink 250 gallons of such water before you're likely to imbibe a cell of Giardia (p. 298).

Reading March of the Microbes, I felt like I was on an extended nature walk with a world-class naturalist, except that this "walk" knows no physical bounds, for it extends from Sierra Nevada ski slopes, where slope operators mix dry powder made of Pseudomonas syringae cells (sold as Snomax) into the snow-machine slurry to facilitate ice-crystal nucleation (p. 263), all the way to miles-deep thermal chimneys on the ocean floor, where Pyrolobus furnarii grows in total darkness at 3000-psi pressure and 113-degree-Celsius temperatures (p.186); and from there to the inside of the 22-gallon rumen of a cow (p. 80), then to the 200-microliter hindgut of the termite, where, again (as in the cow), microbes do the vital work of digesting cellulose, something no higher life form can do.

The interconnectedness of life on earth is a recurring theme in March of the Microbes, as is the sheer adaptability of microbial life. We learn that only microbes can harness molecular nitrogen directly from the atmosphere; only microbes can return it. Only microbes can respire hydrogen sulfide and nitrate to sulfuric acid and nitrite, carving out caverns (such as the famous Carlsbad Caverns) in total darkness. Only microbes can utilize cellulose as a carbon source. Only microbes can survive temperatures well below freezing or significantly above boiling, or live through prolonged exposure to concentrated radiation in a nuclear reactor. Microbes are responsible for the world's great sulfur deposits; the calcium carbonate in the white cliffs of Dover; the 100-meter-thick diatomaceous earth pits in Clark County, Nevada; the massive saltpeter deposits (formed by decomposition of guano) in the coastal ranges of Chile's Atacama Desert. Only Archaea (prokaryotic extremophiles that were once thought to be bacteria) can make methane. And only a bacterium (Alcaligenes eutrophus, for example; there are others) can store 80 percent of its body weight as plastic (polyhydroxybutyrate).

Most bacteria are harmless to humans and other life forms, but March of the Microbes doesn't shy away from pointing out the exceptions. We learn about the (fungal) origins and history of potato blight, the history of syphilis and smallpox (one being a gift of the New World to the Old; the other, vice versa); the curious mode of action of botulism poisoning, and tetanus; the bacterial origins of peptic ulcers; the recent rise of the oddly pathogenic E. coli O157:H7; and the many intricate and interesting ways in which antibiotic resistant bacteria develop drug resistance. An entire chapter is devoted to viruses, which are rather harshly (but deservedly, some would say) villified, despite their contribution to evolution.

If one organism is glaringly absent from March of the Microbes, it would be (drum roll, please) Amoeba proteus, the prototype of the "lowly amoeba," which of course is not so lowly when you consider that the amoeba genome contains 100  times the DNA of a human genome. The freshwater amoeba is, conceptually at least, the progenitor of the modern-day phagocyte (white blood cell); many consider it the "nursery" organism on which Mycobacteria and other pathogens practiced and perfected their anti-phagocytic skills before adapting to warm-blooded animals.

Aside from the lack of an amoeba sighting, March of the Microbes lacks for very little, combining (as it does) refreshingly edifying bits of natural history and ecology with the occasional dash of epidemiology and immunology, plus just the right amount of common-sense chemistry, to provide a deeply satisfying view into the astonishingly diverse, always-full-of -surprises microbial world. It's the kind of book you wish you could never reach the end of, a 306-page Aha moment; a delight from start to finish.

Thursday, June 05, 2014

A Manganese Catalase Fusion Protein

In science, it often happens that finding the answer to a particular mystery only leads to further questions. That's certainly the case with the non-heme/manganese-based catalases in bacteria (which I talked about in a previous post). Regardless of origin, catalases faclitate the breakdown of hydrogen peroxide to water and oxygen. The nearly universal heme-based catalase (found in almost every living thing) comes in large-subunit and small-subunit varieties but is always fairly big (726 amino acids, in E. coli). Manganese catalases, by contrast, are always fairly small: around 276 amino acids. But there's one group of manganese catalases that comes in at 416 to 431 amino acids, more than 50% larger than the "typical" manganese catalase. I wanted to know why. Why are these catalases bigger? 

Finding the answer wasn't hard (I got lucky). But the answer only leads to more questions.

It turns out very few organisms manufacture the "big" manganese catalase. The organisms in question belong to just 3 genera: Rhizobium, Bradyrhizobium, and Rhodopseudomonas. (The closely related Mesorhizobium has a Mn catalase, but it's the "normal" size Mn catalase. Meanwhile, its cousin, Sinorhizobium, has no Mn catalase.)

The first 280 or so amino acids of the Mn catalases made by Rhizobium, Bradyrhizobium, and Rhodopseudomonas align quite well with the normal-size Mn catalases made by other organisms, the only difference being the 140-amino-acid trailer on the end of the "long versions." I used the alignment editor in Mega6 (tantamount to Notepad) to Cut the trailer portion out of one of the sequences and Paste it into the BLAST search field at UniProt.org. A search against the protein database revealed something quite interesting: The trailer portion of the Bradyrhizobium Mn catalase is a 45%-identity match for the enteric-bacteria yciF gene (which is quite a good match considering the phylogenetic distance between E. coli and Bradyrhizobium).

The E. coli yciF gene (top) is a 58% match for the Bradyrhizobium yciF gene (middle), which in turn is a 64% match for the trailer portion of katN (manganese catalase gene) of Bradyrhizobium, bottom.

Further investigation left little doubt that the "big" Mn catalases are indeed fusion proteins: yciF fused to the C-terminal end of the Mn catalase.

But: What in the world is yciF? It turns out to be a very widely distributed, highly conserved protein of uncertain function. The protein has been purified and its crystal structure determined at 2.0 Ã… resolution, but its function is still uncertain. What we do know is that it is stress-inducible (along with other yci-series genes) and seems to bind a metal ligand, probably iron; and it shares structural features with rubrerythrin, a non-heme iron protein implicated in oxidative stress protection in anaerobic bacteria and archaea. In E. coli strains that have Mn catalase, the yciF gene occurs two genes upstream (on the 5' side) of the catalase (katN) gene.

Because yciF is more widely distributed (and more highly conserved) than manganese catalase, and because most Mn catalase producers (including those with the fusion enzyme) have an additional, separate copy of yciF, it seems likely (to me, anyway) that the fusion protein was created by chance in the common ancestor of the Rhizobiales when the original katN gene was laid down by a phage or other mobile genetic element. (In E. coli, katN often occurs near phage genes.) Sinorhizobium lost the combo gene entirely, while Mesorhizobium (which makes a small Mn catalase) either obtained katN on its own or lost the 3' trailer from its fusion protein over time.

Since Bradyrhizobium (and the others) already have a separate yciF gene, it's a mystery why the trailer portion of the fusion gene continues to exist. It might very well provide a favorable enhancement of katN function somehow (maybe exploiting iron in an auxiliary catalytic center). If the trailer's doing nothing useful, it should have disappeared over time. (Maybe it did disappear from other Rhizobiales members, and just hasn't disappeared yet in the three genera that still make the "big" enzyme.) I have a feeling the trailer piece does do something useful. Unfortunately, no one has characterized the Braydrhizobium Mn catalase yet, experimentally. We'll probably have to wait until that happens to find out what the "big" enzyme is capable of.

Wednesday, June 04, 2014

The Other Catalase

Microbiology students are taught from Day One that strict anaerobes (organisms that are killed by exposure to oxygen) lack the enzyme catalase, which breaks down hydrogen peroxide to water and O2. You're already familiar with this enzyme if you've ever poured peroxide on a wound and seen it get foamy (blood is rich in catalase) or if you've used a peroxide bleach on your teeth, which causes saliva to become thick and gummy with microscopic bubbles. Catalase is a nearly universal enzyme, but strict anaerobes lack it (seemingly), because they are so seldom exposed to oxygen or its metabolites.

An unexpected result of genome data mining is the finding that non-heme (and thus non-iron-containing) catalase exists in a wide variety of bacteria once thought to contain no catalases. Manganese-containing catalase was first described in Lactobacillus, an organism that produces no cytochromes (and no porphyrins). The same type of enzyme was later experimentally verified in a thermophilic archeon, Pyrobaculum. It now turns out that many strict anaerobes previously believed to be catalase-free (such as most Clostridium species) contain this enzyme.

Phylogenetic distribution of manganese catalase (click to enlarge). At the top are the spore-forming Bacillus, with non-spore-forming Firmicutes (Aerococcus et al.) in a sub-branch below, and the cyanobacteria (Nostoc and Microcoleus) as an out-group to the Bacillus group. One cyanobacterial species, Cyanothece sp. PCC 7424 (a rice-field isolate), occurs as an out-node amongst the Gammaproteobacteria, indicating possible horizontal gene transfer. Archaeal organisms (Halalkalicoccus etc.) with this enzyme tend to be salt-lovers, although Pyrobaculum (not shown), which thrives in 2% salt, also has it. The phylo-tree was constructed from protein sequence alignments using Mega6 freeware. Node assignments were tested with 500 bootstraps.
Why would an anaerobe need catalase? Short answer: because small oxygenated molecules (like hydrogen peroxide and superoxide anion) are damaging to DNA, typically causing guanine to become oxidized to 8-oxo-guanine (which mispairs with adenine). Also, molecular oxygen irreversibly poisons the nitrogenase enzyme that many Clostridium members (and others) rely on to metabolize atmospheric nitrogen.

You would think that if molecular oxygen is a product of catalase action (and damages nitrogenase), Clostridium would not want to have catalase around in its cytoplasm. But Clostridium doesn't localize its catalase to the cytoplasm. It's a spore-surface enzyme. And this is key to understanding its distribution in nature.

Manganese catalase is rather sparsely distributed. It occurs just in certain taxonomic groups and certain species (see illustration, above). The enzyme seems to have been invented by the spore-forming Firmicutes (Bacillus and Clostridium) as a spore-coat enzyme. Many phylogenetically younger Firmicutes that have lost the ability to form spores also have the enzyme, as do some (not all) members of the Rhizobiales (not shown above), probably as an adaptation to low-iron niches. (The canonical heme-containing form of catalase requires iron.) There is evidence, in fact, that Lactobacillus leads a completely iron-free existence.

Where else do we find manganese catalase? Some (but not all) cyanobacteria have it. This is noteworthy in that the cyanobacteria form a specialized, environmentally hardened sessile cell called an akinete. (It's also noteworthy that both cyanobacteria and certain Clostridium members have nitrogenase.) The cyanobacteria that have the manganese catalase are primarily land-dwellers, however, not marine bacteria. This includes Anabaena (a moss symbiont and rice-paddy dweller), Microcoleus (which occurs in arid soils), and certain Nostoc members (which occur on rocks, in lichens), plus Cyanothece (from rice paddies).

Curiously, very few marine organisms have manganese catalase, the chief exceptions being Pirellula and Planctomyces. (And again, interestingly, Pirellula has a specialized sessile form as well as a motile form.) Many halophilic (salt-loving) archeons have the enzyme; surely those qualify as marine organisms? Not really. Halococcus, Natrinema, etc. are not planktonic; open-ocean waters have too little salt. (These organisms require upwards of 30% salinity, much stronger than the 3.5% salinity of sea water.) The salt-loving archeons live in drying-up seas (like the Dead Sea or Great Salt Lake) and at the edges of beaches, where salt concentrations skyrocket. These are, in effect, "terrestrial marine organisms," if that makes sense.

You can also find manganese catalase in some members of the Gammaproteobacteria (namely certain E. coli strains and some Pseudomonads), though curiously not in Shigella, Yersinia, or the Vibrio family (which is largely marine). The spotty nature of the enzyme's distribution among the enterics and Pseudomonads (which have no specialized sessile form) speaks to a possible horizontal gene transfer scenario.

There is little reason to believe manganese catalase is primordial. It certainly didn't come from the sea. In the first place, the enzyme is absent from Pelagibacter, Vibrio, and other important marine organisms. Secondly, sea water contains surprisingly little manganese (less than a part per billion). Marine organisms that do have catalase tend to have the heme-containing version of the enzyme (which makes sense, in that the oldest photosynthetic organisms long go mastered the art of porphyrin synthesis). Manganese catalase is a terrestrial adaptation, primarily of spore-, heterocyst-, and akinete-formers, with others obtaining the enzyme by lateral gene transfer. It's interesting that the Gammaproteobacteria that have manganese catalase (some enterics and some Pseudomonads) are opportunistic pathogens. They probably find the enzyme useful in combating the respiratory burst of phagocytes.

Tuesday, June 03, 2014

Odd Structures in a Prophage Genome


Contrary to what it looks like, this is not an arrival-gate diagram for an eastern European airport. It's a diagram of secondary structure of a portion of the mRNA for the yoqJ gene in B. subtilis phage SPBc2. Click to enlarge.
The common soil bacterium Bacillus subtilis has long been known to harbor an inducible prophage (virus) called SPBc2. Normally, SPBc2 is dormant in the Bacillus genome, but heat or treatment with mitomycin C will cause the phage to express itself and start a lytic infection cycle. Like many bacterial viruses, phage SPBc2 has a sizable double-stranded DNA genome, consisting (in this case) of 134,416 base pairs encoding 185 genes. It's one of the strangest collection of genes you'll ever want to meet.

The SPBc2 genome is strange, first of all, in its base composition. The following stats come from analysis of the coding regions of the genome:

Base Abundance
A 0.3670
T 0.2825
G 0.1983
C 0.1506

Notice how the purines, A and G (adenine and guanine) are about 30% more abundant than the pyrimidines, C and T (cyotsine and thymine). This is a stark violation of Chargaff's Second Parity Rule (which says A=T and G=C not only for double-stranded DNA but for a single strand, as well).

When purine abundance is tallied on a per-gene basis, you get the following histogram:

Purine abundance for N=185 protein-coding genes in B. subtilis phage SPBc2. Only 8 genes contain less than 50% purine bases.
All but 8 (out of 185) genes have a message-strand purine content above 50%.

And if you look at the gene sequence data, the genes even look funny to the naked eye, containing, as they do, long runs of bases, with funny repeats, like TTGAAAGGAAAAAAAGACGGCCTAAATAAA. One gets the feeling, when  examining the sequence data, that one is looking at tRNA or rRNA data rather than protein-coding sequences. And yet, they really are protein-coding sequences.

But it turns out there's a lot of secondary-structure info in the sequences. (See the picture at the top of this post.) Many of the genes contain self-complementing intragenic regions. I decided to verify this with some custom scripts. First, I had scripts look inside each gene for length-10 nucleotide sequences with a matching reverse-complement sequence further downstream (in the same gene). I expected to find one or two (or ten) hits, based on the fact that any given length-10 sequence of the bases A, G, C, and T can turn up at random in about one in every million base pairs. (Each base is worth about two bits, so a length-10 DNA sequence can encode as much info as a length-20 binary sequence; 2-to-the-20th is a little over a million.) What I found was an astounding 419 complementary length-10 sequences within 89 genes.

When I upped the sequence length to 12, I still found 64 intragenic complement sequences in 38 genes. A length-12 match should happen randomly about once every 16 million base pairs. The SPBc2 genome (as I said) is 134,416 base pairs long.

The high degree of internal complementarity means that when the phage's DNA strands are separated during transcription or replication, they probably assume very particular 3-dimensional structures, and also, the messenger RNA made from the phage's genes are probably rich in secondary structure. Exactly why those structures are needed is anyone's guess. They may protect the virus from restriction nucleases (although frankly this chore is already taken care of by the prophage's own methylases). Or they may attract ribosomes in some special way (many of the genes form tRNA-lookalike structures). They may help with virion packaging. Or they may do nothing special at all. I kind of doubt the latter possibility, though.

Sunday, June 01, 2014

A Gene Heatmap

Lately I've been using the great tools at genomevolution.org plus custom Canvas API scripts to render colorful heatmaps of aligned genes from phylogeneticaly diverse microorganisms. The following graphic is one such.
Glyceraldehyde-3-phosphate dehydrogenase genes of N=134 bacterial species, arranged in order of gene G+C content (high GC at the top). Hot colors are G and C. Cool colors are A and T.

What are we looking at? This is actually a composite rendering of the glyceraldehyde-3-phosphate dehydrogenase genes (DNA sequence info) from 134 bacterial species. Each gene is painted left-to-right (5' to 3') in a strip 4 pixels tall, with hot colors assigned to DNA bases G and C (guanine and cytosine), and cool colors assigned to bases A and T (adenine and thymine). Wherever there's a G or C, it gets painted red or red-orange. Wherever there's A or T, blue or blue-green. Same gene, 134 versions, varying significantly in G+C content. (The gene GC content ranges from a maximum of 69.2% at the top to 29.4% at the bottom.)

Why glyceraldehyde-3-phosphate dehydrogenase (GAPDH)? No real reason, except that it's a fairly universal (indeed, quite ancient) metabolic enzyme, reasonably compact (making possible a rendering that's not super-wide, as it would be for a larger gene), well-delineated genetically (not a fusion protein or an enzyme with multiple isoforms), and probably representative of a good many core metabolic enzymes. This is the enzyme that catalyzes the sixth step of glycolysis (sugar-breakdown). You may recall from Biochem 101 that the breakdown of glucose proceeds by splitting the twice phosphorylated molecule into two 3-carbon pieces. The triose phosphates in turn get phosphorylated by GAPDH before they transfer a phosphate to ADP to yield ATP, the 5-hour energy drink of all cells everywhere.

Alignment of genes was done via ClustalW in MEGA6 freeware. Rendering of the alignment FASTA file took about two seconds, in the browser, using 133 lines of custom JavaScript.