Showing posts with label Pseudomonas. Show all posts
Showing posts with label Pseudomonas. Show all posts

Tuesday, May 06, 2014

Biology's Dirty Little Secret

Biology has a dirty little secret. It's a well-known secret among those who deal with sequenced-genome data intensively, but I suspect many non-biologists are unaware of the problem, which is: Much of the existing genome data (for sequenced genomes, ranging from bacteria to human DNA) is either corrupt or misannotated.

"Junk DNA" probably doesn't exist in living cells. But it certainly exists in published genomes.

A substantial portion of published genome data is suspect, at this point, either because of contamination issues, technical problems surrounding DNA sequencing technology, or faulty gene annotation. An example is the Oryza sativa indica (rice) genome, which inexplicably contains at least 10% of the genome of the bacterium Acidovorax citrulli. There's also a Culex (mosquito) genome with a complete copy of Wolbachia embedded. The genome of Rothia mucilaginosa DY-18 contains over 300 genes incorrectly annotated in antisense orientation (as does the genome of Burkholderia pseudomallei strain 1710b, a truly execrable train-wreck of a genome).

Another example of a genome gone wrong (arguably) is that of the bacterium Ktedonobacter racemifer, which is filled with forward and backward copies of transposases. Incredibly, one in 13 Ktedonobacter genes is a transposase, integrase, or resolvase (and that's not counting the many "hypothetical proteins" with "transposase-like" mentioned in the gene ontology notes). Disregarding the 40% of that organism's genes that are marked as hypothetical proteins, one can say that in Ktedonobacter, one in four genes of known function is a transposase, integrase, or resolvase. (Some of the organism's 4000+ "hypothetical proteins" are actually transposases incorrectly annotated in an antisense orientation.) Common sense says something's amiss.

The misannotation problem is getting worse over time. This graph (from Schnoes et al., 2009, PLoS ONE) depicts the number of sequences (left y-axis, bar graph) found to be correctly annotated in green. Sequences found to be misannotated are shown in red. The bars for each year represent only the sequences deposited into the database in that year. The fraction (right y-axis, line plot) of sequences deposited each year into the Genbank NR database that were misannotated is given by the open nodes, connected by the black line to aid in visualizing the overall trend.

The "dark matter" problem in microbial genetics is widespread and openly acknowledged. At least 20% 28.3% (according to the Joint Genome Institute) of bacterial genes are annotated as "hypothetical protein," and most of these are so annotated because they have no sequence similarity match to any known protein. In many cases, there's no match because many of the sequences are in the wrong reading frame, or have an improperly located start codon (or other serious issues). When Ely and Scott (PLoS ONE, 2014) manually reannotated the genome of the bacterium Caulobacter crescentus, they identified 11 new genes, modified the start site of 113 genes, changed the reading frame of 38 genes, and found that 112 "hypothetical proteins" were actually non-coding DNA (not genes at all). A recent transcriptome analysis of the archaeon Sulfolobus solfataricus resulted in correction of 162 gene annotations and the addition of 80 new open reading frames. But these numbers barely hint at the extent of gene misannotation. In examining the Gene Ontology database (GOSeqLite), Jones et al. found:
Annotations made without use of sequence similarity based methods (non-ISS) had an estimated error rate of between 13% and 18%. Annotations made with the use of sequence similarity methodology (ISS) had an estimated error rate of 49%.
Surprisingly, the use of sequence similarity as a guide to function identification is less reliable than non-SS methods. This is no doubt partly a reflection of the fact that gene databases contain  a great deal of aberrant data. Gene-annotation programs like the widely used Glimmer (Gene Locator and Interpolated Markov Modeler) have to be trained, using a training set. If the training set contains faulty data, it's a classic GIGO situation.

The annotation accuracy problem is getting worse by the year (see graph above). Devos and Valencia estimated in 2001 that misannotation levels could be as high as 37%. More recently, Schnoes et al. (2009) concluded that "function prediction error (i.e., misannotation) is a serious problem in all but the manually curated database Swiss-Prot," and yet Artamonova et al. (2005) found that for five types of annotation entries, even the vaunted UniProt/SwissProt database had an error rate between 33% and 43%. So even the best manually curated database is full of errors.

I've spent many hours examining bacterial genomes and it's been my experience that annotation quality is uniformly poor for all but the best-known genes in the best-curated genomes. By "best-known genes," I mean things like ribosomal protein genes, genes for well-known polymerases and chaperones, well-studied metabolic-pathway genes, and so on. The vast majority of genes encode proteins that have not been studied (and may never be studied) experimentally. All of these are suspect and need to be treated with caution, particularly in high-GC genomes where programs like Glimmer frequently can't distinguish between sense and antisense strands.

As an example: Pseudomonas aeruginosa MPAO1 has a gene, O1Q_25367, encoding a phenol hydroxylase, which is an enzyme required for catabolic breakdown of phenols (the kind of thing many Pseudomonas species are good at). If you run a BLAST search of this gene against the UnitProt.org database, you'll find dozens of good-quality hits against phenol hydroxylases of many organisms (including Rhizobium sp. Pop5, Natronolimnobius innermongolicus JCM 12255, Rhodococcus wratislaviensis IFP 2016, and quite a few others). "Good-quality" here means more than 50% amino acid sequence identities, with E-values of ~10-50. But there's a problem. If you take the DNA sequence for the Pseudomonas phenol hydroxylase and reverse-translate it using the online app at http://web.expasy.org/translate, the resulting protein sequence is a 99% or better match (E=0) for over a hundred Pseudomonas maltose/mannitol transporters. (It's a 100% match in 85 cases.) In other words, the so-called phenol hydroxylase gene is not a phenol hydroxylase gene at all. It's a sugar transporter, backwards. You can do the same trick with the "phenol hydroxylase" of Gordonia terrae C6. Reverse-translate it and it's a much better sugar permease (dozens of E=0 matches) than it is a phenol hydroxylase.

This view of a segment of P. aeruginosa and Streptomyces genomes shows a region of 60% homology (pink band) between the two genomes for two genes. The yellow gene is each case is a "phenol hydroxylase."

How does a program like Glimmer not catch things like this? In this case, Glimmer became confused by an apparent overlapping-gene situation (above). The program found two open reading frames, in the same part of the chromosome, but on opposite strands. Glimmer 3 is supposedly much better at resolving overlaps, but Glimmer 2 (which was used to annotate around half the genomes currently available in public databases) relied on unbelievably crude heuristics to resolve overlap problems. In the above case (perhaps through operator error; these things are configurable, to a degree) Glimmer simply designated each ORF as a legitimate gene, when in fact a BLAST check leaves little doubt the phenol hydroxylase gene is (in reality) a backwards maltose/mannose transporter gene.

My experience has been that Glimmer is very easily confused by high-GC genome data, with organisms like Pseudomonas (typically 63% to 67% GC) showing a greater number of reverse-annotated genes (and no-BLAST-matches "hypothetical proteins") than low-GC organisms like Buchnera, although it should be noted that even E. coli strains tend to contain many suspect annotations. The problem is partly due to the fact that high-GC-content DNA contains relatively few stop codons in the six different reading frames, compared to low-GC DNA (which is rich in stop codons if deciphered the wrong way). Also a problem for Glimmer is the fact that in high-GC genomes, codons tend to contain a purine in the first base position whether they're read forward or in reverse.

I've written before about the fact that codons tend to occur with frequencies roughly equal to that of their reverse-complement twins. This is another confounding factor for Glimmer (codons look the same whether read in the forward direction or the reverse direction), although honestly, I'm beginning to wonder if the codon/anticodon symmetries I've been seeing aren't simply due to widespread reverse annotation of genes (misannotation of a gene's anticoding stand, as with "phenol hydroxylase").

One might naively suppose that it shouldn't be hard to discriminate the "sense" strand accurately, given that so many genes begin with a Shine Dalgarno sequence ahead of the start codon. But in reality, it turns out that not very many organisms make extensive use of SD sequences. Short motifs like GAGG, which occur randomly at a high rate, can be mistaken for a Shine Dalgarno sequence. This only aggravates the false-positive rate.

Erroneous gene annotations are rampant in bacteria, but some authors have suggested that the problem is far worse in eukaryotic genomes. If that's true, we're in trouble.

All of this puts bioinformatics research at a crisis point. Gene discovery algorithms are good (maybe a bit too good) at finding open reading frames but poor at identifying and assigning gene functions, in part because we lack the kind of in-depth understanding of protein folding required for prediction of 3-dimensional structures (and active-site conformations) programmatically, a capability that's sorely needed if we're to progress out of the gene-annotation Stone Age we're now living in. (Protein folding is a Hard Problem requiring supercomputers to untangle.) Maybe in ten years (or twenty?), we'll be able to predict 3D protein structures computationally, and on that basis make better ab initio predictions of protein function. Right now, we have to make do with relatively crude Markov-model pattern recognition software, aided by human intervention, to come up with even a minimally reliable genome annotation. But we have the means, already, of doing much better crosschecks. Some of that can be automated. We just need to have the will to do it.

Thursday, April 24, 2014

Are Overapping Genes Real?

Bacteria belonging to the Pseudomonas family are a perennial favorite among bacteriology instructors (and students) because of the curious ability of some of its members to produce pigments that fluoresce under an ultraviolet light. If you're unlucky enough to get an infected cut on the arm while working in the garden, it's possible your cut will fluoresce under a black light. That's enough of a diagnosis to pronounce the infectious agent. 
Fluorescent colonies of Pseudomonas.

Silby and Levy, investigating the adaptation of the bacterium Pseudomonas fluorescens to soil, uncovered the existence of at least ten antisense genes in P. fluorescens. They went on to demonstrate experimentally that one of the genes, cosA, produces not just antisense RNA but an associated protein. Tellingly, Silby and Levy commented:
These findings suggest that current genome annotations provide an incomplete view of the genetic potential of a given organism.
The implication is that additional antitranscriptome genes remain to be found, not only in Pseudomonas but in other organisms.

There's a good reason they haven't been found yet. Overlapping genes are automatically rejected by many of the annotation programs that are commonly used to find, identify, and label genes in genome sequences. (The oft-used freeware Glimmer 2 program allows you to set the overlap-rejection threshold.) Many yet-to-be-discovered antisense genes have been deliberately and systematically obscured in published genomes.

Still, once in a while such genes do surface. For example, in Pseudomonas stutzeri A1501, we find a pair of overlapping genes at an offset of 3035137 on the chromosome (see illustration below).

Overapping genes in Pseudomonas stutzeri.

The top gene is annotated merely as a "hypothetical protein," while the underlying gene on the opposite strand is an aspartyl-tRNA synthetase. One's normal inclination is to dismiss a hypothetical protein as being unimportant, but this may not be wise. Twenty percent or more of bacterial genes are annotated as hypothetical proteins; common sense says they can't all be unimportant. In fact, in "Transcriptome Analysis of Pseudomonas syringae Identifies New Genes, Noncoding RNAs, and Antisense Activity" by Filiatrault et al. (2010), researchers found that 818 out of 1,646 protein genes in P. syringae annotated as "hypothetical proteins" were expressed under iron-limited conditions. Many (probably most) genes annotated as "hypothetical protein" are quite real and should probably be re-annotated as PUF: "protein of unknown function."

In this case, the "hypothetical protein" shown in yellow (above) turns up medium-strength protein-BLAST hits with other "hypothetical proteins" from other organisms, including a hit with an E-value of 3.0×10-49 in Parasutterella excrementihominis YIT 11859 and a comparable hit on a predicted phosphatase/phosphohexomutase in Rothia mucilaginosa DY-18.

In this particular case, the hypothetical-protein gene lacks a strong upstream Shine Dalgarno sequence (a sequence preceding many genes that helps bind a ribosome to the mRNA). But so too does the gene on the opposite strand. (This is not unusual. The SD sequence is not required for translation and in fact, in about half of bacterial species, a Shine Dalgarno sequence is associated with fewer than 50% of genes.) Hence, the jury's out on whether the antigene is expressed. It could be that no protein is made from the top strand but the gene provides RNA-mediated control of the gene on the bottom strand. We won't know for sure until someone investigates.

In Pseudomonas aeruginosa strain PADK2_CF510, we find another instance of a bidirectional overlapping gene pair (see graphic below). In this case, the gene on the top strand (CF510_06030) encodes the large subunit of an isopropylmalate isomerase. The gene on the bottom strand (CF510_06025, shown in yellow) is annotated as "Flp pilus assembly protein TadG." It could very well be a misannotated non-gene. However, five genes away is FimV (CF510_06060), another pilus-assembly (motility) protein. Moreover, the gene marked TadG has a strong upstream SD sequence containing the canonical GGAGG motif. The gene above it has a weaker GGAAA motif.

P. aeruginosa has an overlap of an isopropylmalate isomerase gene and a gene for a motility protein. The latter is shown in yellow.

In previous posts, I've mentioned (and shown data for) the fact that in the overwhelming majority of protein-encoding genes (across every kind of genome), the first base of a codon tends to be purine-rich. One check of whether a bidi-overlap gene is "real" or not ought to be that the first codon base should be purine rich in both reading directions. This is, in fact, the case for the examples shown above. The aspartyl-tRNA synthetase gene for P. stutzeri has AG1 (1st base, purine) content averaging 59.8%, whereas its bidirectional partner gene ("hypothetical protein") has AG1 = 58.5%. The isopropylmalate isomerase of P. aeruginosa has AG1 = 65.9%, while its antisymmetric partner (TadG) has AG1 = 56.2%.

If you enjoyed this post, please give the URL to your biogeek friends. Thanks!

Tuesday, April 22, 2014

How Antisense Genes Are Discovered

In the past ten years or so, a great deal of research has focused on antisense transcription of genes. Normally, RNA gets transcribed from one strand of DNA only. But it turns out, in many cases RNA also gets transcribed off the opposite strand of DNA (an antisense copy), either at the original gene (so-called cis transcription) or at a copy of the gene some distance away (trans transcription). The latter can be a pseudogene, or a normal copy of the gene.

Antisense transcripts occur very widely not only in human DNA but in bacteria, yeast, and (in fact) every place where scientists have looked, and places where they haven't looked. Some of the most interesting discoveries have happened when researchers weren't specifically looking for antisense transcripts but found them by accident. How does that happen? It happens in experiments involving IVET (in vivo expression technology), an important experimental technique for uncovering new genes.

IVET is a powerful gene manipulation strategy for discovering which genes in an organism (a pathogen, usually) are up-regulated or turned on during host infection. Let's say you're studying a new pathogen and you want to get an idea of which genes, in the pathogen, are turned on during the infection process. First, you need a strain of the organism that's disabled by virtue of lacking a working copy of a particular metabolic enzyme, say an enzyme needed for purine metabolism, e.g. purA. Secondly, you need a vector for inserting a promoterless copy of the working gene into the bacterium. What this usually means is, you need a plasmid (a small extra chromosome; many bacteria have them, and they can often be manipulated in the lab) on which to place a functional purA gene. The gene won't be expressed, however, if it lacks a suitable promoter region on the DNA upstream of the gene. That's good; that's what you want. You want to put a promoterless copy of the good gene on the plasmid, along with (this is crucial) a random chunk of DNA from the pathogen, inserted ahead of purA on the plasmid. In practice, it's easy to create a bunch of plasmids with this arrangement: a working copy of purA, and ahead of it, a random chunk of pathogen DNA. The idea is that you now attempt to infect a lab animal with the bacterium containing the plasmid. If the bacterium establishes infection in the animal, presumably it's because a random chunk of DNA happened to contain a promoter region (and associated downstream genes) that gets turned on during infection. If you now isolate the bacterium from the sick animal, you can look to see what kind(s) of genes got transduced into the bacterium.

IVET is a promoter trap technology for selecting bacterial genes that are specifically induced when bacteria infect a host organism. A plasmid vetor contains a random fragment of the chromosome of the pathogen (red) and a promoterless gene (selective marker, burgundy) that encodes an enzyme required for survival. Pooled plasmid-containing clones are inoculated into the mouse (B). Only those bacteria that contain the selective marker fused to a random gene that is transcriptionally active in the host are able to survive. After a suitable infection period, bacteria that express the marker are isolated from the spleen or other organs. The inclusion of a lacZY mutant gene (blue) allows post-selection screening for promoters that are active only in vivo. What you want are bacteria that are lac-positive only in the host environment, not "constitutive" (always-on).
Exactly this sort of technique was used by Silby, Rainey, and Levy to determine which genes were activated in Pseudomonas during colonization of soil. (The IVET technique can be adapted to any scenario in which an organism differentially expresses genes in its adaptation to a "host" environment, even if the environment is, in fact, a plant, or soil in this case, rather than a mouse.) They were looking to see which genes in Pseudomonas play an essential role in that organism's ability to thrive in soil, and they successfully identified more than 50 promoters (and associated fusions) that come alive during soil colonization. When they looked at 22 "soil genes" that got turned on, they found ten previously undescribed genes that were transcribed in the antisense direction from regions overlapping known genes. They called these ten genes "cryptic fusions" because of their un-annotated existence on the supposedly silent, antisense side of known genes.

Cryptic fusions discovered by Silby et al. are shown in grey, in their antisense orientation to known genes (darker grey).

It's not unusual to find that antisense transcripts are playing a regulatory role. When a gene gets transcribed in both directions, the resulting sense and antisense RNAs can combine (by Watson-Crick pairing) to form a double-stranded RNA product, preventing translation of the RNA into protein. But incredibly, sometimes an antisense RNA transcript encodes a legitimate protein (a protein that gets made off the antisense copy). Silby and Levy documented this for the previously unknown cosA gene in Pseudomonas. It seems likely additional antisense proteins await discovery. (Most studies stop at the level of identifying RNA products.)

The finding of antisense transcripts in IVET experiments is common. One of the authors of the Pseudomonas study (Rainey) had previously published a study of rhizosphere-induced genes in Pseudomonas but had not published the fact that 20% of genes found this way were in an antisense orientation to normal genes. Likewise, a 1996 study of Pseudomonas aeruginosa infection in the mouse (Pseudomonas is an opportunistic pathogen) found antisense activity. In fact, the first-ever paper on IVET (by Mahan et al., 1993) described finding antiscript products.

IVET has uncovered a previously unknown "antitranscriptome" world hidden inside living cells. Until we explore this world fully, we won't know how much undiscovered biology we've left on the table.