Finding novel genes in bacterial communities isolated from the environment
1 Bielefeld University, Center for Biotechnology (CeBiTec) D-33594 Bielefeld Germany
2 Fellowship for Interpretation of Genomes Burr Ridge IL
3 Department of Biology, San Diego State University San Diego, CA
4 Center for Microbial Sciences San Diego, CA
5 Universität Bielefeld, Technische Fakultät D-33594 Bielefeld Germany
6 Universität Bielefeld, Lehrstuhl für Genetik, Fakultät für Biologie D-33594 Bielefeld Germany
*To whom correspondence should be addressed.
Motivation: Novel sequencing techniques can give access to organisms that are difficult to cultivate using conventional methods. When applied to environmental samples, the data generated has some drawbacks, e.g. short length of assembled contigs, in-frame stop codons and frame shifts. Unfortunately, current gene finders cannot circumvent these difficulties. At the same time, the automated prediction of genes is a prerequisite for the increasing amount of genomic sequences to ensure progress in metagenomics.
Results: We introduce a novel gene finding algorithm that incorporates features overcoming the short length of the assembled contigs from environmental data, in-frame stop codons as well as frame shifts contained in bacterial sequences. The results show that by searching for sequence similarities in an environmental sample our algorithm is capable of detecting a high fraction of its gene content, depending on the species composition and the overall size of the sample. The method is valuable for hunting novel unknown genes that may be specific for the habitat where the sample is taken. Finally, we show that our algorithm can even exploit the limited information contained in the short reads generated by 454 technology for the prediction of protein coding genes.
Availability: The program is freely available upon request.
Contact: Lutz.Krause{at}CeBiTec.Uni-Bielefeld.DE