Bioinformatics Vol. 16 no. 5 2000
Pages 451-457
© 2000 Oxford University Press
GeneRAGE: a robust algorithm for sequence clustering and domain detection
1 Computational Genomics Group, Research Programme, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK
Received on January 27, 2000
; revised on March 13, 2000
; accepted on March 13, 2000
Motivation: Efficient, accurate and automatic clustering of large protein sequence datasets, such as complete proteomes, into families, according to sequence similarity. Detection and correction of false positive and negative relationships with subsequent detection and resolution of multi-domain proteins.
Results: A new algorithm for the automatic clustering of protein sequence datasets has been developed. This algorithm represents all similarity relationships within the dataset in a binary matrix. Removal of false positives is achieved through subsequent symmetrification of the matrix using a SmithWaterman dynamic programming alignment algorithm. Detection of multi-domain protein families and further false positive relationships within the symmetrical matrix is achieved through iterative processing of matrix elements with successive rounds of SmithWaterman dynamic programming alignments. Recursive single-linkage clustering of the corrected matrix allows efficient and accurate family representation for each protein in the dataset. Initial clusters containing multi-domain families, are split into their constituent clusters using the information obtained by the multi-domain detection step. This algorithm can hence quickly and accurately cluster large protein datasets into families. Problems due to the presence of multi-domain proteins are minimized, allowing more precise clustering information to be obtained automatically.
Availability: GeneRAGE (version 1.0) executable binaries for most platforms may be obtained from the authors on request. The system is available to academic users free of charge under license.
Contact: ouzounis{at}ebi.ac.uk
* To whom correspondence should be addressed.
![]()
CiteULike
Connotea
Del.icio.us What's this?
This article has been cited by other articles:
![]() |
D. P. Brown Efficient functional clustering of protein sequences using the Dirichlet process Bioinformatics, August 15, 2008; 24(16): 1765 - 1771. [Abstract] [Full Text] [PDF] |
||||
![]() |
S. J. Sammut, R. D. Finn, and A. Bateman Pfam 10 years on: 10 000 families and still growing Brief Bioinform, May 1, 2008; 9(3): 210 - 219. [Abstract] [Full Text] [PDF] |
||||
![]() |
B. E. Suzek, H. Huang, P. McGarvey, R. Mazumder, and C. H. Wu UniRef: comprehensive and non-redundant UniProt reference clusters Bioinformatics, May 15, 2007; 23(10): 1282 - 1288. [Abstract] [Full Text] [PDF] |
||||
![]() |
M. Nikolski and D. J. Sherman Family relationships: should consensus reign?--consensus clustering for protein families Bioinformatics, January 15, 2007; 23(2): e71 - e76. [Abstract] [Full Text] [PDF] |
||||
![]() |
A. Oberai, Y. Ihm, S. Kim, and J. U. Bowie A limited universe of membrane protein families and folds. Protein Sci., July 1, 2006; 15(7): 1723 - 1734. [Abstract] [Full Text] [PDF] |
||||
![]() |
A. Paccanaro, J. A. Casbon, and M. A. S. Saqi Spectral clustering of protein sequences Nucleic Acids Res., March 17, 2006; 34(5): 1571 - 1580. [Abstract] [Full Text] [PDF] |
||||
![]() |
R. L. Marsden, D. Lee, M. Maibaum, C. Yeats, and C. A. Orengo Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space Nucleic Acids Res., February 15, 2006; 34(3): 1066 - 1080. [Abstract] [Full Text] [PDF] |
||||
![]() |
I. Uchiyama Hierarchical clustering algorithm for comprehensive orthologous-domain classification in multiple genomes Nucleic Acids Res., January 25, 2006; 34(2): 647 - 658. [Abstract] [Full Text] [PDF] |
||||
![]() |
L. Goldovsky, P. Janssen, D. Ahren, B. Audit, I. Cases, N. Darzentas, A. J. Enright, N. Lopez-Bigas, J. M. Peregrin-Alvarez, M. Smith, et al. CoGenT++: an extensive and extensible data environment for computational genomics Bioinformatics, October 1, 2005; 21(19): 3806 - 3810. [Abstract] [Full Text] [PDF] |
||||
![]() |
J. E. Donald and E. I. Shakhnovich Determining functional specificity from protein sequences Bioinformatics, June 1, 2005; 21(11): 2629 - 2635. [Abstract] [Full Text] [PDF] |
||||
![]() |
K. Horan, J. Lauricha, J. Bailey-Serres, N. Raikhel, and T. Girke Genome Cluster Database. A Sequence Family Analysis Platform for Arabidopsis and Rice Plant Physiology, May 1, 2005; 138(1): 47 - 54. [Abstract] [Full Text] [PDF] |
||||
![]() |
N. G. Faux, S. P. Bottomley, A. M. Lesk, J. A. Irving, J. R. Morrison, M. G. de la Banda, and J. C. Whisstock Functional insights from the distribution and role of homopeptide repeat-containing proteins Genome Res., April 1, 2005; 15(4): 537 - 551. [Abstract] [Full Text] [PDF] |
||||
![]() |
L. Y. Han, C. Z. Cai, Z. L. Ji, Z. W. Cao, J. Cui, and Y. Z. Chen Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach Nucleic Acids Res., December 7, 2004; 32(21): 6437 - 6444. [Abstract] [Full Text] [PDF] |
||||
![]() |
D. Lindell, M. B. Sullivan, Z. I. Johnson, A. C. Tolonen, F. Rohwer, and S. W. Chisholm Transfer of photosynthesis genes to and from Prochlorococcus viruses PNAS, July 27, 2004; 101(30): 11013 - 11018. [Abstract] [Full Text] [PDF] |
||||
![]() |
V. Veeramachaneni and W. Makalowski Visualizing Sequence Similarity of Protein Families Genome Res., June 1, 2004; 14(6): 1160 - 1169. [Abstract] [Full Text] [PDF] |
||||
![]() |
A. J. Enright, V. Kunin, and C. A. Ouzounis Protein families and TRIBES in genome sequence space Nucleic Acids Res., August 1, 2003; 31(15): 4632 - 4638. [Abstract] [Full Text] [PDF] |
||||
![]() |
C.Z. Cai, L.Y. Han, Z.L. Ji, X. Chen, and Y.Z. Chen SVM-Prot: web-based support vector machine software for functional classification of a protein from its primary sequence Nucleic Acids Res., July 1, 2003; 31(13): 3692 - 3697. [Abstract] [Full Text] [PDF] |
||||
![]() |
S. Mika and B. Rost UniqueProt: creating representative protein sequence sets Nucleic Acids Res., July 1, 2003; 31(13): 3789 - 3791. [Abstract] [Full Text] [PDF] |
||||
![]() |
R. M. R. Coulson and C. A. Ouzounis The phylogenetic diversity of eukaryotic transcription Nucleic Acids Res., January 15, 2003; 31(2): 653 - 660. [Abstract] [Full Text] [PDF] |
||||
![]() |
A. J. Enright, S. Van Dongen, and C. A. Ouzounis An efficient algorithm for large-scale detection of protein families Nucleic Acids Res., April 1, 2002; 30(7): 1575 - 1584. [Abstract] [Full Text] [PDF] |
||||
![]() |
D. Frishman Knowledge-based selection of targets for structural genomics Protein Eng. Des. Sel., March 1, 2002; 15(3): 169 - 183. [Abstract] [Full Text] [PDF] |
||||
![]() |
P. J. Janssen, B. Audit, and C. A. Ouzounis Strain-specific genes of Helicobacter pylori: distribution, function and dynamics Nucleic Acids Res., November 1, 2001; 29(21): 4395 - 4404. [Abstract] [Full Text] [PDF] |
||||
![]() |
N. Wicker, G. Rene Perrin, J. C. Thierry, and O. Poch Secator: A Program for Inferring Protein Subfamilies from Phylogenetic Trees Mol. Biol. Evol., August 1, 2001; 18(8): 1435 - 1441. [Abstract] [Full Text] [PDF] |
||||
![]() |
A. Louis, E. Ollivier, J.-C. Aude, and J.-L. Risler Massive Sequence Comparisons as a Help in Annotating Genomic Sequences Genome Res., July 1, 2001; 11(7): 1296 - 1303. [Abstract] [Full Text] [PDF] |
||||
![]() |
S. Tsoka and C. A. Ouzounis Functional Versatility and Molecular Diversity of the Metabolic Map of Escherichia coli Genome Res., September 1, 2001; 11(9): 1503 - 1510. [Abstract] [Full Text] [PDF] |
||||
![]() |
B. Snel, P. Bork, and M. A. Huynen Genomes in Flux: The Evolution of Archaeal and Proteobacterial Gene Content Genome Res., January 1, 2002; 12(1): 17 - 25. [Abstract] [Full Text] [PDF] |
||||








