Skip Navigation

This Article
Right arrow FREE Full Text (Print PDF) Freely available
Right arrow FREE Full Text (Screen PDF)
Right arrow Alert me when this article is cited
Right arrow Alert me if a correction is posted
Services
Right arrow Email this article to a friend
Right arrow Similar articles in this journal
Right arrow Similar articles in ISI Web of Science
Right arrow Similar articles in PubMed
Right arrow Alert me to new issues of the journal
Right arrow Add to My Personal Archive
Right arrow Download to citation manager
Right arrow Search for citing articles in:
ISI Web of Science (74)
Right arrowRequest Permissions
Google Scholar
Right arrow Articles by Enright, A. J.
Right arrow Articles by Ouzounis, C. A.
Right arrow Search for Related Content
PubMed
Right arrow PubMed Citation
Right arrow Articles by Enright, A. J.
Right arrow Articles by Ouzounis, C. A.
Social Bookmarking
 Add to CiteULike   Add to Connotea   Add to Del.icio.us  
What's this?

Bioinformatics Vol. 16 no. 5 2000
Pages 451-457
© 2000 Oxford University Press

GeneRAGE: a robust algorithm for sequence clustering and domain detection

Anton J. Enright 1 and Christos A. Ouzounis 1,*

1 Computational Genomics Group, Research Programme, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK

Received on January 27, 2000 ; revised on March 13, 2000 ; accepted on March 13, 2000

Motivation: Efficient, accurate and automatic clustering of large protein sequence datasets, such as complete proteomes, into families, according to sequence similarity. Detection and correction of false positive and negative relationships with subsequent detection and resolution of multi-domain proteins.

Results: A new algorithm for the automatic clustering of protein sequence datasets has been developed. This algorithm represents all similarity relationships within the dataset in a binary matrix. Removal of false positives is achieved through subsequent symmetrification of the matrix using a Smith–Waterman dynamic programming alignment algorithm. Detection of multi-domain protein families and further false positive relationships within the symmetrical matrix is achieved through iterative processing of matrix elements with successive rounds of Smith–Waterman dynamic programming alignments. Recursive single-linkage clustering of the corrected matrix allows efficient and accurate family representation for each protein in the dataset. Initial clusters containing multi-domain families, are split into their constituent clusters using the information obtained by the multi-domain detection step. This algorithm can hence quickly and accurately cluster large protein datasets into families. Problems due to the presence of multi-domain proteins are minimized, allowing more precise clustering information to be obtained automatically.

Availability: GeneRAGE (version 1.0) executable binaries for most platforms may be obtained from the authors on request. The system is available to academic users free of charge under license.

Contact: ouzounis{at}ebi.ac.uk

* To whom correspondence should be addressed.


Add to CiteULike CiteULike   Add to Connotea Connotea   Add to Del.icio.us Del.icio.us    What's this?


This article has been cited by other articles:


Home page
BioinformaticsHome page
D. P. Brown
Efficient functional clustering of protein sequences using the Dirichlet process
Bioinformatics, August 15, 2008; 24(16): 1765 - 1771.
[Abstract] [Full Text] [PDF]


Home page
Brief BioinformHome page
S. J. Sammut, R. D. Finn, and A. Bateman
Pfam 10 years on: 10 000 families and still growing
Brief Bioinform, May 1, 2008; 9(3): 210 - 219.
[Abstract] [Full Text] [PDF]


Home page
BioinformaticsHome page
B. E. Suzek, H. Huang, P. McGarvey, R. Mazumder, and C. H. Wu
UniRef: comprehensive and non-redundant UniProt reference clusters
Bioinformatics, May 15, 2007; 23(10): 1282 - 1288.
[Abstract] [Full Text] [PDF]


Home page
BioinformaticsHome page
M. Nikolski and D. J. Sherman
Family relationships: should consensus reign?--consensus clustering for protein families
Bioinformatics, January 15, 2007; 23(2): e71 - e76.
[Abstract] [Full Text] [PDF]


Home page
Protein Sci.Home page
A. Oberai, Y. Ihm, S. Kim, and J. U. Bowie
A limited universe of membrane protein families and folds.
Protein Sci., July 1, 2006; 15(7): 1723 - 1734.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
A. Paccanaro, J. A. Casbon, and M. A. S. Saqi
Spectral clustering of protein sequences
Nucleic Acids Res., March 17, 2006; 34(5): 1571 - 1580.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
R. L. Marsden, D. Lee, M. Maibaum, C. Yeats, and C. A. Orengo
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space
Nucleic Acids Res., February 15, 2006; 34(3): 1066 - 1080.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
I. Uchiyama
Hierarchical clustering algorithm for comprehensive orthologous-domain classification in multiple genomes
Nucleic Acids Res., January 25, 2006; 34(2): 647 - 658.
[Abstract] [Full Text] [PDF]


Home page
BioinformaticsHome page
L. Goldovsky, P. Janssen, D. Ahren, B. Audit, I. Cases, N. Darzentas, A. J. Enright, N. Lopez-Bigas, J. M. Peregrin-Alvarez, M. Smith, et al.
CoGenT++: an extensive and extensible data environment for computational genomics
Bioinformatics, October 1, 2005; 21(19): 3806 - 3810.
[Abstract] [Full Text] [PDF]


Home page
BioinformaticsHome page
J. E. Donald and E. I. Shakhnovich
Determining functional specificity from protein sequences
Bioinformatics, June 1, 2005; 21(11): 2629 - 2635.
[Abstract] [Full Text] [PDF]


Home page
Plant Physiol.Home page
K. Horan, J. Lauricha, J. Bailey-Serres, N. Raikhel, and T. Girke
Genome Cluster Database. A Sequence Family Analysis Platform for Arabidopsis and Rice
Plant Physiology, May 1, 2005; 138(1): 47 - 54.
[Abstract] [Full Text] [PDF]


Home page
Genome ResHome page
N. G. Faux, S. P. Bottomley, A. M. Lesk, J. A. Irving, J. R. Morrison, M. G. de la Banda, and J. C. Whisstock
Functional insights from the distribution and role of homopeptide repeat-containing proteins
Genome Res., April 1, 2005; 15(4): 537 - 551.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
L. Y. Han, C. Z. Cai, Z. L. Ji, Z. W. Cao, J. Cui, and Y. Z. Chen
Predicting functional family of novel enzymes irrespective of sequence similarity: a statistical learning approach
Nucleic Acids Res., December 7, 2004; 32(21): 6437 - 6444.
[Abstract] [Full Text] [PDF]


Home page
Proc. Natl. Acad. Sci. USAHome page
D. Lindell, M. B. Sullivan, Z. I. Johnson, A. C. Tolonen, F. Rohwer, and S. W. Chisholm
Transfer of photosynthesis genes to and from Prochlorococcus viruses
PNAS, July 27, 2004; 101(30): 11013 - 11018.
[Abstract] [Full Text] [PDF]


Home page
Genome ResHome page
V. Veeramachaneni and W. Makalowski
Visualizing Sequence Similarity of Protein Families
Genome Res., June 1, 2004; 14(6): 1160 - 1169.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
A. J. Enright, V. Kunin, and C. A. Ouzounis
Protein families and TRIBES in genome sequence space
Nucleic Acids Res., August 1, 2003; 31(15): 4632 - 4638.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
C.Z. Cai, L.Y. Han, Z.L. Ji, X. Chen, and Y.Z. Chen
SVM-Prot: web-based support vector machine software for functional classification of a protein from its primary sequence
Nucleic Acids Res., July 1, 2003; 31(13): 3692 - 3697.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
S. Mika and B. Rost
UniqueProt: creating representative protein sequence sets
Nucleic Acids Res., July 1, 2003; 31(13): 3789 - 3791.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
R. M. R. Coulson and C. A. Ouzounis
The phylogenetic diversity of eukaryotic transcription
Nucleic Acids Res., January 15, 2003; 31(2): 653 - 660.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
A. J. Enright, S. Van Dongen, and C. A. Ouzounis
An efficient algorithm for large-scale detection of protein families
Nucleic Acids Res., April 1, 2002; 30(7): 1575 - 1584.
[Abstract] [Full Text] [PDF]


Home page
Protein Eng Des SelHome page
D. Frishman
Knowledge-based selection of targets for structural genomics
Protein Eng. Des. Sel., March 1, 2002; 15(3): 169 - 183.
[Abstract] [Full Text] [PDF]


Home page
Nucleic Acids ResHome page
P. J. Janssen, B. Audit, and C. A. Ouzounis
Strain-specific genes of Helicobacter pylori: distribution, function and dynamics
Nucleic Acids Res., November 1, 2001; 29(21): 4395 - 4404.
[Abstract] [Full Text] [PDF]


Home page
Mol Biol EvolHome page
N. Wicker, G. Rene Perrin, J. C. Thierry, and O. Poch
Secator: A Program for Inferring Protein Subfamilies from Phylogenetic Trees
Mol. Biol. Evol., August 1, 2001; 18(8): 1435 - 1441.
[Abstract] [Full Text] [PDF]


Home page
Genome ResHome page
A. Louis, E. Ollivier, J.-C. Aude, and J.-L. Risler
Massive Sequence Comparisons as a Help in Annotating Genomic Sequences
Genome Res., July 1, 2001; 11(7): 1296 - 1303.
[Abstract] [Full Text] [PDF]


Home page
Genome ResHome page
S. Tsoka and C. A. Ouzounis
Functional Versatility and Molecular Diversity of the Metabolic Map of Escherichia coli
Genome Res., September 1, 2001; 11(9): 1503 - 1510.
[Abstract] [Full Text] [PDF]


Home page
Genome ResHome page
B. Snel, P. Bork, and M. A. Huynen
Genomes in Flux: The Evolution of Archaeal and Proteobacterial Gene Content
Genome Res., January 1, 2002; 12(1): 17 - 25.
[Abstract] [Full Text] [PDF]



Disclaimer:
Please note that abstracts for content published before 1996 were created through digital scanning and may therefore not exactly replicate the text of the original print issues. All efforts have been made to ensure accuracy, but the Publisher will not be held responsible for any remaining inaccuracies. If you require any further clarification, please contact our Customer Services Department.