ProbCons: Probabilistic consistency-based multiple sequence alignment - PubMed (original) (raw)

Comparative Study

ProbCons: Probabilistic consistency-based multiple sequence alignment

Chuong B Do et al. Genome Res. 2005 Feb.

Abstract

To study gene evolution across a wide range of organisms, biologists need accurate tools for multiple sequence alignment of protein families. Obtaining accurate alignments, however, is a difficult computational problem because of not only the high computational cost but also the lack of proper objective functions for measuring alignment quality. In this paper, we introduce probabilistic consistency, a novel scoring function for multiple sequence comparisons. We present ProbCons, a practical tool for progressive protein multiple sequence alignment based on probabilistic consistency, and evaluate its performance on several standard alignment benchmark data sets. On the BAliBASE, SABmark, and PREFAB benchmark alignment databases, ProbCons achieves statistically significant improvement over other leading methods while maintaining practical speed. ProbCons is publicly available as a Web resource.

PubMed Disclaimer

Figures

Figure 1.

Basic pair-HMM for sequence alignment between two sequences, x and y. State M emits two letters, one from each sequence, and corresponds to the two letters being aligned together. State Ix emits a letter in sequence x that is aligned to a gap, and similarly state Iy emits a letter in sequence y that is aligned to a gap. Finding the most likely alignment according to this model by using the Viterbi algorithm corresponds to applying Needleman-Wunsch with appropriate parameters. The logarithm of the emission probability function p(.,.) at M corresponds to a substitution scoring matrix, while affine gap penalty parameters can be derived from the transition probabilities δ and ε (Durbin et al. 1998).

Figure 2.

Column reliability plot for 1csy_ref1 from BAliBASE, Reference 1. The red line and solid regions indicate the predicted and actual proportion of correct pairwise matches at each alignment position, respectively. All column reliability values have been multiplied by 100. Below, the actual ProbCons alignment is shown with core block residues highlighted in green. Note that only pairwise matches in core block regions of the BAliBASE alignment are considered correct when computing the “actual” proportion of correct pairwise matches; however, some residues outside of the core block regions may also be alignable. Thus, regions in which predicted homology exceeds actual homology do not necessarily indicate overprediction of homology by the aligner.

Cited by

Three duplication events and variable molecular evolution characteristics involved in multiple GGPS genes of six Solanaceae species.
Li F, Wei CY, Qiao C, Chen Z, Wang P, Wei P, Wang R, Jin L, Yang J, Lin F, Luo Z. Li F, et al. J Genet. 2016 Jun;95(2):453-7. doi: 10.1007/s12041-016-0634-1. J Genet. 2016. PMID: 27350691 No abstract available.
Dynamically evolving novel overlapping gene as a factor in the SARS-CoV-2 pandemic.
Nelson CW, Ardern Z, Goldberg TL, Meng C, Kuo CH, Ludwig C, Kolokotronis SO, Wei X. Nelson CW, et al. Elife. 2020 Oct 1;9:e59633. doi: 10.7554/eLife.59633. Elife. 2020. PMID: 33001029 Free PMC article.
The Escherichia coli RlmN methyltransferase is a dual-specificity enzyme that modifies both rRNA and tRNA and controls translational accuracy.
Benítez-Páez A, Villarroya M, Armengod ME. Benítez-Páez A, et al. RNA. 2012 Oct;18(10):1783-95. doi: 10.1261/rna.033266.112. Epub 2012 Aug 13. RNA. 2012. PMID: 22891362 Free PMC article.
GUIDANCE2: accurate detection of unreliable alignment regions accounting for the uncertainty of multiple parameters.
Sela I, Ashkenazy H, Katoh K, Pupko T. Sela I, et al. Nucleic Acids Res. 2015 Jul 1;43(W1):W7-14. doi: 10.1093/nar/gkv318. Epub 2015 Apr 16. Nucleic Acids Res. 2015. PMID: 25883146 Free PMC article.
Structural and molecular basis of interaction of HCV non-structural protein 5A with human casein kinase 1α and PKR.
Sudha G, Yamunadevi S, Tyagi N, Das S, Srinivasan N. Sudha G, et al. BMC Struct Biol. 2012 Nov 13;12:28. doi: 10.1186/1472-6807-12-28. BMC Struct Biol. 2012. PMID: 23148689 Free PMC article.

References

1. Altschul, S.F. 1991. Amino acid substitution matrices from an information theoretic perspective. J. Mol. Biol. 219: 555-565. - PMC - PubMed
1. Altschul, S.F., Carroll, R.J., and Lipman, D.J. 1989. Weights for data related by a tree. J. Mol. Biol. 207: 647-653. - PubMed
1. Altschul, S.F., Madden, T.L., Schaffer, A.A., Zhang, J., Zhang, Z., Miller, W., and Lipman, D.J. 1997. Gapped BLAST and PSI-BLAST: A new generation of protein database search programs. Nucleic Acids Res. 25: 3389-3402. - PMC - PubMed
1. Attwood, T.K. 2002. The PRINTS database: A resource for identification of protein families. Brief. Bioinform. 3: 252-263. - PubMed
1. Bateman, A., Coin, L., Durbin, R., Finn, R.D., Hollich, V., Griffiths-Jones, S., Khanna, A., Moxon, M.M., Sonnhammer, E.L., Studholme, D.J., et al. 2004. The Pfam protein families database. Nucleic Acids Res. 32: D138-D141. - PMC - PubMed

WEB SITE REFERENCES

1. http://probcons.stanford.edu; ProbCons alignment tool.

Publication types

MeSH terms

LinkOut - more resources

Full Text Sources
Other Literature Sources
- H1 Connect - Access expert opinions and insights on biomedical research.
- The Lens - Patent Citations Database

ProbCons: Probabilistic consistency-based multiple sequence alignment - PubMed (original) (raw)