The InterPro protein families database: the classification resource after 15 years - PubMed (original) (raw)

. 2015 Jan;43(Database issue):D213-21.

doi: 10.1093/nar/gku1243. Epub 2014 Nov 26.

Hsin-Yu Chang 1, Louise Daugherty 1, Matthew Fraser 1, Sarah Hunter 1, Rodrigo Lopez 1, Craig McAnulla 1, Conor McMenamin 1, Gift Nuka 1, Sebastien Pesseat 1, Amaia Sangrador-Vegas 1, Maxim Scheremetjew 1, Claudia Rato 1, Siew-Yit Yong 1, Alex Bateman 1, Marco Punta 1, Teresa K Attwood 2, Christian J A Sigrist 3, Nicole Redaschi 3, Catherine Rivoire 3, Ioannis Xenarios 4, Daniel Kahn 5, Dominique Guyot 5, Peer Bork 6, Ivica Letunic 6, Julian Gough 7, Matt Oates 7, Daniel Haft 8, Hongzhan Huang 9, Darren A Natale 9, Cathy H Wu 10, Christine Orengo 11, Ian Sillitoe 11, Huaiyu Mi 12, Paul D Thomas 12, Robert D Finn 13

Affiliations

PMID: 25428371
PMCID: PMC4383996
DOI: 10.1093/nar/gku1243

The InterPro protein families database: the classification resource after 15 years

Alex Mitchell et al. Nucleic Acids Res. 2015 Jan.

Abstract

The InterPro database (http://www.ebi.ac.uk/interpro/) is a freely available resource that can be used to classify sequences into protein families and to predict the presence of important domains and sites. Central to the InterPro database are predictive models, known as signatures, from a range of different protein family databases that have different biological focuses and use different methodological approaches to classify protein families and domains. InterPro integrates these signatures, capitalizing on the respective strengths of the individual databases, to produce a powerful protein classification resource. Here, we report on the status of InterPro as it enters its 15th year of operation, and give an overview of new developments with the database and its associated Web interfaces and software. In particular, the new domain architecture search tool is described and the process of mapping of Gene Ontology terms to InterPro is outlined. We also discuss the challenges faced by the resource given the explosive growth in sequence data in recent years. InterPro (version 48.0) contains 36,766 member database signatures integrated into 26,238 InterPro entries, an increase of over 3993 entries (5081 signatures), since 2012.

PubMed Disclaimer

Figures

Figure 1.

InterPro matches for UniProtKB entry Q3JCG5 showing predicted protein family membership, domains and sites.

Figure 2.

Detailed InterPro member database match data for UniProtKB entry Q3JCG5.

Figure 3.

Number of entries provided by InterPro and its member databases per year.

Figure 4.

The InterPro Domain Architecture tool add/remove domains pop-up window. The list of domains can be refined using either the search box (A) or drop down menu (B). Domains can be added or removed from the query using plus or minus buttons (C). The number of copies of a particular domain to add to the query is indicated (D). Selecting the Apply button (E) performs the query.

Figure 5.

The InterPro Domain Architecture tool showing the results of searching with a VIT and 14-3-3 domain. Checking the ‘Order sensitivity’ option (A) means that domain order is taken into account in the results section (B). The domains can be reordered by dragging and dropping their graphical representations (C), or removed from the query by dragging them to the dustbin (D) or clicking on the [x] icon next to their name and accession (E). The InterPro accession string (F) summarizes the domain architecture composition.

Figure 6.

Growth of the manually-annotated Swiss-Prot and automatically annotated TrEMBL sections of UniProtKB over the last decade.

Cited by

A genome assembly of decaploid Houttuynia cordata provides insights into the evolution of Houttuynia and the biosynthesis of alkaloids.
Huang P, Li Z, Wang H, Huang J, Tan G, Fu Y, Liu X, Zheng S, Xu P, Sun M, Zeng J. Huang P, et al. Hortic Res. 2024 Jul 30;11(9):uhae203. doi: 10.1093/hr/uhae203. eCollection 2024 Sep. Hortic Res. 2024. PMID: 39308792 Free PMC article.
A near-complete chromosome-level genome assembly of looseleaf lettuce (Lactuca sativa var. crispa).
Zhang B, Xue Y, Liu X, Ding H, Yang Y, Wang C, Xu Z, Zhou J, Sun C, Tang J, Li D. Zhang B, et al. Sci Data. 2024 Sep 4;11(1):961. doi: 10.1038/s41597-024-03830-y. Sci Data. 2024. PMID: 39231996 Free PMC article.
Comparative transcriptomics identifies genes underlying growth performance of the Pacific black-lipped pearl oyster Pinctada margaritifera.
Dorant Y, Quillien V, Le Luyer J, Ky CL. Dorant Y, et al. BMC Genomics. 2024 Jul 24;25(1):717. doi: 10.1186/s12864-024-10636-0. BMC Genomics. 2024. PMID: 39049022 Free PMC article.
SNP and Structural Study of the Notch Superfamily Provides Insights and Novel Pharmacological Targets against the CADASIL Syndrome and Neurodegenerative Diseases.
Papageorgiou L, Papa L, Papakonstantinou E, Mataragka A, Dragoumani K, Chaniotis D, Beloukas A, Iliopoulos C, Bongcam-Rudloff E, Chrousos GP, Kossida S, Eliopoulos E, Vlachakis D. Papageorgiou L, et al. Genes (Basel). 2024 Apr 23;15(5):529. doi: 10.3390/genes15050529. Genes (Basel). 2024. PMID: 38790158 Free PMC article.
SAFPred: synteny-aware gene function prediction for bacteria using protein embeddings.
Urhan A, Cosma BM, Earl AM, Manson AL, Abeel T. Urhan A, et al. Bioinformatics. 2024 Jun 3;40(6):btae328. doi: 10.1093/bioinformatics/btae328. Bioinformatics. 2024. PMID: 38775729 Free PMC article.

References

1. Finn R.D., Bateman A., Clements J., Coggill P., Eberhardt R.Y., Eddy S.R., Heger A., Hetherington K., Holm L., Mistry J., et al. Pfam: the protein families database. Nucleic Acids Res. 2014;42:D222–D2230. - PMC - PubMed
1. Attwood T.K., Coletta A., Muirhead G., Pavlopoulou A., Philippou P.B., Popov I., Romá-Mateo C., Theodosiou A., Mitchell A.L. The PRINTS database: a fine-grained protein sequence annotation and analysis resource—its status in 2012. Database. 2012;10:bas019. - PMC - PubMed
1. Sigrist C.J.A., de Castro E., Cerutti L., Cuche B.A., Hulo N., Bridge A., Bougueleret L., Xenarios I. New and continuing developments at PROSITE. Nucleic Acids Res. 2013;41:D344–D347. - PMC - PubMed
1. Bru C., Courcelle E., Carrère S., Beausse Y., Dalmar S., Kahn D. The ProDom database of protein domain families: more emphasis on 3D. Nucleic Acids Res. 2005;33:D212–D215. - PMC - PubMed
1. Lees J.G., Lee D., Studer R.A., Dawson N.L., Sillitoe I., Das S., Yeats C., Dessailly B.H., Rentzsch R., Orengo C.A. Gene3D: multi-domain annotations for protein sequence and comparative genome analysis. Nucleic Acids Res. 2014;42:D240–D245. - PMC - PubMed

The InterPro protein families database: the classification resource after 15 years - PubMed (original) (raw)

The InterPro protein families database: the classification resource after 15 years

Abstract

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Substances

Grants and funding

LinkOut - more resources

Full Text Sources

Other Literature Sources