UniProt

UniProt
Research center	EMBL-EBI, UK; SIB, Switzerland; PIR, US.
Primary citation	UniProt Consortium
Access
Data format	Custom flat file, FASTA, GFF, RDF, XML.
Website	www.uniprot.org; www.uniprot.org/news/
Download URL	www.uniprot.org/downloads & for downloading complete data sets ftp.uniprot.org
Web service URL	Yes – JAVA API see info here & REST see info here
Tools
Web	Advanced search, BLAST, ClustalO, bulk retrieval/download, ID mapping
Miscellaneous
License	Creative Commons Attribution-NoDerivs
Versioning	Yes
Data release; frequency	8 weeks
Curation policy	Yes – manual and automatic. Rules for automatic annotation generated by database curators and computational algorithms.
Bookmarkable; entities	Yes – both individual protein entries and searches

UniProt is a freely accessible database of

Washington, DC

, USA.

The UniProt consortium

The UniProt consortium comprises the

ExPASy (Expert Protein Analysis System) servers that are a central resource for proteomics tools and databases. PIR, hosted by the National Biomedical Research Foundation (NBRF) at the Georgetown University Medical Center in Washington, DC, US, is heir to the oldest protein sequence database, Margaret Dayhoff's Atlas of Protein Sequence and Structure, first published in 1965.^[2] In 2002, EBI, SIB, and PIR joined forces as the UniProt consortium.^[3]

The roots of the UniProt databases

Each consortium member is heavily involved in protein database maintenance and annotation. Until recently, EBI and SIB together produced the Swiss-Prot and TrEMBL databases, while PIR produced the Protein Sequence Database (PIR-PSD).

protein sequence

coverage and annotation priorities.

Swiss-Prot was created in 1986 by

post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recognizing that sequence data were being generated at a pace exceeding Swiss-Prot's ability to keep up, TrEMBL (Translated EMBL Nucleotide Sequence Data Library) was created to provide automated annotations for those proteins not in Swiss-Prot. Meanwhile, PIR maintained the PIR-PSD and related databases, including iProClass

, a database of protein sequences and curated families.

The consortium members pooled their overlapping resources and expertise, and launched UniProt in December 2003.[10]

Organization of the UniProt databases

UniProt provides four core databases: UniProtKB (with sub-parts Swiss-Prot and TrEMBL), UniParc, UniRef and Proteome.

UniProtKB

UniProt Knowledgebase (UniProtKB) is a protein database partially curated by experts, consisting of two sections: UniProtKB/Swiss-Prot (containing reviewed, manually annotated entries) and UniProtKB/TrEMBL (containing unreviewed, automatically annotated entries).^[11] As of 22 February 2023^[update], release "2023_01" of UniProtKB/Swiss-Prot contains 569,213 sequence entries (comprising 205,728,242 amino acids abstracted from 291,046 references) and release "2023_01" of UniProtKB/TrEMBL contains 245,871,724 sequence entries (comprising 85,739,380,194 amino acids).^[12]

UniProtKB/Swiss-Prot

UniProtKB/Swiss-Prot is a manually annotated, non-redundant protein sequence database. It combines information extracted from scientific literature and

biocurator-evaluated computational analysis. The aim of UniProtKB/Swiss-Prot is to provide all known relevant information about a particular protein. Annotation is regularly reviewed to keep up with current scientific findings. The manual annotation of an entry involves detailed analysis of the protein sequence and of the scientific literature.^[13]

Sequences from the same gene and the same species are merged into the same database entry. Differences between sequences are identified, and their cause documented (for example alternative splicing, natural variation, incorrect initiation sites, incorrect exon boundaries, frameshifts, unidentified conflicts). A range of sequence analysis tools is used in the annotation of UniProtKB/Swiss-Prot entries. Computer-predictions are manually evaluated, and relevant results selected for inclusion in the entry. These predictions include post-translational modifications, transmembrane domains and topology, signal peptides, domain identification, and protein family classification.^[13]^[14]

Relevant publications are identified by searching databases such as PubMed. The full text of each paper is read, and information is extracted and added to the entry. Annotation arising from the scientific literature includes, but is not limited to:^[10]^[13]^[14]

Protein and gene names
Function
Enzyme-specific information such as catalytic activity, cofactors and catalytic residues
Subcellular location
Protein-protein interactions
Pattern of expression
Locations and roles of significant domains and sites
substrate
- and cofactor-binding sites
Protein variant forms produced by natural genetic variation,
proteolytic
processing, and post-translational modification

Annotated entries undergo quality assurance before inclusion into UniProtKB/Swiss-Prot. When new data becomes available, entries are updated.

UniProtKB/TrEMBL

UniProtKB/TrEMBL contains high-quality computationally analyzed records, which are enriched with automatic annotation. It was introduced in response to increased dataflow resulting from genome projects, as the time- and labour-consuming manual annotation process of UniProtKB/Swiss-Prot could not be broadened to include all available protein sequences.

EMBL-Bank/GenBank/DDBJ nucleotide sequence database

are automatically processed and entered in UniProtKB/TrEMBL. UniProtKB/TrEMBL also contains sequences from

Ensembl, RefSeq and CCDS.^[15] Since 22 July 2021 it also includes predicted with AlphaFold tertiary and Alphafold-multimer can even do quaternary^[16] structures.^[17]

UniParc

UniProt Archive (UniParc) is a comprehensive and non-redundant database, which contains all the protein sequences from the main, publicly available protein sequence databases.[18] Proteins may exist in several different source databases, and in multiple copies in the same database. In order to avoid redundancy, UniParc stores each unique sequence only once. Identical sequences are merged, regardless of whether they are from the same or different species. Each sequence is given a stable and unique identifier (UPI), making it possible to identify the same protein from different source databases. UniParc contains only protein sequences, with no annotation. Database cross-references in UniParc entries allow further information about the protein to be retrieved from the source databases. When sequences in the source databases change, these changes are tracked by UniParc and history of all changes is archived.

Source databases

Currently UniParc contains protein sequences from the following publicly available databases:

DDBJ/GenBank
nucleotide sequence databases

Ensembl

European Patent Office (EPO)
FlyBase: the primary repository of genetic and molecular data for the insect family Drosophilidae (FlyBase)
H-Invitational Database (H-Inv)
International Protein Index (IPI)
Japan Patent Office (JPO)
Protein Information Resource (PIR-PSD)
Protein Data Bank (PDB)
Protein Research Foundation (PRF)^[19]
RefSeq
Saccharomyces Genome Database (SGD)
The Arabidopsis Information Resource (TAIR)
TROME^[20]
US Patent Office
(USPTO)
UniProtKB/Swiss-Prot, UniProtKB/Swiss-Prot protein isoforms, UniProtKB/TrEMBL
Vertebrate and Genome Annotation Database
(VEGA)
WormBase

UniRef

The UniProt Reference Clusters (UniRef) consist of three databases of clustered sets of protein sequences from UniProtKB and selected UniParc records.^[21] The UniRef100 database combines identical sequences and sequence fragments (from any organism) into a single UniRef entry. The sequence of a representative protein, the accession numbers of all the merged entries and links to the corresponding UniProtKB and UniParc records are displayed. UniRef100 sequences are clustered using the CD-HIT algorithm to build UniRef90 and UniRef50.^[21]^[22] Each cluster is composed of sequences that have at least 90% or 50% sequence identity, respectively, to the longest sequence. Clustering sequences significantly reduces database size, enabling faster sequence searches.

UniRef is available from the UniProt FTP site.

Funding

UniProt is funded by grants from the National Human Genome Research Institute, the National Institutes of Health (NIH), the European Commission, the Swiss Federal Government through the Federal Office of Education and Science, NCI-caBIG, and the US Department of Defense.^[11]

References

PMID 25348405
.

^ Dayhoff, Margaret O. (1965). Atlas of protein sequence and structure. Silver Spring, Md: National Biomedical Research Foundation.

^ "2002 Release: NHGRI Funds Global Protein Database". National Human Genome Research Institute (NHGRI). Archived from the original on 24 September 2015. Retrieved 14 April 2018.

PMID 12230036
.

PMID 12520019
.

PMID 12520024
.

PMID 8594581
.

PMID 10812477
.

ISSN 1660-9824
.

^
PMID 15036160
.

^
PMID 19843607
.

^ "UniProtKB/Swiss-Prot Release 2023_01 statistics". web.expasy.org. Retrieved 31 March 2023.

^ ^a ^b ^c "How do we manually annotate a UniProtKB entry?". UniProt. September 21, 2011. Archived from the original on Dec 13, 2013. Retrieved 14 April 2018.

^
PMID 14681372
.

^ "Where do the UniProtKB protein sequences come from?". UniProt. September 21, 2011. Archived from the original on Dec 15, 2013. Retrieved 14 April 2018.

PMID 34762488. Archived
from the original on 30 Mar 2024 – via PMC.

^ Hassabis, Demis (22 July 2022). "Putting the power of AlphaFold into the world's hands". Deepmind. Archived from the original on 24 July 2021. Retrieved 24 July 2021.

PMID 15044231. Archived
(PDF) from the original on Mar 30, 2024.

^ "Protein Research Foundation".

^ ftp://ftp.isrec.isb-sib.ch/pub/databases/trome^{[permanent dead link]}

^
PMID 17379688
.

PMID 11294794
.

External links

Wikidata has the property:
UniProt protein ID (P352) (see uses)

UniProt

v
t
e
Bioinformatics
Databases

Sequence databases: GenBank, European Nucleotide Archive, DNA Data Bank of Japan and China National GeneBank

Secondary databases: UniProt, database of protein sequences grouping together Swiss-Prot, TrEMBL and Protein Information Resource

Other databases:
Gene Ontology

Specialised genomic databases: BOLD, Saccharomyces Genome Database, FlyBase, VectorBase, WormBase, Rat Genome Database, PHI-base, Arabidopsis Information Resource, GISAID and Zebrafish Information Network

Software

BLAST

Bowtie

Clustal

EMBOSS

HMMER

MUSCLE

PANGOLIN

SAMtools

SOAP suite

TopHat

Other

Server:
ExPASy

Rosalind (education platform)

Institutions

Broad Institute

Computational Biology Department
(CBD)

Microsoft Research - University of Trento Centre for Computational and Systems Biology (COSBI)

Database Center for Life Science (DBCLS)

DNA Data Bank of Japan (DDBJ)

European Bioinformatics Institute (EMBL-EBI)

European Molecular Biology Laboratory (EMBL)

Flatiron Institute

J. Craig Venter Institute (JCVI)

Max Planck Institute of Molecular Cell Biology and Genetics (MPI-CBG)

US National Center for Biotechnology Information (NCBI)

Japanese Institute of Genetics

Netherlands Bioinformatics Centre (NBIC)

Philippine Genome Center (PGC)

Scripps Research

Swiss Institute of Bioinformatics (SIB)

Wellcome Sanger Institute

Whitehead Institute

Organizations

African Society for Bioinformatics and Computational Biology (ASBCB)

Australia Bioinformatics Resource (EMBL-AR)

European Molecular Biology network (EMBnet)

International Nucleotide Sequence Database Collaboration (INSDC)

International Society for Biocuration (ISB)

International Society for Computational Biology (ISCB)
Student Council (ISCB-SC)

Institute of Genomics and Integrative Biology (CSIR-IGIB)

Japanese Society for Bioinformatics (JSBi)

Meetings

Basel Computational Biology Conference‎ ([BC²])

European Conference on Computational Biology (ECCB)

Intelligent Systems for Molecular Biology (ISMB)

International Conference on Bioinformatics (InCoB)

International Conference on Computational Intelligence Methods for Bioinformatics and Biostatistics (CIBB)

ISCB Africa ASBCB Conference on Bioinformatics

Pacific Symposium on Biocomputing (PSB)

Research in Computational Molecular Biology (RECOMB)

File formats

CRAM format

FASTA format

FASTQ format

NeXML format

Nexus format

Pileup format

SAM format

Stockholm format

VCF format

GFF format

Related topics

Computational biology

List of biobanks

List of biological databases

Molecular phylogenetics

Sequencing

Sequence database

Sequence alignment

Category

Commons

Retrieved from "https://en.wikipedia.org/w/index.php?title=UniProt&oldid=1219093979"

[1] PMID 25348405
.

[dayhoff-2] Dayhoff, Margaret O. (1965). Atlas of protein sequence and structure. Silver Spring, Md: National Biomedical Research Foundation.

[3] "2002 Release: NHGRI Funds Global Protein Database". National Human Genome Research Institute (NHGRI). Archived from the original on 24 September 2015. Retrieved 14 April 2018.

[pmid12230036-4] PMID 12230036
.

[pmid12520019-5] PMID 12520019
.

[pmid12520024-6] PMID 12520024
.

[7] PMID 8594581
.

[Bairoch2000-8] PMID 10812477
.

[9] ISSN 1660-9824
.

[pmid15036160-10] 
PMID 15036160
.

[pmid19843607-11] 
PMID 19843607
.

[SPstats-12] "UniProtKB/Swiss-Prot Release 2023_01 statistics". web.expasy.org. Retrieved 31 March 2023.

[faq45-13] "How do we manually annotate a UniProtKB entry?". UniProt. September 21, 2011. Archived from the original on Dec 13, 2013. Retrieved 14 April 2018.

[pmid14681372-14] 
PMID 14681372
.

[faq37-15] "Where do the UniProtKB protein sequences come from?". UniProt. September 21, 2011. Archived from the original on Dec 15, 2013. Retrieved 14 April 2018.

[16] PMID 34762488. Archived
from the original on 30 Mar 2024 – via PMC.

[17] Hassabis, Demis (22 July 2022). "Putting the power of AlphaFold into the world's hands". Deepmind. Archived from the original on 24 July 2021. Retrieved 24 July 2021.

[pmid15044231-18] PMID 15044231. Archived
(PDF) from the original on Mar 30, 2024.

[19] "Protein Research Foundation".

[20] tp://ftp.isrec.isb-sib.ch/pub/databases/trome^{[permanent dead link]}

[pmid17379688-21] 
PMID 17379688
.

[pmid11294794-22] PMID 11294794
.

[1]

[2]

[3]

[11]

[12]

[13]

[14]

[10]

[15]

[16]

[17]

[19]

[20]

[21]

[22]