Uploaded August 2026 | Updated September 2026, 2 weeks ago
Information about protein sequences (& their corresponding proteins) is housed in UniProt.
UniProt consists of...
- UniProtKB: UniProt KnowledgeBase. It includes annotated entries-one per protein-encoding gene.
- Further split into:
- Swiss-Prot, which has reviewed data
- TrEMBL, which as unreviewed data
- Swiss-Prot (sp)- reviewed, manually curated (more reliable, more info, often something people have actually experimented with a lab)
- TrEMBL (Translated EMBL Nucleotide Sequence Data Library)(tr)- unreviewed, autocurated (less known, be more cautious, most stuff from less-studied organisms is here)
- UniRef: UniProt Reference Clusters. It includes data clustered by % sequence identity, then represented by a reference entry.
- Further split into:
- UniRef100: 100% identity
- UniRef 90 90% identity - 58% reduction in database size
- UniRef 50 50% identity - 79% reduction in database size
- Helps you search protein sequence data more efficiently
- Clusters are giving accession codes containing "UniRef100β, "UniRef90_", or "UniRef100_", followed by the code of the reference entry.
- UniParc: UniProt Archives. It includes ALL non-redundant protein sequence data. Data in UniParc is the least reliable, but it ensures everything is stored permanently somewhere.
Much more here: ebi.ac.uk/training/online/courses/uniprot-quick-tour/the-uniprot-databases
If you're looking for information about a specific protein, you probably want to be in UniProtKB, which holds annotated entries. Each entry contains information about all versions (isoforms and variants) of the protein made from a single gene. The amount of annotation varies from entry to entry, and is reflected in the annotation score.
Pro Tip: You can BLAST against these databases to quickly find a wider variety of sequences similar to a query sequence.
More on annotations here: thebumblingbiochemist.com/glossary/annotation & youtu.be/z2vyZ6bfGkg
More on databases and when & how to use them here: bit.ly/databases_guide & youtube.com/playlist?list=PLfHM6aSvn5eI
More on Uniprot: bit.ly/uniprotprotparam
More on other tools for working with proteins: thebumblingbiochemist.com/365-days-of-science/proteintools
more about all sorts of things: #365DaysOfScience All (with topics listed) π bit.ly/2OllAB0 or search blog: thebumblingbiochemist.com ββ
ββ
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
Information about protein sequences (& their corresponding proteins) is housed in UniProt.
UniProt consists of...
- UniProtKB: UniProt KnowledgeBase. It includes annotated entries-one per protein-encoding gene.
- Further split into:
- Swiss-Prot, which has reviewed data
- TrEMBL, which as unreviewed data
- Swiss-Prot (sp)- reviewed, manually curated (more reliable, more info, often something people have actually experimented with a lab)
- TrEMBL (Translated EMBL Nucleotide Sequence Data Library)(tr)- unreviewed, autocurated (less known, be more cautious, most stuff from less-studied organisms is here)
- UniRef: UniProt Reference Clusters. It includes data clustered by % sequence identity, then represented by a reference entry.
- Further split into:
- UniRef100: 100% identity
- UniRef 90 90% identity - 58% reduction in database size
- UniRef 50 50% identity - 79% reduction in database size
- Helps you search protein sequence data more efficiently
- Clusters are giving accession codes containing "UniRef100β, "UniRef90_", or "UniRef100_", followed by the code of the reference entry.
- UniParc: UniProt Archives. It includes ALL non-redundant protein sequence data. Data in UniParc is the least reliable, but it ensures everything is stored permanently somewhere.
Much more here: ebi.ac.uk/training/online/courses/uniprot-quick-tour/the-uniprot-databases
If you're looking for information about a specific protein, you probably want to be in UniProtKB, which holds annotated entries. Each entry contains information about all versions (isoforms and variants) of the protein made from a single gene. The amount of annotation varies from entry to entry, and is reflected in the annotation score.
Pro Tip: You can BLAST against these databases to quickly find a wider variety of sequences similar to a query sequence.
More on annotations here: thebumblingbiochemist.com/glossary/annotation & youtu.be/z2vyZ6bfGkg
More on databases and when & how to use them here: bit.ly/databases_guide & youtube.com/playlist?list=PLfHM6aSvn5eI
More on Uniprot: bit.ly/uniprotprotparam
More on other tools for working with proteins: thebumblingbiochemist.com/365-days-of-science/proteintools
more about all sorts of things: #365DaysOfScience All (with topics listed) π bit.ly/2OllAB0 or search blog: thebumblingbiochemist.com ββ
ββ
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem





