the bumbling biochemist
The Protein Data Bank (PDB) - a users guide
updated
more on size exclusion chromatography: http://bit.ly/sizeexclusionchromatography
more on protein chromatography: http://bit.ly/proteincleaning
More about the Bibel lab at LMU and what we do: https://bibellab.lmu.build/
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
Part 1: youtu.be/5w8_zW396MA
More on enzyme catalysis: http://bit.ly/enzymecatalysis
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
More on: AKTA flow paths and injection valve modes – controlling the plumbing on an FPLC! youtu.be/Wn0dRwwocOc
More on running size exclusion chromatography (practical look at running gel filtration on an AKTA PURE)
youtu.be/KJ4v1FwfyP4
More on connecting a SEC column: youtu.be/o0ya3EpM86o
more on size exclusion chromatography: http://bit.ly/sizeexclusionchromatography
more on protein chromatography: http://bit.ly/proteincleaning
More about the Bibel lab at LMU and what we do: https://bibellab.lmu.build/
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
Enzyme fundamentals
* Key properties of enzymes:
* Catalyze (speed up) a reaction in both directions
* By reducing the activation energy required for the reaction to occur
* DO NOT change overall thermodynamic favorability - cannot change equilibrium constants, can only increase the speed at which you can reach that equilibrium
* Don’t get used up in the process
* Lower the activation energy required for the reaction to occur (ΔG‡) by lowering the energy of the Transition State (TS) (more below)
* Enzymes often bind to helper molecules called cofactors that provide additional functional groups to work with
* Can be metals or small organic molecules
* Transition states vs reaction intermediates:
* You will not have to draw any transition states but should know what they are & why they’re important
* You also will not have to draw any energy coordinate diagrams but should be comfortable with what they represent
* Transition states (TS) are fleeting moments (shorter than length of bond vibration) in which bonds are being broken & formed
* Involve partial covalent bonds & partial charges
* Can’t be isolated but we know they have to exist
* High energy points
* Peaks on reaction coordinate diagrams
* Often represented with brackets & the symbol ‡
* The TS with the highest activation energy represents the slowest step in a reaction → will be the rate-determining step (RDS)
* Reaction intermediates are (at least theoretically) isolatable molecules with full covalent bonds
* Lower energy points
* Valleys (troughs) on reaction coordinate diagrams
* Enzymes bind (and thus stabilize) the transition state (TS) the most tightly
* Interactions (noncovalent IMFs & sometimes covalent bonds) with the TS provide “binding energy” that lowers the free energy of the TS → decreases the activation barrier and thus allows the reaction to happen faster
Part 2: youtu.be/Q2Q48UO00W4
More on enzyme catalysis: http://bit.ly/enzymecatalysis
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
youtu.be/KJ4v1FwfyP4
More on connecting a SEC column: youtu.be/o0ya3EpM86o
more on size exclusion chromatography: http://bit.ly/sizeexclusionchromatography
more on protein chromatography: http://bit.ly/proteincleaning
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
This is the set we used, but this isn’t a paid add or anything!: Bio-Rad Gel Filtration Standard #1511901: Pkg of 6 vials, lyophilized mix of thyroglobulin, bovine γ-globulin, chicken ovalbumin, equine myoglobin, and vit B12, MW 1,350–670,000, pI 4.5–6.9
We used it to compare our class mix to and next I plan to run my real proteins of interest (MDH) to compare to!
More on running size exclusion chromatography (practical look at running gel filtration on an AKTA PURE)
youtu.be/KJ4v1FwfyP4
More on connecting a SEC column: youtu.be/o0ya3EpM86o
more on size exclusion chromatography: http://bit.ly/sizeexclusionchromatography
more on protein chromatography: http://bit.ly/proteincleaning
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
More on connecting a SEC column: youtu.be/o0ya3EpM86o
more on protein chromatography: http://bit.ly/proteincleaning
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
The key steps:
1) prime and wash the pump
2) go to column handling to find the pre-column pressure limit for the column you're using
3) go to manual → execute manual instructions and set that pre-column pressure limit (5.0 for a Superdex 200 10/300 GL)
4) set the column position to the column you want, "down flow"
5) set a low flow rate (e.g. 0.5 ml/min for a Superdex 200 10/300 GL)
6) wait until drops come out of the in line
7) when a drop is on the tip, screw in the in line (be sure the tip of the tubing is slightly out when you go in, keep pressure on it when tightening, then tug gently to ensure it's snug)
8) wait for drops to come out the bottom of the column
9) attach the out line (be sure the tip of the tubing is slightly out when you go in, keep pressure on it when tightening, then tug gently to ensure it's snug)
10) check the monitor
Running size exclusion chromatography (practical look at running gel filtration on an AKTA PURE)
youtu.be/KJ4v1FwfyP4
more on size exclusion chromatography resins: http://bit.ly/sizeexclusionchromatography
more on protein chromatography: http://bit.ly/proteincleaning
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
Unmodeled regions were physically there when the data were collected (i.e. that region of the protein or other macromolecule was not missing), but the atoms in the region haven’t left sufficient evidence for scientists to know exactly where they are.
If a letter is grayed out in the sequence and you can’t select it (or see it!), this is because it is in an unmodeled region. You can also see this information in the PDB entry’s “Structure Summary” & “Sequence” Tabs. Look for an “Unmodeled” track. Gray regions indicate that residues there are not modeled.
Sometimes whole amino acids in proteins are left unmodeled. And sometimes you can just make out the backbone (less flexible) but can’t make out the side chain so it’s only partially modeled or an alanine is modeled on instead of the amino acid that’s actually there since alanine is pretty “generic” structure-wise. More on that in my alanine post. http://bit.ly/alaninechirality
Bottom line thing to keep in mind is that scientists can make out details better in some areas of the map than in others, and thus can model those regions more confidently. These regions have a low B-factor (aka temperature or displacement factor). On the other hand, regions with a high B-factor have more "iffy" data - it's less clear how to model in the underlying atoms. If you view a structure in something like PyMol, Chimera, or even the PDB web browser now, you can color the model by B-factor. Typically it’s displayed as a heat map with red corresponding to high B factor. You tend to see a lot of red on the periphery of the protein a blue inside.
No matter where in a molecule some feature is located, even if it has a low B factor, you want to make sure you actually view the maps yourself before relying on the model if you are trying to interpret some claim about it, or make some hypothesis or something.
Before you judge a structure based on its resolution, know that resolution isn’t the only thing that matters and even a low resolution structure can provide valuable information, especially if used in concert with high res info. Such as, for example, placing models from high res crystal structures of individual complex components into a low res cryo map of the whole complex.
This video covers basic navigation and features, and using Uniprot and PDB in combination with PyMOL to find and show a protein with key residues in stick. youtu.be/CeTTmm5DI_k
More on PyMOL: bit.ly/pymolintro
more on the PDB: bit.ly/pdbstructures & youtu.be/5wMwlsChz98
more on UniProt: bit.ly/uniprotprotparam & youtu.be/f75f6QCe1gA
Structural biology resources - guide to my favorite software, articles, books, lectures, websites
downloadable version: bit.ly/structure_guide ; YouTube: youtu.be/OoWpQ_2pLac
* If you need a refresher on X-ray crystallography, start here: http://bit.ly/xraycrystallography2
* and I made a page on my blog that has links to all my structural biology posts: bit.ly/structural_biology
* and here's a link to my YouTube structural biology playlist: youtube.com/playlist?list=PLUWsCDtjESrGhwVxsRbTJdL-BEsN60RCs
* here are some key posts:
* resolution: blog: bit.ly/structure_resolution ; YouTube: youtu.be/Ijm-nDzLWhA youtu.be/1DZClUKowsY and this intro video: YouTube: youtu.be/t5WeMkY0skU
* understanding crystal structures: bit.ly/crystalstructuremodels & youtu.be/YK3VkqD2o2s
* intro to PDB, crystal structure entries, crystal contents, resolution: bit.ly/pdbstructures & youtu.be/Re2gwi-_OEw & youtu.be/IZtHsUFbye
* maps: blog: bit.ly/xraymaps ; YouTube: youtu.be/Smi9yXbsesg
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
blog: bit.ly/desalt_buffer_exchange
Imagination you have a turtle race with a crowd of turtles en route and then partway through the race a cheetah and a crowd of tortoises come in. The cheetah’s gonna beat all those tortoises, but since the turtles had such a head start, it’ll finish in a crowd of them. And if you stop the race before the tortoises come in you get that cheetah in a crowd of turtles (the starting racers) instead of in a crowd of tortoises (the partners it came in with).
Now imagine that, instead of a cheetah, you have one of those giant tortoises which are so awesome. They don’t move faster than the turtles, but they get to take a shortcut. They’ll still finish ahead of the turtles they come in with that don’t get to take the shortcut, but they’ll finish with the starting crowd. This is the concept behind buffer exchange
You have a column filled with resin (little beads) and those bead have “secret tunnels” - winding pores that only molecules small enough to enter them can go through. Molecules (be they proteins, DNA, RNA, etc.) that can’t fit get to go around (shortcut!) but molecules that can fit (like salts, free ATP, etc.) have to go the long way. And because they have to travel further, they take longer to go through.
Before I talked about how I use size exclusion chromatography (SEC) during protein purification as a “polishing step” to separate the protein I want from proteins that are different sizes. In that case, I want good resolution (ability to separate similar-sized things) - so I want a “long racecourse” so that the molecules encounter more beads, with smaller things getting slowed down more and more and you can catch the outgoing molecules on their way out by taking fractions. If you want good resolution, you want a long, skinny column with little beads
But if you want more of a yes/no resolution (like salt vs protein) you want a short, fat column and big beads.
G-25 is the name of the resin, which is a form of “Superdex” - the dex is for dextran - a sugar that, along with agarose, forms the beads’ gel mesh. These versions are small, perfect for the small volumes I’m using, but they also have bigger ones, including ones you can use for proteins. In undergrad, I used PD-10 columns for buffer exchange of proteins - those are smaller volumes of the HiPrep desalting column I’m currently using.
The DESALTING part comes in handy for things like ion exchange chromatography (IEX), where you get a protein to bind to a column based opposite-charge attractions, then you add increasing levels of salt to outcompete it (when a salt dissolves it breaks into its component ions (charged particles) - e.g. table salt (sodium chloride, NaCl) becomes Na⁺ + Cl⁻. This can leave you with a very salty (though hopefully also very pure) protein.
To remove that excess salt you could use dialysis, where you put the protein in a semipermeable-membraned pouch and put it in a bunch of low salt buffer - protein stays in the pouch & salt flows out until concentrations of salt equalize inside & out. Cheap & does the trick, but takes a while… you can instead use a desalting column which you “pre-fill” with the liquid that you want your protein to come out with. Also good for removing small competitors like imidazole that you might use for affinity chromatography.
more on other size exclusion chromatography resins: http://bit.ly/sizeexclusionchromatography
more on dialysis: bit.ly/proteindialysis
more on protein chromatography: http://bit.ly/proteincleaning
more on radiolabeling: http://bit.ly/radiolabelings
more on oligo synthesis: http://bit.ly/2We8e8W
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
#scicomm #biochemistry #molecularbiology #biology #sciencelife #science #realtimechem
PS – sorry about the rocking noise. It seems to be coming through the ceiling in our lab so I can't do anything about it.
blog: http://bit.ly/minipreps
So we want to purify our plasmid DNA (pDNA), but that plasmid is in bacterial cells which have their own genomic DNA (gDNA) as well as a bunch of proteins and stuff. So how do we isolate the pDNA?
Clearly, we’ll need to get it out of the cells, which we can do by breaking the cells open (LYSIS). But before you break the cells open, you want to make sure they have an ideal environment to come out into. The cell membrane provides a nice natural barrier to allow you to swap out the external environment - so take advantage of it!
Right now, the cells are in the liquid media you grew them in (often Lysogeny Broth, LB) – in addition to the antibiotic for selection, the media contains salts, peptides (from the tryptone - trypsin-digested casein protein), vitamins etc. (from the yeast extract), all sorts of stuff the bacteria want/need to survive and thrive (and make lots of copies of your plasmid). http://bit.ly/bacterialmedia
If our goal is to purify the pDNA, we might as well get rid of as much of that stuff we don’t want now, while the cell membrane provides a convenient natural barrier (especially because some of that stuff could interfere with future steps). We also want to minimize the volume we’re working with. Right now, the cells are “spread out” in excess media but if we “spin them down” by centrifugation, we can separate the cells (heavy) from the media (light).
So we stick the tubes in a centrifuge, which spins really fast, pelleting the cells. We then pour off the liquid (SUPERNATANT) but keep the pellet which has our cells.
Now we resuspend them in a smaller volume of a more ideal solution. We’re going to be using a series of different solutions we refer to as “buffers.” “Buffer” is in reference to their pH-stabilizing ability which comes from component(s) that can both give protons (act as an acid) and take protons (act as a base). But we often use the term “buffer” pretty broadly to refer to different pH-stabilized “salt waters” http://bit.ly/phbuffers
The first buffer we’ll use in the miniprep is RESUSPENSION BUFFER (“P1” in this QIAprep kit (this isn’t a paid endorsement or anything, just what our lab uses) has:
∙ Tris – this acts as the actual pH buffer (pH stabilizer) – it’s here because we’re not ready to change the pH yet
∙ glycerol – this is an osmotic balancer – osmosis refers to the movement of water from where water is more concentrated (there’s less other stuff around) to where water is less concentrated (there’s more other stuff in between the water molecules). Since there’s a lot of stuff in the cells, if there isn’t a lot of stuff outside the cells, water will flood in. So we add glycerol to increase the amount of stuff outside the cells and thus keep the cells from bursting before we’re ready. Later we’ll burst them with detergent (synthetic soap), not pressure differences
∙ RNaseA – this degrades RNA (and is why P1 has to be stored in the fridge) – we don’t want the RNA, so might as well break it down so it doesn’t clog things up or try to tag along (a lot of times in the lab we’re trying to protect from RNaseA, because this RNA chewer is all over the place (for example, bacteria often secrete it to destroy foreign RNA (like from viruses) before they can invade), but here our usual enemy is our friend because we’re both on team DNA! (RNaseA won’t chew DNA because it works in a sneaky way that gets the RNA to attack itself using its 2’ oxygen (which DNA doesn’t have - hence the d in DeoxyriboNucleic Acid) bit.ly/rnaseadepc
note: with this kit you have to add the RNaseA to the buffer (and then check the check box on top of the bottle to signal you did)
∙ EDTA – this is a chelating agent (metal-biter) - it binds divalent cations (molecules with a +2 charge) like Mg²⁺ or Ca²⁺ which DNases (DNA chewers) need – prevents DNases from degrading the plasmid (RNaseA doesn’t need cation cofactors so it can still work and chew away the RNA).
Finished in comments
Choose a pore size* depending on how gunky your sample is & what you’re trying to filter out. The smaller the pore size, the more you’ll filter out, but the harder it will be to push things through.
- for “sterilization” use 0.2 or 0.22 um - this won’t filter out all viruses, but it will filter out bacteria, fungi, etc.
- for general filtration purposes where you’re just trying to filter out particulates, 0.45 um is a common go-to
- For gunkier solutions, you might want to go higher, such as 1 μm or so, which is good for insect cell lysates (0.45 usually works fine for bacteria lysates which are typically less gunky)
If you need to use a lower pore size but have something gunkier you will likely have to use multiple filters and a lot of muscle. You can also use a caulk dispenser to help push if you have one. But be sure not to push so hard you damage the filter.
You can also do sequential filtrations - start with a bigger pore size to get out big stuff then take the pre-filtered solution and put it through a finer filter.
Whatever you do, make sure to get out extra air that can make your pushing life more difficult. Turn the syringe over, remove the filter, push the air out (careful not to push the sample out), reattach, and now try pushing.
Choose a membrane type based on your solvent type and what you don’t (or do) want to get stuck to it).
We often use PES (polyethersulfone) or PVDF (polyvinylidene difluoride), also sometimes CA (cellulose acetate), or MCE (mixed cellulose esters) for aqueous (water-based) solutions.
These work well and have very low protein binding (especially PES and PVDF) which is good when your filtering lysates and protein solutions where you’re trying to save the proteins. If you want to remove proteins (such as trying to minimize RNase contamination) or don’t care about them, you can use nitrocellulose (cellulose nitrate) (which has high protein-binding and is often what we use in western blots and slot blots to capture proteins).
Honestly, for most intents and purposes we just use these ones interchangeably and I’m not sure when one type might be better than another but if you have specific concerns about pH or stuff like that, check out compatibility charts like these:
“Chemical Compatibility of Filter Components”, Millipore, 2020 emdmillipore.com/Web-CA-Site/en_CA/-/CAD/ShowDocument-Pronet?id=201510.399&usg=AOvVaw3h0KMcgRcLW-ZMsoV9AlbV
“Membrane Filter Chemical Compatibility Chart”, Tisch Scientific scientificfilters.com/membrane-filter-chemical-compatibility-chart
If you are working with an organic solvent-based solution, however, you’ll need to use a different type of membrane. Check compatibility charts to see what’s appropriate for your purposes, but if you’re using DMSO, nylon might be the way to go. It’s what I use after learning the hard way that DMSO will dissolve some of those other membranes… Nylon is another one of those that binds proteins, though, so don’t use it for protein solutions.
Choose the filter diameter (13mm, 25mm, 33mm, etc.) based on the volume of your sample, how much you care if you lose some, and how impatient you are. The larger the diameter (and therefore the surface area, aka the Effective Filtration Area (EFA)) the faster it will go, but you will lose more because some of the solution always gets stuck on the surface and in the casing (if you’re wondering why you have less coming out than you started with, that’s why - the hold-up volume!) If you’re really worried about losing stuff they make tiny, super cute, 4mm filters that go on the syringes.
And you can also start by pulling up a little air before you pull up your sample - this will make it so that you have a little pocket of air above your sample that you can use to push all the liquid through the syringe without some getting stuck in the stopcock.
You may also need to worry about “extractables” - stuff like trace metals that can dissolve from the membrane and come out with your sample, contaminating it. If you’re really worried about this you can pre-wash the filter but that’s not common. andyjconnelly.wordpress.com/2016/09/28/syringe-filters/
Finished in comments
- add 1μL per nmol (look on tube) for 1mM or 10μL per nmol for 100μM
- make dilutions of the stock to your working concentration & aliquot if you’re going to be using lots of times
- avoids contaminating stock, minimizes freeze-thaws, lets you keep at higher concentration which is more stable
- don’t do huge dilutions - instead dilute step-wise
blog: bit.ly/oligo_working
more in video & figures and these great web pages I found
IDT Tips for resuspending and diluting your oligonucleotides idtdna.com/pages/education/decoded/article/tips-for-resuspending-and-diluting-your-oligonucleotides
Nolan Speicher, former IDT associate.
Published Mar 31, 2017
Revised/updated Sep 22, 2017
My oligos have arrived: Now what? Resuspension, dilution, storage, and other tips
idtdna.com/pages/education/decoded/article/my-oligos-have-arrived-now-what-
Elisabeth Wagner, PhD, Manager of Scientific Applications Support, IDT.
Published Jan 14, 2014Revised/updated Sep 10, 2016
Storing oligos: 7 things you should know idtdna.com/pages/education/decoded/article/storing-oligos-7-things-you-should-know
Nan Pazdernik, PhD, Science Writer, IDT
Nolan Speicher, former IDT associate.
Published Jun 20, 2017Revised/updated Dec 3, 2020
Millipore Sigma: Oligonucleotide Handling & Stability: sigmaaldrich.com/US/en/technical-documents/protocol/genomics/dna-and-rna-purification/oligonucleotide-handling-and-stability
more on concentrations and C1V1=C2V2: blog form: http://bit.ly/c1v1equalsc2v2 ; YouTube: youtu.be/JbtTwDOVyOo
more on dimensional analysis: http://bit.ly/dimensionalanalysising & youtu.be/KQMA0aAGfP4
more on how these things are synthesized: http://bit.ly/solidstateoligo
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
More on PyMOL: bit.ly/pymolintro
more on the PDB: bit.ly/pdbstructures & youtu.be/5wMwlsChz98
more on UniProt: bit.ly/uniprotprotparam & youtu.be/f75f6QCe1gA
more on PKU: http://bit.ly/pkustoryandscience
And more on the expansion to extended, 12-character, PDB IDs here: blog: thebumblingbiochemist.com/365-days-of-science/pdbids ; YouTube: youtu.be/U-MEtWaFNdI
Acevedo, R., & Procko, K. (Eds.). (2025). Seeing the Invisible: Learning to Teach with Biomolecular Visualization. The University of Texas at Austin. https://doi.org/10.15781/m4zd-7691
(1)
My chapter: https://utexas.pressbooks.pub/molviz-education/chapter/a-guide-to-pdb-structures-and-the-rcsb-pdb/
Structural biology resources - guide to my favorite software, articles, books, lectures, websites
downloadable version: bit.ly/structure_guide ; YouTube: youtu.be/OoWpQ_2pLac
* If you need a refresher on X-ray crystallography, start here: http://bit.ly/xraycrystallography2
* and I made a page on my blog that has links to all my structural biology posts: bit.ly/structural_biology
* and here's a link to my YouTube structural biology playlist: youtube.com/playlist?list=PLUWsCDtjESrGhwVxsRbTJdL-BEsN60RCs
* here are some key posts:
* resolution: blog: bit.ly/structure_resolution ; YouTube: youtu.be/Ijm-nDzLWhA youtu.be/1DZClUKowsY and this intro video: YouTube: youtu.be/t5WeMkY0skU
* understanding crystal structures: bit.ly/crystalstructuremodels & youtu.be/YK3VkqD2o2s
* intro to PDB, crystal structure entries, crystal contents, resolution: bit.ly/pdbstructures & youtu.be/Re2gwi-_OEw & youtu.be/IZtHsUFbye
* maps: blog: bit.ly/xraymaps ; YouTube: youtu.be/Smi9yXbsesg
PyMOLWiki: pymolwiki.org/index.php/Main_Page
Tutorial videos:
Chris Berndsen – SeamDock: loom.com/share/490bd9ffda1a40fe99984e78a6f8d647?sid=2fc27dde-a35d-4a13-86c3-67cc465a2c49
Swanson Does Science, PyMOL for Beginners - video 1: orientation – Using the mouse: youtu.be/wiKyOF-pGw4?si=iaBNERt-h0tP5hzC&t=597 min 10-11:20
The bumbling biochemist: PyMol tips to not lose orientations & styling: save views, scenes, & selections youtu.be/-Ct9OTMjG7Q
The bumbling biochemist: PyMol tips: hide → unselected; copy to object; extract to object; export molecule, etc. youtu.be/AUfZNrsXza4 Shows how to show monomer with NAD & malate
Molecular memory PyMOL 101
* PyMOL 101 Lesson 2: Basic Selection, Show, Hide, and Actions Menus (Carbonic Anhydrase Active Site)
youtu.be/UN8cj7omiCM?si=YE75E_Vm2js0WrIz
* PyMOL 101, Lesson 4: Color, Transparency, and Alternate Renderings
youtu.be/v4QYuYGOBD8?si=4Yw7rKvYK27wunnd
* PyMOL 101, Lesson 5: Select by Distance, the Mouse Control Menu, and Labeling
youtu.be/1MLvFLWbj1A?si=V8KRaYQtaWsGmzCI
* PyMOL 101, Lesson 6: Introduction to the measurement wizard & polar contacts, the N- & C-terminus
youtu.be/YrWOohtQTuk?si=4FvsErG8U064RG8q
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
blog form: http://bit.ly/bohreffect
More on hemoglobin cooperativity: youtu.be/Y4gaUbIcsvA
The textbook-y explanation, and then a deeper, bumblier look. The Bohr effect describes how low pH (acidity) lowers the affinity of hemoglobin for oxygen, making hemoglobin more likely to offload oxygen in areas of low pH, which for reasons I’ll get into, tissues in need of oxygen tend to have. So how does it work? Well first, what is hemoglobin?
Hemoglobin is this protein in your blood that picks up oxygen from your lungs and takes it all the way to your toes. We’ve talked about it a lot in terms of when it doesn’t work - like in sickle cell disease where mutations cause it to clump up and block blood vessels, cutting off circulation to tissues and organs, and causing pain and organ damage. http://bit.ly/sicklecelldiseases
But we haven’t talked much about hemoglobin when it *does* work. So that’s what I want to tell you about today - how hemoglobin “knows” when to take or give up oxygen - and how we know how it “knows.” It’s a really cool protein - it has 4 subunits which can each take an oxygen and they work as a team so that one binding makes it easier for the others to bind and one letting go makes it easier for the others to let go. This is called cooperativity - and we’ll talk more about what causes it in a bit. But this kind of take-all or dump-all approach with respect to oxygen means that it has to be tightly regulated so that it doesn’t all get dumped too soon. And the hemoglobin doesn’t just retake what it dumps. And, as we’ll see, the Bohr effect helps with both of these.
Often “respiration” is used to describe breathing, but biochemists often talk about respiration in terms of the processes that take place in your cells to use the oxygen (O₂) you get when you breath to make energy. The basic idea of cellular respiration is that you breathe in oxygen, combine it with breakdown products of things like glucose (blood sugar) to make the energy storage molecule ATP, and generate carbon dioxide (CO₂) that you breathe out as a waste product in the process. That has to occur in cells all throughout your body. So oxygen has to be able to get to all of those cells. And you have CO₂ being produced in all of those cells. But you only have oxygen entering your body in one place - your lungs.
Your lungs might not seem that big, but they use their allocated space wisely. They have a branching structure with the branches ending in tiny grape-like parts called alveoli which are covered with tiny little blood vessels called capillaries. They’re tiny in terms of diameter, but huge in terms of surface area, so they can let lots of oxygen in. It’s an easy route for free oxygen to get in - it just has to diffuse through the cell membrane, which is easy for it because it’s so small. Diffusion’s basically just random molecule moving - it leads to a net movement of molecules from where they’re more cramped (areas of high concentration) to places where they’re less cramped (areas of low concentration) until the concentration is equal everywhere. The molecules still move around randomly but because their movement is random, for each molecule “moving left” there’s another one “moving right” so there’s no net movement and no net change in concentration anywhere once a system reaches equilibrium.
When it comes to reporting concentrations of gases, people frequently talk in terms of partial pressures. The “partial” comes from the fact that when you have a gas or a mixture of gases, the total pressure is proportional to the number of gas molecules, not their identity (e.g. either 1000 CO₂ gas molecules OR 1000 O₂ gas molecules would produce the same pressure). So, instead of counting individual gas molecules, you can get information about how many gas molecules there are by measuring the pressure. And since the identity of the gas doesn’t matter, you can add the pressure contribution of different gases (their partial pressures) together to get the total pressure (e.g. a mixture of 1000 CO₂ gas molecules AND 1000 O₂ gas molecules TOGETHER would have a total pressure proportional to 2000 gas molecules).
Finished in comments
* 2,3-BisPhosphoGlycerate (BPG) is produced in erythrocytes (red blood cells) in an offshoot of glycolysis (Rapoport Leubering Cycle or Shunt)
* 2,3-BPG binds to the central channel that is open in the T state (high oxygen affinity) (but not the R state (low oxygen affinity))
* Binding of 2,3-BPG stabilizes the T state, decreasing affinity for oxygen (shifting the curve to the right)
* Fetal hemoglobin has a lower affinity for 2,3-BPG, so it has its affinity decreased less (gets shifted right LESS)
- so it has a higher affinity for oxygen than adult hemoglobin in the presence of 2,3-BPG (left-shifted)
- this lets it efficiently bind and release oxygen in the low-oxygen conditions it’s exposed to
* Production of 2,3-BPG is upregulated (more is made) at high altitudes (where there’s less oxygen available in the lungs) to allow for roughly the same amount of oxygen to be delivered to tissues as if you were at sea level
* Small decrease in oxygen saturation in lungs - pick up a little less
* But big decrease in oxygen saturation in tissues – drop off a lot
* 2,3-BPG is classified as an allosteric, heterotropic, negative modulator
external figures: http://biocheminfo.com/2020/04/07/rapoport-leubering-cycle-or-shunt-synthesis-of-23-bisphosphoglycerate ; themedicalbiochemistrypage.org/hemoglobin-and-myoglobin/#google_vignette ; LHcheM, CC BY-SA 3.0 creativecommons.org/licenses/by-sa/3.0, via Wikimedia Commons
Much more here: http://bit.ly/sicklecelldiseases & youtu.be/q5CzzD6OpLQ
More on the Bohr effect: blog form: http://bit.ly/bohreffect ; YouTube: youtu.be/kiHDr7toR4U
hemoglobin animation tutorial: Interactive tutorial from by Eric Martz and Frieda Reichsman: bioinformatics.org/jmol-tutorials/jtat/hemoglobin/6reg/chapter.htm
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
We ran out of time - of course, so I added an (optional) PART 2 afterwards to cover a few more of the techniques students expressed interest in and/or I knew were in their papers before uploading it for them. Figured I’d upload it on YouTube too in case anyone else might find it useful. Apologies for the poor audio. There was a problem with the in-class room recording, so I had to just use the Zoom audio.
If you want to know about any of these techniques or others: thebumblingbiochemist.com/lets-talk-science/techniques & youtube.com/playlist?list=PLUWsCDtjESrGlFUCaizromCCk1-z2-GMQ
solution)
* The more protons there are, the more acidic, and the lower the pH
* The fewer protons there are, the more basic/alkaline, and the higher the pH
The pKa is the pH at which half the copies of an acidic site is deprotonated (and thus in conjugate base state) and half is protonated (and this in conjugate acid form)
* At a pH below the pI, there’s more protons available to take, so more conjugate acid
* At a pH above the pI, there are fewer protons available, so more conjugate base
The pKR is just the pKa of an amino acid R-group (side chain)
The pI (isoelectric point) is the pH at which a molecule is net neutral (has no net charge)
* At a pH below the pI, a molecule is positively-charged
* At a pH above the pI, a molecule is negatively-charged
pI tells you about charge. pKa does NOT. Instead, it tells you about protonation. And at a single site. The molecule may be negative, neutral, or positive when that site is protonated.
The pI comes from the combination of acid/base groups in a molecule
The pH comes from the contributions of all acids/bases in a solution
Something can only act as an acid in its protonated state.
But the stronger the acid is, the less likely it will be to be in that state!
So the more likely it is to be in the deprotonated state, where it can only act as a base!
If the molecule is neutral in its protonated state (conjugate acid), it will be negatively-charged in its deprotonated state (conjugate base).
But if the molecule is positively-charged in its protonated state (conjugate acid) it will be neutral in its deprotonated state (conjugate base).
The happier a molecule is to be in a state, the more likely it will be in that state. Resonance and inductive effects can make molecules happy, so resonance etc. that stabilizes one state will make it favorable, even if it comes with charge.
We call amino acids basic or acidic based on what their neutral form acts as - even if their predominant form acts the opposite! So, for example, you pretty much always find arginine in its protonated, acid state because its neutral form is a stronger base, so we call it basic (and show it as blue). On the other hand, you pretty much always find aspartate in its deprotonated, base state because its neutral form is a stronger acid, so we call it acidic (and show it as red)
The protonated form will always be the conjugate acid and the deprotonated form will always be the conjugate base. The stronger the acid, the more likely you are to find it in the conjugate base form (the result of “acting” as an acid meaning giving up a proton). The weaker the acid, the more likely you are to find it in the conjugate acid form (because it’s bad at acting as an acid meaning giving up a proton).
We classify amino acids as “acidic” or “basic” based on their neutral forms.
“Acidic” amino acids have neutral conjugate acid forms and negative conjugate base forms.
“Basic” amino acids have neutral conjugate base forms and positive conjugate acid forms.
more on pH, pKa, and the Henderson-Hasselbalch equation: http://bit.ly/phbuffers & youtu.be/HnmuGTHW3dg
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
Entropy is freedom to move or "disorder." If you imagine a bunch of water molecules moving around in "bulk water," where there is high entropy, each time you take a picture of them, the different molecules will be in different places (microstates) so it would seem disordered. But, if water is stuck in place such as in a solvation shell around a hydrophobic molecule (kinda like a skin), the molecules would still be in the same place each time you took a picture. Entropy is favorable. Molecules like to move around. The hydrophobic effect lets more water molecules move around by reducing the surface area of the hydrophobic thing that needs to be covered by that skin, so fewer water molecules need to be in the solvation shell, and more water molecules can be in the "bulk water." The way it does that is by pushing the hydrophobic/nonpolar things together, like sweeping up a mess into a compact pile.
More on the hydrophobic effect: youtu.be/VoIxAmcMsb8
blog form: http://bit.ly/hydrophobiceffectPSA & bit.ly/dontignorewater
Quick overview: youtu.be/g1_i_qwfXxM
blog form (also has static graphics): http://bit.ly/bindingaffinityavidity
The whole premise of biochemistry is that molecules interact to do things. For example(s), the protein enzyme (reaction mediator) DNA Polymerase links together nucleotides (DNA letters) to copy DNA; a protein called tubulin assembles itself into structural supports and molecular conveyor belts in your cells; and proteins called antibodies bind to foreign molecules (like viral proteins) and call for help, etc. Pretty awesome, right? But in order for any of this to happen, the molecules have to first bind one another. Which means they have to
1) come into contact with one another and
2) like each other (first enough to bind and then enough to stay bound)
We sometimes call binding partners “ligands” (and sometimes we call one partner a “receptor” and another a “ligand”) - they can be anything from proteins to nucleic acids (DNA or RNA) to “small molecules” (things like pharmaceutical drugs, etc.). The higher the concentration of the partner (the more copies of it there are in some space), the more likely they are to come into contact with one another. And the more they like each other, the more likely they are to stick (and stay stuck) if they do contact one another. Therefore, the amount of sticking (and how much stuck you’ll find if you look) depends on concentration and binding strength.
Another way to think of binding partners is as really really tiny people on dates. The concentration is like how likely they are to run into potential partners (are you in Antarctica? or at a speed dating session?). And affinity is like how likely they are to get married and not get divorced.
Even if the concentration changes, that doesn’t change how much the partners “like each other” (Prince charming is just as charming if you meet him at the bar or on an ice floe). In biochemical terms, the binding strength is constant (at least for a given set of conditions (same temperature, salt concentration, etc.) because it’s a property of the binding partners themselves. And we call this “binding strength” AFFINITY.
We can measure binding affinity by altering the concentrations, measuring the binding, and fitting it to an equation that takes into account the contribution of concentration and “hides it” so you can see the constant part - the affinity! If that didn’t make sense, bear with me and I’ll get into more detail, but the end result is we get a value called the dissociation constant, abbreviated Kd. This value tells us what concentration of one binding partner (we can call it B if you want) would lead to half of its binding partner (we’ll call A) being bound at equilibrium (i.e. once the rates of binding and unbinding have stabilized and the mixture has found its happy ratio of bound & unbound).
The higher the affinity (the more sticky they are for one another) the less ligand is required to reach that value- higher affinity is like thinking the partner’s prince charming - you’ll take him whenever you find him - so lower Kd. But if you think there’s still someone better out there you might “hold off” unless there’s so many of that okay-ish match that you “give in” - so higher Kd.
This is a really important, though potentially confusing, concept to remember:
higher affinity → lower Kd
lower affinity → higher Kd
If you’re wondering why we don’t use the association constant, Ka, which is the inverse of Kd (so Kd = 1/Ka) it’s because that’s not in concentration units - look at the figures if you’re interested, but for now let’s get back to the marriages (and divorces) (and re-marriages)(and re-divorces…)
What’s really happening is that each time 2 molecules collide, they have a certain probability of sticking together. And then, depending on how much they like each other, they can stay stuck for various lengths of time. The more they like each other, the higher the affinity, so they’ll stay stuck. But, if you have a lower affinity, they’ll keep coming apart, so you’ll need more standing by to take their place. So a higher affinity corresponds to a lower Kd.
And what does Kd come from? As an equilibrium constant (more on what this means in a second), Kd is a sort of “endpoint” measurement.
finished in comments
(note: video typo-fixed (sorry!), added an analogy, refreshed, and recut mainly from past videos)
blog form: bit.ly/dontignorewater & bit.ly/charge_intuition
Water molecules are really “sticky” towards each other because they’re highly polar - basically atoms (like the 2 hydrogens and the oxygen in H₂O) have smaller parts called subatomic particles - positive-charged protons & neutral neutrons in a dense central nucleus with a cloud of negatively-charged electrons whizzing around. Atoms can form bonds by sharing electrons, but they don’t always share fair. Oxygen is much more electronegative (electron-hogging) than hydrogen, so it pulls their shared electrons closer to itself, making the O partly negative (δ-) and the Hs partly positive (δ+). And opposite charges attract, so the H’s of one water molecule can hang out with the O of another. Each water molecule can form up to 4 “hydrogen bonds” (H-bonds) with other molecules.
Unlike the strong covalent bonds holding the Hs to the O in each individual water molecule, these inter-molecular bonds are weaker, so they can stick and unstick. As long as water molecules have sufficient energy to temporarily break free of the bonds, the water molecules can move around & explore, breaking and forming interactions with other water molecules as they travel. These water molecules can occupy many different “states” and the term we use to describe this is high entropy (aka “disorder” or “randomness”)
But they can’t interact readily with hydrophobic molecules, which are characterized by being nonpolar (electrons are evenly distributed so there aren’t partly or fully charged regions) and thus don’t offer tantalizing charge opportunities. So each water molecule that has to be next to part of a hydrophobe has part of its stickiness “hidden” and is limited in its binding opportunities - it can occupy fewer “states” and thus has lower entropy (is less disordered)
We use a term called free energy, G, to describe how “comfy” a molecule is - it takes into account entropy (S) (that disorder) as well as something called enthalpy (H), which has to do with bond energy. Molecules interact spontaneously in ways that make them comfier (reduce the free energy) and we describe this using the equation
ΔG = ΔH - TΔS
This says that the change in (abbreviated delta, Δ) free energy equals the change in enthalpy (do the new interactions have more or less energy than the old ones) minus temperature (in Kelvin) times the change in entropy (do the molecules have more freedom now?) Negative G is “good” (means a reaction is favorable) - and you can get to it if the new bonds are much less energetic (easier to hold together) and/or the molecules gain freedom of movement.
Say you have a sea of water molecules and you toss in a hydrophobe. Some of the water molecules will have to hang out with it - there’s no getting around that - and because there’s only so much space around a water molecule, it’ll have to break up some of its water-water bonds to do this. And this requires putting in energy (without getting a better bond in return) so you have a + ΔH
Now imagine you keep dropping in hydrophobes. Each time a water molecule swaps an interaction with water for an interaction with the hydrophobe, it has to “spend energy,” so you keep racking up “enthalpic penalties.” and it loses binding opportunities - so you have “entropic penalties” as well. But, if those hydrophobes all cluster together, fewer water molecules will have to give up the “better” opportunities offered by water.
But the hydrophobes have no intrinsic desire to clump themselves together - instead it is the water molecules around them that kind of shepherd them together - by “reaching out” to other water molecules in their network (remember each water molecule can form up to 4 H-bonds), they draw together (kind like how surface tension can lead to drops of water staying spherical - when hydrophobes combine, water molecules get released from the clathrate cage, leading to an increase in entropy (+ ΔS). And this offsets the energy you have to put in (ΔH) to break up the individual cages when you merge them. So the hydrophobic effect is ENTROPY-driven. And powerful - it’s the driving force of protein folding! more here: http://bit.ly/hydrophobiceffectPSA
Finished in comments
bit.ly/biochemistrystartersguide
Be able to recognize (both structure and shorthand) and draw the following functional groups:
Carbonyl
Hydroxyl (alcohol)/hydroxylate
Thiol (sulfhydryl)/thiolate
Amine
Amide
Carboxylic acid/carboxylate (-COOH/COO-)
Aldehyde (CHO)
Ketone
Ester
Ether
Thioester
Thioether
Methyl (Me)
Acyl
Acetyl (Ac)
Phenyl (Ph)
Phosphoryl (Ⓟ)
You should also know their properties & reactivity (which you can figure out if you know their structure!)
More to help you start out in biochemistry: thebumblingbiochemist.com/starters-guide-to-biochemistry & thebumblingbiochemist.com/365-days-of-science/biochemistry-things-worth-memorizing
Blog form (with figures): bit.ly/proteinstructure & bit.ly/protein_domains_motifs
note: adapted from past posts
for more detail, see http://bit.ly/allaminoacids & http://bit.ly/aminoacidstoproteins & http://bit.ly/peacepeptide
and I have a whole page on my website dedicated to protein stuff (just added today’s post to it): thebumblingbiochemist.com/lets-talk-science/amino-acids
Proteins are made up of long “polypeptide chains” of letters called amino acids linked backbone-wise through peptide bonds. There are 20 (common) genetically-specified amino acids, each with a generic backbone with to allow for linking up as well as unique side chains (aka “R groups” that stick off like charms from a charm bracelet). These chains fold up into functional proteins whose structure complements the tasks they carry out. They have several layers of structure that come from
1. the the order of amino acids linked up in the chain (primary structure)
2. how the backbones of those amino acids interact through hydrogen bonds to give common motifs like α-helices and β-strands (secondary structure).
3. How the side chains interact to give you structure on top of that structure (tertiary structure) and then
4. how different chains sometimes interact to give you quaternary structure (not all proteins have multiple chains and only those that do have quaternary structure
"Domain" and "motif,” which are both just ways we can refer to specific parts of a protein that are somehow "interesting" - functionally, structurally, evolutionarily, etc.
More on domains and motifs: bit.ly/protein_domains_motifs ; YouTube: youtu.be/7ejb6P6Fo-8
Protein domains are a bit like rooms in an apartment - they can be distinct rooms (structural domains) (think wall-separated rooms) and/or places in the house where you do specific things (functional domains) (think different areas of a studio apartment).
Domains are typically largish regions of a protein that serve a specific function (functional motif) &/or are structurally “independent*” (structural motif)
*by structurally-independent, I mean that they can often fold by themselves and be separated from the rest of the protein and still carry out their function (so you can study them on their own)
Protein motifs are a bit like furniture in a room - these may or may not have a known function ( e.g. zinc fingers typically bind DNA), but are often evolutionarily conserved & shared by many proteins.
examples: zinc finger, β turn, helix-turn-helix , β hairpin
Motifs may be located in a larger “domain” (e.g. a zinc finger motif in a protein’s DNA-binding domain)
more about protein structure: bit.ly/proteinstructure
resources:
UniProt: uniprot.org
Pfam: http://pfam.xfam.org
InterPro: ebi.ac.uk/interpro
Bernstein, N. K., Williams, R. S., Rakovszky, M. L., Cui, D., Green, R., Karimi-Busheri, F., Mani, R. S., Galicia, S., Koch, C. A., Cass, C. E., Durocher, D., Weinfeld, M., & Glover, J. N. (2005). The molecular architecture of the mammalian DNA repair enzyme, polynucleotide kinase. Molecular cell, 17(5), 657–670. doi.org/10.1016/j.molcel.2005.02.012
more on primary & secondary structure: bit.ly/proteinstructure ; YouTube: youtu.be/FFAhrp3EEoM
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
a non-exhaustive list of some key prefixes, suffixes, roots, etc. you may see (some lots!)
the bumbling biochemist (Bri Bibel), last updated 8/20/25
downloadable version: bit.ly/biochemistry_word_parts
blog: bit.ly/biochemwordparts
First things first – prefixes!
In addition to metric prefixes…
* mono-: single, one
* e.g. monomer (a single unit, a molecule acting by itself)
* bi/di (2), tri (3), tetr/quartr (4), pent (5), hex (6), sept (7), oct (8), non (9), deci (10)…
* oligo-: few, little
* e.g. oligonucleotide (a short nucleic acid chain, such as a PCR primer); oligopeptide (a short chain of amino acids)
* poly-: many
* e.g. polymer (a long chain of linked-together monomers), such as a polypeptide (a long chain of amino acids – a protein)
* multi-: multiple
* e.g. multimer (typically used to refer to a protein with multiple subunits/chains)
* pleio-: more
* e.g. pleiotropic (doing or affecting multiple things, potentially a drug doing more than you want)
* hypo-: under/below (remember hypo, below)
* e.g. hypoactive (less active than normal), hypotonic (having lower tonicity)
* hyper-: over/above (remember hyper, over)
* e.g. hyperactive (more active than normal), hypertonic (having higher tonicity)
* epi-: over/above
* e.g. epigenetics (field dealing with genetic information that is “above” the level of the DNA sequence, such as DNA modifications (histone methylation, etc.))
* infra-: below
* e.g. infrared (wavelengths of light with frequencies below those of red light)
* ultra-: above
* e.g. ultraviolet (wavelengths of light with frequencies above those of violet light); ultracentrifuge (a centrifuge that spins at speeds above those of a conventional centrifuge)
* supra-: beyond, over, above
* e.g. supramolecular (made up of multiple molecules (beyond a single molecule))
* meso-: middle, moderate
* e.g. mesothermic (living at moderate temperature)
* iso-: equal, similar
* e.g. isotonic (having equal tonicity); isomers (molecules that are similar but differ in some way, such as having different shapes)
* ortho-: normal, straight, correct
* e.g. orthosteric (binding in the same site as the “normal” binding partner)
* allo-: other
* e.g. allosteric (binding at another site (not the “normal” binding site) and/or something happening at one site that has effects at another site)
* endo- & ento-: inside (remember enside)
* e.g. endogenous (something that is made by an organism rather than taken in from outside, such as hormones the body produces)
* exo- & ecto-: outside (remember exo, external)
* e.g. exogenous (something that is added “artificially” or from the outside, such as pharmaceutical drugs someone takes)
* auto-: self
* ex. autotroph (an organism able to make its own food)
* inter-: between
* e.g. intermolecular interactions (interactions between different molecules)
* intra-: within
* e.g. intramolecular interactions (interactions between different parts of the same molecule)
* trans-: across
* e.g. transmembrane proteins (proteins that go across/span membranes)
* cis-: same
* e.g. cis-regulatory elements (regulatory, “noncoding” regions of DNA such as promoters on the same piece of DNA of a gene that regulate expression of that gene)
* hetero-: different, other
* e.g. heterogenous (made up of multiple things and/or not consistently distributed); heteromer (a complex with different components)
* homo-: same
* e.g. homogenous (made up of a single thing and/or consistently distributed); homomer (a complex with multiple copies of a single component)
* amphi-/ampho-/ambi-: both
* e.g. amphoteric (able to act as both an acid and a base); amphiphilic (having both hydrophilic and hydrophobic parts)
* per-: through
* e.g. permeable (allowing things through)
* semi-: partly
* e.g. semipermeable (allowing some things through but not others)
* pan-: all
* e.g. pan-genome analysis (an experiment looking across the entire
finished below
solution)
* The more protons there are, the more acidic, and the lower the pH
* The fewer protons there are, the more basic/alkaline, and the higher the pH
The pKa is the pH at which half the copies of an acidic site is deprotonated (and thus in conjugate base state) and half is protonated (and this in conjugate acid form)
* At a pH below the pI, there’s more protons available to take, so more conjugate acid
* At a pH above the pI, there are fewer protons available, so more conjugate base
The pKR is just the pKa of an amino acid R-group (side chain)
The pI (isoelectric point) is the pH at which a molecule is net neutral (has no net charge)
* At a pH below the pI, a molecule is positively-charged
* At a pH above the pI, a molecule is negatively-charged
pI tells you about charge. pKa does NOT. Instead, it tells you about protonation. And at a single site. The molecule may be negative, neutral, or positive when that site is protonated.
The pI comes from the combination of acid/base groups in a molecule
The pH comes from the contributions of all acids/bases in a solution
Something can only act as an acid in its protonated state.
But the stronger the acid is, the less likely it will be to be in that state!
So the more likely it is to be in the deprotonated state, where it can only act as a base!
If the molecule is neutral in its protonated state (conjugate acid), it will be negatively-charged in its deprotonated state (conjugate base).
But if the molecule is positively-charged in its protonated state (conjugate acid) it will be neutral in its deprotonated state (conjugate base).
The happier a molecule is to be in a state, the more likely it will be in that state. Resonance and inductive effects can make molecules happy, so resonance etc. that stabilizes one state will make it favorable, even if it comes with charge.
We call amino acids basic or acidic based on what their neutral form acts as - even if their predominant form acts the opposite! So, for example, you pretty much always find arginine in its protonated, acid state because its neutral form is a stronger base, so we call it basic (and show it as blue). On the other hand, you pretty much always find aspartate in its deprotonated, base state because its neutral form is a stronger acid, so we call it acidic (and show it as red)
more on pH, pKa, and the Henderson-Hasselbalch equation: http://bit.ly/phacidbase & http://bit.ly/phbuffers
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
Blog form: bit.ly/proteinstructure
Note: recut & refreshed from past videos for my 2025 Biochemistry course
More about Ramachandran plots: youtu.be/fli3CVXJAyo
resources: Ramachandran tutorial: http://bioinformatics.org/molvis/phipsi & http://tinyurl.com/RamachandranPrinciple
& Tutorial:Ramachandran Plot Inspection - Proteopedia, life in 3D proteopedia.org/w
Tutorial:Ramachandran_Plot_Inspection
http://bioinformatics.org/molvis/phipsi/?fbclid=IwAR2O2hRdLXtazAqPStaiUAQueCvrWuWGLOcJ7uW32_0m_dMtmBVVwlFw33s
proteopedia.org/w/Tutorial:Ramachandran_Plot_Inspection
proteopedia.org/w/Ramachandran_Plot
posts on proteins and amino acids: thebumblingbiochemist.com/lets-talk-science/amino-acids
YouTube channel on amino acids: youtube.com/playlist?list=PLUWsCDtjESrFQoCEsEmZX6NxnwlHzjHZ6
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
Note: recut & refreshed from past videos
All protein backbones have limited flexibility because the peptide bonds linking them together get stabilized by resonance (electron delocalization where “extra” electrons are shared among more than just 2 atoms), which can only happen if Ca, N, & O are in the same plane, so you end up with a chain of planes where you can only rotate at certain places in the backbone (C-Cα (psi) & Cα-N (phi)). And even those rotations are restricted by steric hindrance (you can’t have atoms colliding with one another so bulky side chains restrict movement more). http://bit.ly/aminoacidstoproteins
But even that twisting is restricted, depending on the nature of the side chain because of “steric hindrance” - that’s basically a fancy way of saying 2 things can’t be in the same place at once (even if they’re super super small). Bulky things need more space, leaving them with fewer available ways to move without hitting something - like the atoms of the peptide backbone. So bulkier side chain → more steric hindrance
The thing about glycine is that its side chain is just an H - which is pretty damn small - movement-wise it’s like there’s barely anything there at all! As a result glycine has very low steric hindrance, so glycine residues are very flexible (remember residue’s just what we call an amino when it’s in a peptide chain so has lost that water-equivalent (I don’t mean to harp on about this I was just confused about it for a really long time but embarrassed to ask!))
So glycine’s smallness lets its backbone take on awkward angles that would be major no-nos for other amino acids. You can see this if you look at a Ramachandran Plot. When we’re solving a crystal structure (more later) we often check that the angles are geometrically solid & one of the things you’ll see in the “report card” for a structural model is “Ramachandran outliers” - atoms in the model that have suspicious angles. Usually it’ll be reported as “non-glycine,” “non-proline” Ramachandran outliers - basically glycine can “break the normal rules” because it’s so small (its side chain’s just an H) so it’s ok to find it at weird angles. Proline can also break the rules, but instead of being able to move lots more ways, like glycine, it just has “different rules” - it’s restricted to different angles. Here’s the link for the paper in the figure: doi.org/10.1002/prot.10286
Glycine’s “loosey-goosey-ness” makes it good for flexible regions of proteins BUT bad for places you need strong structure. So it’s often found in sharp turns leading into or out of more orderly structures like helices & sheets.
It’s typically only found in small amounts in protein, though it is the most abundant in the weird triple-helices of the protein collagen that helps make our skin stretchy but sturdy. But it’s found a lot of other places too. In its free form it acts as a neurotransmitter - a chemical messenger relaying news throughout the brain. And it is a member of the antioxidant tripeptide glutathione, which helps control oxidation status in our bodies.
The other Ramachandran weirdo is Proline. Proline’s backbone N can’t form them H-bonds. Normally, the generic backbone offers 2 locations for H bonding. The carbonyl (C=O) provides an H-bond acceptor in the form of the O and the amino group provides an H-bond donor in the form of the N-H. Therefore, backbones can interact through H-bonds to give a protein its “secondary structure” (common “structural motifs” like helixes, sheets, etc) http://bit.ly/insulindiabetes
BUT Proline’s N doesn’t have this H because it’s “been replaced” by a bond to side chain. So it can’t act as a donor. Thus, it doesn’t want to form α-helixes, and if it’s in them it’ll make them kinky. Proline can also make other places kinky because its side chain contortion “locks” the N-Ca bond in place, leading to limited backbone flexibility - even limited-er than usual!
resources: Ramachandran tutorial: http://bioinformatics.org/molvis/phipsi & http://tinyurl.com/RamachandranPrinciple
& Tutorial:Ramachandran Plot Inspection - Proteopedia, life in 3D proteopedia.org/w/Tutorial:Ramachandran_Plot_Inspection
http://bioinformatics.org/molvis/phipsi/?fbclid=IwAR2O2hRdLXtazAqPStaiUAQueCvrWuWGLOcJ7uW32_0m_dMtmBVVwlFw33s
proteopedia.org/w/Tutorial:Ramachandran_Plot_Inspection
proteopedia.org/w/Ramachandran_Plot
My typical strategy:
1. Find molar ratio A– (since it’s easiest to find directly from H-H)
2. Use that value to find fraction A– or HA
3. Convert to %, concentration, etc. as needed
PDF: drive.google.com/file/d/17DJ9pX1Bw7chOE7o6sVA4M30VPwKwYt4/view?usp=drive_link
blog: http://bit.ly/phbuffers
blog: http://bit.ly/phbuffers
note: mostly adapted from past videos
If you want to know what proportion is protonated or deprotonated, you don’t get that directly. Instead, you get a molar ratio…
The Henderson-Hasselbalch equation is:
pH = pKa + log([A⁻]/[HA])
We can rearrange that to get the molar ratio (MR):
[A⁻]/[HA] = 10^(pH-pKa)
This is telling us parts A⁻ per 1 part HA. But what we usually want is parts A⁻ per *total* parts. To get that, we need to divide the molar ratio by the molar ratio + 1.
So,
fraction A⁻ = MR/(1+MR)
or
A⁻ = 10^(pH-pKa)/(1+ 10^(pH-pKa))
and
fraction HA = 1/(1+MR)
or
HA = 1/(1+ 10^(pH-pKa))
For example, if 10^(pH-pKa) = 0.5 that doesn’t mean 1/2 is deprotonated. Instead, it means that 0.5/(0.5+1) = 1/3 is deprotonated!
If you want percents, just multiply by 100%
Or if you want concentrations, multiply by the total concentration.
If you want mol, multiply those concentrations by the volume
And always always always remember to do a quick reasonability check after each calculation you do. Does your answer make sense?
The pKa is the pH where 1/2 is protonated, 1/2 deprotonated (e.g. the molar ratio is 1) If you’re above the pKa, less than half should protonated because there are fewer protons around. For each 1 pH unit away from the pKa you get 10-fold more of one of the forms (i.e. 10-times more deprotonated for each pH unit above the pKa and 10-times less deprotonated for each pH unit below the pKa).
For more on the Henderson-Hasselbalch equation: http://bit.ly/phbuffers YouTube: youtu.be/VRJV2FTOUeM
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
I obtained a PhD from Cold Spring Harbor Laboratory (CSHL)'s Graduate School of Biological Sciences in Cold Spring Harbor, New York, graduating in October 2021. There, in the lab of Dr. Leemor Joshua-Tor, I used a combination of biochemical, biophysical, and structural methods to gain a better understanding of how gene expression is regulated through the RNA interference (RNAi) pathway (you can learn more about RNAi under "Let's Talk Science").
I then (April 2022-July 2023) carried out postdoctoral research using biochemistry and chemical biology to explore translation (protein-making) in Danica Fujimori's lab at the University of California, San Francisco (UCSF). I then (July 2023 - July 2025) was a Visiting Professor of biochemistry at St. Mary's College of California.
I graduated from Saint Mary's College of California (SMC) in 2016 with a B.S. in Biology. While at SMC, I became actively involved in biochemistry research, studying enzyme kinetics in the lab of Dr. Jeffrey Sigman. I find science fascinating and my goal is to help make science accessible to everyone, so that they can share in the wonder. I am interested in exploring the use of various forms of science communication and advocating for the advancement of women and underrepresented minorities in STEM (Science, Technology, Engineering, and Mathematics). I served as a Student Ambassador for the International Union of Biochemistry and Molecular Biology (IUBMB), a founding member of the IUBMB Trainee Initiative, and an Ambassador for Trainees.
You can follow me (@bumblingbiochemist.bsky.social on Bluesky. You can also catch Broadcasts of the Bumbling Biochemist on Facebook, Instagram and YouTube.
LinkedIn Profile: linkedin.com/in/bbibel
More about me and my philosophy: http://bit.ly//pushelectronsnotppl thebumblingbiochemist.com/about/about-me
When it comes to amino acids, some key things are to:
- Know their abbreviations (3-letter and 1-letter)
- Know (w/o looking at structures) & identify their functional groups
- Classify the amino acids’ R groups as:
* polar/nonpolar
* hydrophobic/hydrophilic
* acidic/basic/neutral (non-ionizable)
* aromatic/aliphatic
- Know which amino acids…
* are nucleophilic
* are weirdos (and why)
This prepares you to see how amino acid structure directly influences protein structure & function
Much more on amino acids: http://bit.ly/allaminoacids ; YouTube: youtube.com/watch?v=Os6VVovCR8U
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
Note: Recut and refreshed from past video
blog form: http://bit.ly/bbproteinchromatography & bit.ly/aktainaction
A few common types used for proteins are affinity chromatography (AC), ion exchange chromatography (IEX) & size exclusion chromatography (SEC) and they use different resins. If the protein likes the resin (solid phase) more than it likes the liquid it came in with it’ll stick to the column. It might like the column because it’s oppositely-charged (this is the basis behind ION EXCHANGE CHROMATOGRAPHY (IEX)). The protein might also like the column more because it has some, more specific, special feature (like an engineered tag) that matches a special feature sticking off of the resin beads. This is how AFFINITY CHROMATOGRAPHY works. (note: the beads are usually porous - they have little tunnels running through them - and the affinity groups can stick out into these tunnels as well so you have more binding opportunities)
Both IEX & AC rely on the protein you want sticking to the column, while the other proteins flow through, then competing your protein off with salts and/or mimics or changing the pH to change the charge. But In SIZE EXCLUSION CHROMATOGRAPHY (SEC) you don’t want the protein to stick to the resin. Instead, you separate proteins by making smaller ones travel further because they can enter secret tunnels in the resin beads that big proteins can’t get into.
AKTA’s kinda like the “Google” of the protein chromatography world in that it basically dominates the market and if you say AKTA other protein-purifiers know what you mean. The AKTA takes our protein sample and pumps it onto a column (which it’s gotten ready by flowing a bunch of buffer (pH-stable salt water) to “equilibrate” it. It then washes the column with the buffers we tell it to. We have 2 system pumps so you can use 2 different buffers that send liquid first into a mixing chamber so you can mix them if you want to make a gradient for a gradient elution to introduce the “competitor” that will push your protein off the column (e.g. have a no salt & a high salt or a no imidazole & high imidazole (for His tags) you can mix). Or you can just use 1 for an “isocratic elution” like for SEC when you don’t need to change the buffer.
The AKTA also allows us to control the flow rate (much easier than trying to twiddle with the stopcock in gravity flow). As the name implies, it *can* go fast, but you don’t always want it to or you’ll crush the resin in the column! Each column has different maximum flow rates. For the SEC columns I use, I typically run at ~0.7mL/min, which is actually pretty slow… And the fastest columns I run are only ~4mL/min. The times when the pumps are working their hardest is when doing pump washes. During those it’s pumping at 20mL/min, but it’s not going through any columns so you don’t have to worry about hurting them.
And, just in case, the system has pressure monitors at the entrance and exit of the columns. If the pre-column pressure (pressure going in) is too high it can damage the column hardware (the cylinder itself) & if the delta column pressure (difference between pressure going in & going out) is too high it could the resin in the column and/or the filter on top of the column are clogging up and generating dangerous pressure that can hurt the resin. So the AKTA will stop and alert you.
The liquid flows through lots of little tubes that offer different flow paths. Which path the liquid takes is dictated by lots and lots of valves to go with those lots and lots of little tubes. It’s kinda like a subway system that can change the tracks. So we can direct liquid into different columns and, when it comes out of the columns into the waste or a fractionater which collects it to “keep.”
We choose which fractions we actually want to keep based on the chromatogram. This is where we see the evidence of our protein coming out in the form of a peak in the 280nM wavelength absorbance. Proteins (in particular tryptophan, tyrosine, and phenylalanine) absorb that type of light so you can tell when protein’s elute because they “steal” that wavelength from the light spectrum. A UV monitor on the path between the bottom of the column & the fractionater measures this. And the computer shows this to us as a peak. more here: bit.ly/2yzyi4w
FPLC looks a lot like a related technique, HPLC. HPLC stands for High Performance Liquid Chromatography. HPLC uses higher pressures but lower flow rates. It’s usually used for small chemical compounds and sturdier beads that can withstand those high pressures.
Finished in comments
I learned this method from a postdoc, Elad, in grad school and will forever be grateful for him taking the time to walk me through things and show me how to set up calculations. I know not everyone is so lucky to have a mentor like that, so I hope this can help serve a similar role to someone(s).
Link to blog post with download links and more: bit.ly/SLICworkflow
Direct links:
Spreadsheet: lmu0-my.sharepoint.com/:x:/g/personal/bri_bibel_lmu_edu/EQSClWehvDxCnCgJ-BZjc9QBFkJOEFfCQ6hK6r9c4Wy-KQ?e=0YnlCB
Written guide: lmu0-my.sharepoint.com/:w:/g/personal/bri_bibel_lmu_edu/EZKhZaHa7gpCkJM9M2JcTLoB0WgwSZXd2Hd_TS5XAW83DQ?e=v85V99
You will need to adapt it to your polymerase, concentrations, etc. but hopefully this will serve as a starting point. And then you can optimize.
As always, I'm doing this solely for the love of helping others and paying it forward. So I sincerely hope these resources can help!
Resources:
More on SLIC: blog: bit.ly/SLICworkflow ; YouTube: youtu.be/Wy_kCoyBdZc
More on molecular cloning: http://bit.ly/molecularcloningguide & youtube.com/playlist?list=PLUWsCDtjESrESdmbna9aUt-VXwaQWMONu
Here’s the original SLIC paper: Li, M.Z., Elledge, S.J. (2012). SLIC: A Method for Sequence- and Ligation-Independent Cloning. In: Peccoud, J. (eds) Gene Synthesis. Methods in Molecular Biology, vol 852. Humana Press. doi.org/10.1007/978-1-61779-564-0_5
Here’s a nice blog post from Addgene on SLIC: Plasmids 101: Sequence and Ligation Independent Cloning (SLIC), By Mary Gearing, 2015 blog.addgene.org/plasmids-101-sequence-and-ligation-independent-cloning
Here’s a comparison of SLIC, Gibson, and a few other methods: J5 manual, The SLIC, Gibson, CPEC, and SLiCE assembly methods (and GeneArt® Seamless, In-Fusion® Cloning) j5.jbei.org/j5manual/pages/22.html
More on PCR: http://bit.ly/pcrtrain & youtu.be/GZSLfECgW3Q
PCR playlist: youtube.com/playlist?list=PLUWsCDtjESrGDJ0GTWcdeDBCvBdU7FdUy
Optimizing PCR reactions: blog: bit.ly/pcrspecificity ; YouTube: youtu.be/vLauaHFixQs
Promega has a nice guide on optimizing your PCR: promega.com/resources/guides/nucleic-acid-analysis/pcr-amplification/#general-considerations-for-pcr-optimization-6f575242-fe99-4fcb-88c5-b205ad7becc7
More on agarose gel electrophoresis: bit.ly/agarosegelcompare & youtu.be/vbuxf3rcMxg
More on nucleic acid spin columns: http://bit.ly/spincolumns & youtu.be/fz2OpjxQKKM
more on transformation & heat shock: blog: http://bit.ly/transformheatshock ; YouTube: youtu.be/3C6X2a7xWVw
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
More here: bit.ly/why_program
In addition to Python itself, here are key Python modules, programs, etc. to install, learn, etc.
• Numpy – for dealing with numerical data
• Pandas – for working with data tables
• Matplotlib – for plotting
• Seaborn – for plotting (based on matplotlip but sleeker)
• Plotly for interacive graphs
• Jupyter notebook/Jupyter lab – for organizing & working interactively with code
• Google colab – for organizing & working interactively with code
• Anaconda – for installing packages and managing virtual environments
• Github – for finding and saving code
• Google – for when you get stuck!
• Basic command line – bash etc. for navigating directories (folders), running scripts, etc.
Python basics
• data types: strings, floats, integers, lists, dictionaries, tuples, etc.
- converting between types
• indexing, slicing, splitting
- beware that indexing for most things in Python starts with 0 and doesn't include the end
o e.g. [0:2] will get you the first and second values
• functions
- a lot of the time there will be optional arguments that make the docstrings look scary but really you don't need to worry about because they have default values set
- *args and **kwargs lets a function take an unknown number of input values or key-value pairs
o can use * or ** when giving input to give a function a list (*) or dictionary(**) to unpack
• scope (global, local, etc.)
- when you're in a function you are "isolated" from the rest of the file so you need to pass it what you want it to have in the arguments & then have the function return back to you what you want
• for loops & while loops
- check out enumerate function for iterating
• Boolean statements (e.g. True/False)
- beware that = is to assign values, but == is to check whether things are equal
• comments (start with #) and docstrings (surrounded by ' ' or ''' ''')
- docstrings are basically like the “manuals” for using a function
• use help(thing_you_want_help_with) to get docstrings
Python basics +
• lambda functions
- another one of those things that looks super scary but is really super helpful, especially when working with pandas
• list comprehension
- lets you make lists from lists (filtering out and/or modifying the values)
• string formatting (check out f strings)
• regular expressions
- search for specific strings
Key modules
Modules contain extra functions, etc. that aren’t included in core Python & are tailored for various specialized purposes (graphing, calculating, etc.)
• pandas (for working with data tables - think Excel on steroids)
- import pandas as pd
• numpy (for working with arrays of numbers, plus a bunch of helpful science/math functions)
- import numpy as np
• matplotlib (for plotting things)
- import matplotlib.pyplot as plt
- %matplotlib inline
• seaborn (for plotting things more prettily & easily)
- import seaborn as sns
• plotly
- import seaborn as sns
• Biopython (for working with DNA and protein sequences, etc.)
Here’s a link to CopyClip apps.apple.com/us/app/copyclip-clipboard-history/id595191960?mt=12
And info about PC ones - zapier.com/blog/best-clipboard-managers
more (hopefully) helpful random practical lab tips & tricks: bit.ly/lab_tricks_page
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
blog form: http://bit.ly/wikieditingintro
slides: bit.ly/wikipedia_slides
If you don’t know where to start, check out Wikipedia’s tutorial. It will help you learn the ropes. (And don’t worry, you can practice in a “Sandbox” before you’re ready to take things live). en.wikipedia.org/wiki/Help:Introduction
Also, visit the Community Portal. en.wikipedia.org/wiki/Wikipedia:Community_portal
At The Teahouse, you can ask your novice questions without judgement. en.wikipedia.org/wiki/Wikipedia:Teahouse
When you make a new page, it gets on editors’ radars, and people will quickly pitch in to help (such as fixing the title of my Mol* article (from Molstar, which I had put because I was having trouble changing it). To keep it on your radar, it also automatically gets added to your watchlist, which you can find in the upper right corner, and which provides a running log of pages you’ve edited and/or interested in.
To help your article get more reach, find relevant pages and add links to it. You can also add categories to your page (the easy way to do this is click on the icon with 3 horizontal lines and select the categories option) and ad your page to relevant links.
For a break, check out Wikipedia’s article on Unusual articles! en.wikipedia.org/wiki/Wikipedia:Unusual_articles
And, if you’re looking to learn something random, click the “Random article” button on left-hand menu. Each time you do, you’ll get something new!
Here’s more I wrote a while back . . .
Wikipedia is free to edit, so we each can play a part, so today, as I work on 2 articles I’m really excited about, I’m re-sharing a guide for those who don’t know where to start! Like it or not, Wikipedia is one of the main sources people turn to for information - on everything from people, to companies, to molecules. So it’s really important that that information is accurate - and the accuracy of the information depends upon all of us because the content is written and edited by everyday folks like us who’ve decided to take to the keyboard, volunteer a little time, and become a “Wikipedia editor.” ⠀
⠀
Don’t worry, it’s not a huge time investment (unless you want it to be) - you can edit as much or as little as you want - create whole new articles or just fix some punctuation. The biggest investment time-wise is getting started and getting used to working in the Wikiverse. ⠀
⠀
Because all the content is added by volunteers, each with their own interests, the information available tends to skew towards those peoples’ interests. So, there might be a ton on some obscure video game and barely anything on your favorite chemical. And, super importantly, there’s typically a lot more on Caucasians rather than people of color and men more than women (fewer than 18% of Wikipedia biographies are about women as per https://lat.ms/2wQESGA )⠀
⠀
And this brings me to how I first got started editing - about 7 years ago I saw an article about a UK physicist Dr. Jess Wade, who had devoted herself to the cause of increasing representation for female scientists on Wikipedia, writing literally hundreds of new articles a year. She was recently interviewed by Emily Kwong on NPR Shortwave, http://n.pr/3nw0Hjh which gave me the extra push I needed to work on a couple of new articles.
⠀
I didn’t set my sights as big as Dr. Wade, but figured I could maybe write an article or too. My first article (in March 2020) was on Virginia Minnich, who discovered an abnormal form of hemoglobin, hemoglobin E that can cause a blood disorder: http://bit.ly/2OneZGh en.wikipedia.org/wiki/Virginia_Minnich ⠀
Finished in comments
1. A header/identifying line that starts with a carat, and then provides information about the following sequence (accession codes for various databases, gene and/or protein name, species it comes from, etc.) followed by
2. The corresponding sequence, using IUBMB/IUPAC conventions for 1-letter amino acid and nucleotide abbreviations, typically broken into 60-80 characters per line
There can be multiple sequences per FASTA file–they’re distinguished from one another because they each start with a carat. The header line shouldn’t have any hard returns (new line characters), but the sequence can. If you need to copy it somewhere without line breaks, here’s a quick way: removelinebreaks.net
Blog: bit.ly/fastaformat
FASTA files are plain text files. They may end in a variety of extensions (.fasta, .fas, .fa, .fna, .ffn, .faa, .mpfa, .frn) but can be opened in any text editor
- Different file extensions may correspond to different sequence types (i.e. nucleic acids, proteins, multiple proteins), but the main ones are generic: .fasta, .fas, .fa
- There’s also fastq which has sequencing quality info. I dealt with those a lot during my postdoc
You can download FASTA sequences for proteins from UniProt and for nucleic acids from NCBI Nucleotide and associated databases (GenBank, RefSeq, etc.). The download buttons are in various places in the nucleic acid databases, but often on the upper right of an entry.
From UniProt, when you’re on an entry page, go to the Sequence Selection and you will see a button that somewhat confusingly says Download. I say confusingly because, instead of downloading something, it takes you to a screen with the FASTA-formatted text. Just right click and “Save As” if you want to download it, or you can copy and paste it somewhere you desire, which is often enough.
Alternatively, you can download multiple sequences from your basket
You can make any sequence* into FASTA format by adding a carat-ed line above it with a name or description of your choice
*Following IUBMB/IUPAC conventions for 1-letter abbreviations
This includes,
For proteins:
- X = any amino acid residue
- * = translation stop
- - = gap (indeterminate length)
For nucleic acids:
- N = any nucleotide residue
- Y = pYrimidine (C, T, or U)
- R = puRine (A or G)
- K = G, T, or U (bases that have Ketones)
- M = A or C (bases with aMino groups)
- - = gap (indeterminate length)
- some other weird ones too, but Wikipedia has a nice table: en.wikipedia.org/wiki/FASTA_format#Sequence_representation
When you get a sequence from a database, however, there will be more information in that header line.
For UniProt (UniProtKB), the header line will follow the following format:
carat (which YouTube doesn't allow...) db|UniqueIdentifier|EntryName ProteinName OS=OrganismName OX=OrganismIdentifier [GN=GeneName] PE=ProteinExistence SV=SequenceVersion
note: the GeneName is optional and isn’t actually bracketed
It starts with telling you which database it’s coming from. UniProtKB (UniProt Knowledgebase) consists of 2 databases: UniProtKB/Swiss-Prot (abbreviated sb) and UniProtKB/TrEMBL (abbreviated tb). Basically,
sp: Swiss-Prot - manually curated (more reliable, more info, often something people have actually experimented with a lab)
tr: TREMBL - auto curated (less known, be more cautious, most stuff from less-studied organisms is here)
Much more here: ebi.ac.uk/training/online/courses/uniprot-quick-tour/the-uniprot-databases
Then it gives you some identifying information, followed by a weird one, “Protein Existence.” Basically, when something pops up in nucleic acid sequencing data, it might seem like it’d make a protein but no one’s actually seen it. That could just be because no one’s cared to look (and/or had tools to do so) or it could be because it isn’t actually a functional protein.
PE = protein existence (# from 1 (most) to 5 (least) evidence it exists)
UniProt scores PE this way:
1. Experimental evidence at protein level (people have detected the actual protein)2. Experimental evidence at transcript level (people have sequenced the mRNA)3. Protein inferred from homology (its sequence is similar to proteins known to exist)4. Protein predicted (it has a reasonable open reading frame near a predicted ribosome binding site (RBS), etc.)5. Protein uncertain…
more here: uniprot.org/help/protein_existence
Finished in comments
MUCH more here: bit.ly/crisprcasscience
CRISPR genome editing strategies include:
* Conventional - Cas makes double-stranded DNA (dsDNA) break & leaves cell to fix
* NHEJ: Non-Homologous End Joining stitches the pieces together
* Typically causes “uncontrolled” insertions & deletions (indwells) leading to gene inactivation/“knockout”
* HDR: Homology-Directed Repair swaps in an alternative template with ends that match the cut site region
* Inefficient and limited mainly to actively-replicating cells
* “Second generation” CRISPR/Cas genome editing systems use a nickase Cas that makes a ssDNA break and is attached to a DNA-modifying enzyme
* Base editing: Alters a single nucleotide base for point mutations
* Typically through a deaminase domain
* Prime editing: Sticks in a new/alternate sequence provided as part of an extended guide RNA called a pegRNA (prime editing guide RNA)
* Uses an attached reverse transcriptase (RT) to make a DNA copy of the pegRNA-provided template
CRISPR genome editing challenges include:
* Editing accurately and efficiently
* Targeting desired sequences (e.g. overcoming PAM limitations)
* Expanding the range of edits (e.g. single base changes vs large insertions)
* Preventing off-target activity (unwanted edits)
* Delivering the machinery
* Targeting specific cell types
* Evading the immune system
Solutions include:
* Using nickase Cas proteins instead of fully-active ones to minimize off-target damage and cellular responses
* Changing the Cas to ones that are more efficient, higher fidelity, smaller, less immunogenic, and/or have more desirable PAMs
* “Mining” bacterial genomes to find new naturally-occurring Cas proteins
* Altering existing Cas proteins (through targeted mutations and/or directed evolution)
* Modifying the guide RNAs to decrease off-target binding
* Working to improve delivery systems (e.g. different LNP formulations, cell-specific targeting factors)
CRISPR isn’t just for genome editing anymore:
* Enzymes can be attached to catalytically-inactivated (“dead”) Cas proteins to allow for things like:
* Transcriptional regulation
* via transcription factors and/or epigenetic changes through chromatin remodelers
* CRISPRi, CRISPRa, CRISPRon, CRISPRoff
* Cas variants (e.g. Cas13, Cas7-11) can be used to target RNA for modification &/or inhibition
* CRISPR/Cas can be used for diagnostic purposes - Cas proteins can get activated to cleave reporter RNAs by binding to targeted sequences (such as viral sequences)
Recommended reading:
Doudna and Charpentier’s game-changer article:
Jínek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., & Charpentier, E. (2012). A programmable dual-rna–guided dna endonuclease in adaptive bacterial immunity. Science, 337(6096), 816-821. doi.org/10.1126/science.1225829
And a 2014 article from them giving perspective on the findings leading up to it:
Doudna, J. A.; Charpentier, E. The New Frontier of Genome Engineering with CRISPR-Cas9. Science 2014, 346 (6213), 1258096–1258096. doi.org/10.1126/science.1258096.
Some nice review articles, etc.:
Pacesa, M.; Pelea, O.; Jinek, M. Past, Present, and Future of CRISPR Genome Editing Technologies. Cell 2024, 187 (5), 1076–1100. doi.org/10.1016/j.cell.2024.01.042.
CRISPR 101: Cytosine and Adenine Base Editors. blog.addgene.org/single-base-editing-with-crispr (accessed 2025-05-17).
Macarrón Palacios, A.; Korus, P.; Wilkens, B. G. C.; Heshmatpour, N.; Patnaik, S. R. Revolutionizing in Vivo Therapy with CRISPR/Cas Genome Editing: Breakthroughs, Opportunities and Challenges. Front. Genome Ed. 2024, 6. doi.org/10.3389/fgeed.2024.1342193.
.
thebumblingbiochemist.com/365-days-of-science/cps1crispr
Penn press release: World’s First Patient Treated with Personalized CRISPR Therapy. pennmedicine.org/news/news-releases/2025/may/worlds-first-patient-treated-with-personalized-crispr-therapy?fbclid=IwY2xjawKW6-BleHRuA2FlbQIxMQBicmlkETFBdmIycVR6MDdjRzhjajRNAR4K0BajQkZ985aj0_opsILZ5pu_JFgGlp0mU6TC7STAtSZAF_fPVatGRC4cfw_aem_EBFMh_TEI1xNX5YP0xM8QQ (accessed 2025-05-19).
Musunuru, K. . . . Ahrens-Nicklas, R. C. et al. Patient-Specific In Vivo Gene Editing to Treat a Rare Genetic Disease. N Engl J Med 2025. doi.org/10.1056/nejmoa2504747.
More on CRISPR from me: http://bit.ly/crisprdoudna
More on amino acid metabolism and the urea cycle: youtu.be/Ce78en-LEcE & youtu.be/6X1uC6PDFmY
Urea cycle graphic: doi.org/10.5281/zenodo.14188989
Resources and references:
CRISPR 101: Cytosine and Adenine Base Editors. blog.addgene.org/single-base-editing-with-crispr (accessed 2025-05-17).
Pacesa, M.; Pelea, O.; Jinek, M. Past, Present, and Future of CRISPR Genome Editing Technologies. Cell 2024, 187 (5), 1076–1100. doi.org/10.1016/j.cell.2024.01.042.
Macarrón Palacios, A.; Korus, P.; Wilkens, B. G. C.; Heshmatpour, N.; Patnaik, S. R. Revolutionizing in Vivo Therapy with CRISPR/Cas Genome Editing: Breakthroughs, Opportunities and Challenges. Front. Genome Ed. 2024, 6. doi.org/10.3389/fgeed.2024.1342193.
Zhao, Z.; Shang, P.; Mohanraju, P.; Geijsen, N. Prime Editing: Advances and Therapeutic Applications. Trends in Biotechnology 2023, 41 (8), 1000–1012. doi.org/10.1016/j.tibtech.2023.03.004.
Anzalone, A. V.; Randolph, P. B.; Davis, J. R.; Sousa, A. A.; Koblan, L. W.; Levy, J. M.; Chen, P. J.; Wilson, C.; Newby, G. A.; Raguram, A.; Liu, D. R. Search-and-Replace Genome Editing without Double-Strand Breaks or Donor DNA. Nature 2019, 576 (7785), 149–157. doi.org/10.1038/s41586-019-1711-4.
Murray, J. B.; Harrison, P. T.; Scholefield, J. Prime Editing: Therapeutic Advances and Mechanistic Insights. Gene Ther 2025, 32 (2), 83–92. doi.org/10.1038/s41434-024-00499-1.
Chehelgerdi, M.; Chehelgerdi, M.; Khorramian-Ghahfarokhi, M.; Shafieizadeh, M.; Mahmoudi, E.; Eskandari, F.; Rashidi, M.; Arshi, A.; Mokhtari-Farsani, A. Comprehensive Review of CRISPR-Based Gene Editing: Mechanisms, Challenges, and Applications in Cancer Therapy. Molecular Cancer 2024, 23 (1), 9. doi.org/10.1186/s12943-023-01925-5.
Lazzarotto, C. R. et al. CHANGE-Seq-BE Enables Simultaneously Sensitive and Unbiased in Vitro Profiling of Base Editor Genome-Wide Activity. bioRxiv March 30, 2024, p 2024.03.28.586621. doi.org/10.1101/2024.03.28.586621.
Lichter-Konecki, U.; Caldovic, L.; Morizono, H.; Simpson, K.; Ah Mew, N.; MacLeod, E. Ornithine Transcarbamylase Deficiency. In GeneReviews®; Adam, M. P., Feldman, J., Mirzaa, G. M., Pagon, R. A., Wallace, S. E., Amemiya, A., Eds.; University of Washington, Seattle: Seattle (WA), 1993.
Longo, N.; and Holt, R. J. Glycerol Phenylbutyrate for the Maintenance Treatment of Patients with Deficiencies in Enzymes of the Urea Cycle. Expert Opinion on Orphan Drugs 2017, 5 (12), 999–1010. doi.org/10.1080/21678707.2017.1405807.
Diez-Fernandez, C.; Häberle, J. Targeting CPS1 in the Treatment of Carbamoyl Phosphate Synthetase 1 (CPS1) Deficiency, a Urea Cycle Disorder. Expert Opinion on Therapeutic Targets 2017, 21 (4), 391–399. doi.org/10.1080/14728222.2017.1294685.
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
PS - if your microwave is tall enough, you can also (carefully) melt it that way too
More random lab tips & tricks: bit.ly/lab_tricks_page & youtube.com/playlist?list=PLUWsCDtjESrFEAWZCRKJL7sMc6a_KgfLU
More about agar plates: http://bit.ly/agarbakery & youtu.be/5TlkpkngDFo
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com
An electrical heater hooked up to the steam jacket provides the heat energy needed to get the water to boil into steam that enters the chamber. You might think that the water needs to get to 100°C - that’s the boiling point of water, right? But actually, you need to get it to ~121°C (at least once the pressure builds to where you want it)! Before you go WTF? let me explain. 100°C is the boiling point of water at sea level. But pressure in an autoclave is much higher, which makes it much harder for water molecules to break free from one another. They have to wiggle a lot harder in order to overcome the pressure, and the harder they wiggle, the higher the temperature. So autoclaves are able to get really hot without just having all the liquids you’re trying to sterilize evaporate into nothingness. This is a similar concept to how pressure cookers work.
A great thing about autoclaves is that steam can get places your scrub brush bristles can’t which is great for cleaning awkward-shaped and porous things. BUT the autoclave is kinda like hand sanitizer in that it disinfects but it doesn’t clean - so you’ve gotta do some scrubbing first to get off the grimy gunk. This is especially important when you have tissue culture flasks where you often get a crusty ring of dead cells at the top of where the cells are swirling around. So, I manually scrubbed all 30 of those flasks and then loaded them up in a giant dishwasher. Only after both these washes are they ready to go in the autoclave.
Another thing you have to do before you run an autoclave - LOOSEN THE LIDS! You have to keep lids loose for a similar reason to why you don’t want to get hand sanitizer hot - gases take up more space than liquids. When molecules have enough energy to break free from other molecules and become a gas, it has enough energy to run away fast! But if there is no escape, the pressure will build up & potentially burst catastrophically - or at least the lid will get pretty permanently stuck onto your bottle. I found this out the hard way in undergrad when I didn’t loosen a lid quite enough. The lid got stuck on and we ultimately had to smash the bottle open with a hammer…. not good. So then I overcorrected and over-unscrewed the cap - and all the agar (a sort of sugar gel we use for making Petri dish bacterial homes) inside boiled out… not good. So, lids loose but not too loose!
Note: for alcohol-based hand sanitizers, the explosion risk is “high” because alcohols are volatile (they have low boiling points so they boil easily).
Another another thing to do before running the autoclave is stick on some autoclave tape. This tape has patterns that change color when they get hot. These patterns are usually stripes or tape brand logos, but I’ve always thought they should make autoclave tape with jokes - or at least motivational quotes or something… Kickstarter? Anyways, these color changes typically involve some chemical decomposing in an irreversible reaction. So they tell you something got hot enough. But not for how long it stayed hot enough - typically you need ~15-20 min of time at that heat. So tape’s not a sure-fire thing and there are other indicators you can use if you need to really really make sure - like if you’re sterilizing stuff at a hospital or something.
But tape’s cheap, it lets you leave a “note” on each of your things that they’ve been sterilized, and if your autoclave’s pretty reliable you can (hopefully) get away with it - especially when your autoclave digitally monitors what it’s doing. But, kinda like how your microwave heats the outer part of your food better so you might see your hot pocket steaming but then take a bite and the center’s cold, the autoclave heats the outside first because that’s what the steam sees first. So you need to give the heat time to transfer throughout the entire object, and you want to stick the tape close to the center if possible.
Finished in comments
MUCH more here: bit.ly/pdbstructures & http://bit.ly/thepdb
And more on the expansion to extended, 12-character, PDB IDs here: blog: thebumblingbiochemist.com/365-days-of-science/pdbids ; YouTube: youtu.be/U-MEtWaFNdI
Here’s some text adapted from a past post
If you’ve ever read an article discussing the structure of a protein (what it “looks like” at the atomic scale) or seen a picture of a protein model, you might have seen something like: PDB ID 2hhb, 2hho. Each of those “accession codes” is a specific name for a structure. In these cases, 2hhb is a structure of deoxyhemoglobin; 1hho is a structure of oxyhemoglobin. These codes (which would now be better written as pdb_00002hhb and pdb_00002hho) are more than just nicknames. In one sense they’re more like the structural biology equivalent of citing the photographer. Except they’re way cooler because they goes way beyond just giving credit where credit’s due! If you search the Protein Data Bank (PDB) for that name, it’ll pop right up and let you explore it. In 3D! You can actually play around with rotating it, coloring it different ways, etc. instead of just looking at a static snapshot. And you can find out more information about it (how it was “solved,” how reliable it is, etc.) as well as do a lot more with it, especially if you’re in the field and you know what you’re doing. So I thought I’d make a video walking you through how you can make the most of it even if you aren’t a hard-core structural biologist.
Let me step back a sec and explain what “structural biology” is because I didn’t even hear of the term until college, but now that I’m in the field I can sometimes forget that most people aren’t and might not know the term. more on structural biology here: http://bit.ly/cryoemxray but basically it’s the sub-field of biology that deals with trying to figure out what macromolecules (things like proteins, DNA, RNA, and mix-and-matched complexes of those components) look like (their form). And how that form fits with what they do (their functions).
To get an idea of why this might be relevant, think of how a spoon is good for scooping ice cream whereas and a knife is good for cutting cake. Macro means large, but these molecules are only large in comparison to other molecules. Compared to us, they’re tiny! So tiny that they’re invisible to our eyes, and even to our conventional microscopes. Therefore, we have to use fancy-dancy techniques like X-ray crystallography, cryo-electron microscopy (cryo-EM), and nuclear magnetic resonance (NMR) to visualize them.
The details are complicated and vary from technique to technique, but a key thing to know about all of them is that they don’t give you the actual atomic positions of each of the atoms that make up the molecule (individual carbons, hydrogens, oxygens, nitrogens, etc.). Instead they just give evidence of whereabouts those atoms are and then you have to use math-y stuff (or at least the computer does) to generate a “model” of the structure, placing in the atoms to fit the evidence. This allows scientists to generate “models” of the atomic structure of the thing they were looking at. And when someone does this, we say they’ve “solved the structure” of that protein or complex or whatever it was.
If they then want to publish a paper talking about that structure, they have to deposit the “coordinates” for it in the PDB. These coordinates are like addresses for each of the atoms in the model - plus some extra info like how confident the depositors are in the position. Basically, the structure-solvers have to upload enough data that everyone can evaluate the validity of the model for themselves and see if they can find any “hidden treasures” in it (features that the depositors might have missed, like evidence for a bound metal ion). And the way that anyone can look at it for themselves is through the PDB, so I want to tell you a bit more about how to use it.
Some of the specific features and viewing options will depend on what technique was used to generate the data. I haven’t done NMR, or cryo-EM (although everyone else in my lab has, so I hear about it a lot), but I have done X-ray crystallography, so I’m most familiar with that. Lucky for me, most of the structures in the PDB were solved using crystallography (although cryo-EM has gained steam in recent years thanks to technological advances).
much more on crystallography here: http://bit.ly/xraycrystallography2
Finished in comments
thebumblingbiochemist.com/365-days-of-science/pdbids
You can adapt old codes by adding “pdb_0000” to the beginning of them.
For example, 9cn3 becomes pdb_00009nc3
They expect they’ll run out of the 4-character ones by 2029, and then only issue the long ones, but the long ones work already for the ones that were originally issued 4-character ones. So you can use them with your molecular visualization software, etc.
But, the legacy PDB file format (.pdb) is incompatible with the longer codes. So, instead, you need to use the improved, PDBx/mmCIF (.cif) format. This format (Protein Data Bank Exchange macromolecular Crystallographic Information Frame) is already broadly-integrated and adopted as the wwPDB (worldwide PDB) standard. Its structure is much more flexible than the “legacy PDB” format allows the files to hold more information (bigger, more complex structures, and maps not just models).
You can learn much more about all this from official PDB sources.
- Webinar: Supporting Extended PDB IDs by the RCSB PDB, March 13, 2025: youtube.com/watch?v=qXTCJ21cr9o
- PDB File formats: wwpdb.org/documentation/file-formats-and-the-pdb
- Extended PDB ID With 12 Characters wwpdb.org/documentation/new-format-for-pdb-ids
- mmCIF user guide mmcif.wwpdb.org/docs/user-guide/guide.html
And if you want to learn more about the PDB, they’ve got a lot of info, as do I: bit.ly/pdbstructures & http://bit.ly/thepdb
As well as much more structural biology content: bit.ly/structural_biology & youtube.com/playlist?list=PLUWsCDtjESrGhwVxsRbTJdL-BEsN60RCs
As always - hope this helps. And happy transitioning!
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: http://thebumblingbiochemist.com


