A Practical Guide to Training a Small Language Model: Tokenizers, Training, and Real-World Pitfalls @socallinuxexpo
A Practical Guide to Training a Small Language Model: Tokenizers, Training, and Real-World Pitfalls  @socallinuxexpo
Uploaded April 2026 | Updated September 2026, 1 week ago
Talk by David vonThenen

socallinuxexpo.org/scale/23x/presentations/practical-guide-training-small-language-model-tokenizers-training-and-real

This session walks through how to build and train your own Small Language Model using open source tools. You'll learn how to select datasets, reuse BERT's tokenizer, design a smaller student model, and apply knowledge distillation to preserve accuracy. We'll share code, show common pitfalls, and help you deploy an efficient, real-world SLM that runs fast on modest hardware.
A Practical Guide to Training a Small Language Model: Tokenizers, Training, and Real-World PitfallsPorous by DesignDoug Comer: Software Distribution Now And Then - Why And How The Internet Changed EverythingPlanetNix Opening Ceremony Day 1Whats Cooking? Recipes For A Successful Developer PlatformZero Trust for Linux Admins with Open-Source IAMRoom 104 Sunday Mar. 08 - SCaLE 23xWelcome to SunSecConBallroom B Sunday Mar. 08 - SCaLE 23xYoud better start believing in supply chains because youre in oneRoom 101 Sunday Mar. 08 - SCaLE 23xBuilding a Postgres DBaaS with open source
Southern California Linux Expo |

A Practical Guide to Training a Small Language Model: Tokenizers, Training, and Real-World Pitfalls

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER