LLM Building Blocks & Transformer Alternatives @SebastianRaschka
LLM Building Blocks & Transformer Alternatives  @SebastianRaschka
Uploaded October 2025 | Updated September 2026, 2 weeks ago
Resources:
- Understanding and Coding the KV Cache in LLMs from Scratch: magazine.sebastianraschka.com/p/coding-the-kv-cache-in-llms
- The Big Architecture Comparison: magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison
- Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers
- Reasoning From Scratch book: https://mng.bz/Nwr7

Description:
Learn the core components of modern transformer-based large language models (LLMs) and the practical techniques that make inference faster and cheaper.
We walk through Grouped-Query Attention (GQA), Multi-Head Latent Attention (MLA), and Sliding Window Attention (SWA), show where Mixture of Experts (MoE) fits into today’s architectures, and finish with a look at promising alternatives and hybrid models beyond standard transformers.

Chapters:

00:00 Intro
01:13 Main theme: bigger models and cheaper inference
02:13 Grouped-Query Attention (GQA)
05:44 Multi-Head Latent Attention (MLA)
09:51 Sliding Window Attention (SWA)
13:57 Mixture of Experts
17:01 LLM and transformer alternatives


#LLM #Transformers #DeepLearning #MachineLearning #Inference #MoE #GQA #MLA #SWA #KVCache
LLM Building Blocks & Transformer AlternativesL19.3 RNNs with an Attention MechanismL8.9 Softmax Regression   Code Example Using PyTorchDeep Learning News #4, Feb 20 2021L18.2: The GAN Objective13.4.3 Feature Permutation Importance Code Examples (L13: Feature Selection)L13.9.3 AlexNet in PyTorchL11.7 Weight Initialization in PyTorch   Code ExampleL8.1 Logistic Regression as a Single-Layer Neural NetworkThe Three Elements of PyTorchL6.2 Understanding Automatic Differentiation via Computation GraphsL17.3 The Log-Var Trick
Sebastian Raschka |

LLM Building Blocks & Transformer Alternatives

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER