Uploaded January 2026 | Updated September 2026, 2 hours ago
Day 17/42: What Is Latency?
Yesterday, we explained inference.
Today, we talk about the wait.
Latency is the delay between your question and the full answer.
There are two parts:
Time to first token: when text starts appearing
Time between tokens: how fast it keeps flowing
Fast models feel smarter.
Slow ones feel broken, even if they’re correct.
This is why speed matters as much as quality in real products.
Missed Day 16? Start there.
Tomorrow, we control randomness: temperature.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#Latency #LLM #AIExplained #short
Day 17/42: What Is Latency?
Yesterday, we explained inference.
Today, we talk about the wait.
Latency is the delay between your question and the full answer.
There are two parts:
Time to first token: when text starts appearing
Time between tokens: how fast it keeps flowing
Fast models feel smarter.
Slow ones feel broken, even if they’re correct.
This is why speed matters as much as quality in real products.
Missed Day 16? Start there.
Tomorrow, we control randomness: temperature.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#Latency #LLM #AIExplained #short










