TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis) @YannicKilcher
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)  @YannicKilcher
Uploaded December 2025 | Updated September 2026, 2 weeks ago
Paper: arxiv.org/abs/2511.08923

Abstract:
Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation and forfeits its potential parallelizability. We introduce TiDAR, a sequence-level hybrid architecture that drafts tokens (Thinking) in Diffusion and samples final outputs (Talking) AutoRegressively - all within a single forward pass using specially designed structured attention masks. This design exploits the free GPU compute density, achieving a strong balance between drafting and verification capacity. Moreover, TiDAR is designed to be serving-friendly (low overhead) as a standalone model. We extensively evaluate TiDAR against AR models, speculative decoding, and diffusion variants across generative and likelihood tasks at 1.5B and 8B scales. Thanks to the parallel drafting and sampling as well as exact KV cache support, TiDAR outperforms speculative decoding in measured throughput and surpasses diffusion models like Dream and Llada in both efficiency and quality. Most notably, TiDAR is the first architecture to close the quality gap with AR models while delivering 4.71x to 5.91x more tokens per second.

Authors: Jingyu Liu, Xin Dong, Zhifan Ye, Rishabh Mehta, Yonggan Fu, Vartika Singh, Jan Kautz, Ce Zhang, Pavlo Molchanov

Links:
Homepage: ykilcher.com
Merch: ykilcher.com/merch
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: ykilcher.com/discord
LinkedIn: linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution (Paper Explained)Tree of Thoughts: Deliberate Problem Solving with Large Language Models (Full Paper Review)Titans: Learning to Memorize at Test Time (Paper Analysis)I created an AI-powered Social Network[ML News] GPT-3 learns to edit | Google Pathways | Make-A-Scene | CLIP meets GamePhysics | DouBlindLearning Rate Grafting: Transferability of Optimizer Tuning (Machine Learning Research Paper Review)Unsupervised Brain Models - How does Deep Learning inform Neuroscience? (w/ Patrick Mineault)Dynamic Inference with Neural Interpreters (w/ author interview)RWKV: Reinventing RNNs for the Transformer Era (Paper Explained)I BUILT A FULLY AUTOMATIC MANSPLAINER[ML News] Stable Diffusion Takes Over! (Open Source AI Art)
Yannic Kilcher |

TiDAR: Think in Diffusion, Talk in Autoregression (Paper Analysis)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER