Session 1: LLM Scaling and the Role of Synthetic Data @ColumbiaSEAS
Session 1: LLM Scaling and the Role of Synthetic Data  @ColumbiaSEAS
Uploaded June 2025 | Updated September 2026, 3 weeks ago
By Tatsunori Hashimoto, Stanford University: Scaling up language models has been a key driver of the recent, dramatic improvements in their capabilities. Despite the significant empirical successes of scaling up pre-training in the past 5 years, the future of this approach has become uncertain: large base models no longer show the same types of jumps in benchmark performance, and new forms of scaling (’test-time scaling’) have been proposed to take its place. Does the data inefficiency of pretraining pose fundamental challenges to scaling? Will scaling up inference compute suffice for future capability gains? This talk will cover a few initial investigations into these questions, in the hopes of better understanding whether and how LLM scaling will continue.
Session 1: LLM Scaling and the Role of Synthetic DataSession 3: A New Paradigm for Learning Distribution ShiftSymposium on AI & Sports | Lightning Talks, Part 1SESSION 5: Columbia CryptoEconomics Workshop 2024Senior Design Expo 2025CCAE 2025: Practical & Theoretical Trends in Research, Development & Design of Electrical DrivesWhy Fusion Was Always 30 Years AwaySports AI Innovation Symposium | Opening KeynoteCCE Day 2: Session 7 - Lightning TalksSESSION 4: Columbia CryptoEconomics Workshop 2024SESSION 1C: Columbia CryptoEconomics Working SessionColumbia Space Initiative Reaches New Heights
Columbia Engineering |

Session 1: LLM Scaling and the Role of Synthetic Data

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER