Uploaded November 2025 | Updated September 2026, 2 hours ago
Synthetic data might be the most misunderstood topic in AI right now. Is it a cheat code for training better models or a trap that slowly collapses model diversity? Here's what Letitia, one of the sharpest minds in VLMs and fresh from her PhD, the answer is way more interesting than a simple yes or no.
Small models? Synthetic data can be a superpower because you get clean, structured, high quality examples without the chaos of the real web. Her team even trained 1B and 8B models on only synthetic German text and outperformed the same models trained purely on organic data. But for large models, diversity becomes king and the messiness of real world text is irreplaceable.
Where things really get interesting is vision language models. Most image captions online are basically “a dog on a bench,” which teaches models nothing about the visual richness in the image. For VLMs, synthetic captions aren’t just helpful; they’re necessary. You can inject detail, attributes, context, relationships, and give the model a reason to look deeper. And that changes everything.
#syntheticdata #visionlanguagemodels #llm
Synthetic data might be the most misunderstood topic in AI right now. Is it a cheat code for training better models or a trap that slowly collapses model diversity? Here's what Letitia, one of the sharpest minds in VLMs and fresh from her PhD, the answer is way more interesting than a simple yes or no.
Small models? Synthetic data can be a superpower because you get clean, structured, high quality examples without the chaos of the real web. Her team even trained 1B and 8B models on only synthetic German text and outperformed the same models trained purely on organic data. But for large models, diversity becomes king and the messiness of real world text is irreplaceable.
Where things really get interesting is vision language models. Most image captions online are basically “a dog on a bench,” which teaches models nothing about the visual richness in the image. For VLMs, synthetic captions aren’t just helpful; they’re necessary. You can inject detail, attributes, context, relationships, and give the model a reason to look deeper. And that changes everything.
#syntheticdata #visionlanguagemodels #llm










