Uploaded February 2023 | Updated September 2026, 50 minutes ago
Understanding and Improving Compositional Generalization
Ben Bogin
Pre-trained language models perform well on a large variety of question-answering tasks, but they still often fail in the compositional generalization setup, where models are tested on unseen compositions of reasoning skills.
In this talk, I will go over three research directions that address this challenge. In the first part, I will show how we can improve generalization and interpretability with a compositional model architecture that recursively computes outputs and representations over grounded latent trees. In the second part, I will describe our finding of a key factor that makes sequence to sequence models fail to output unseen structures, and I will show how we can leverage this insight to improve generalization. Finally, I will show how large language models, such as GPT-3, cope with this challenge, showing that at least for now, scaling the number of parameters often improves generalization, but is still far from solving it.
Understanding and Improving Compositional Generalization
Ben Bogin
Pre-trained language models perform well on a large variety of question-answering tasks, but they still often fail in the compositional generalization setup, where models are tested on unseen compositions of reasoning skills.
In this talk, I will go over three research directions that address this challenge. In the first part, I will show how we can improve generalization and interpretability with a compositional model architecture that recursively computes outputs and representations over grounded latent trees. In the second part, I will describe our finding of a key factor that makes sequence to sequence models fail to output unseen structures, and I will show how we can leverage this insight to improve generalization. Finally, I will show how large language models, such as GPT-3, cope with this challenge, showing that at least for now, scaling the number of parameters often improves generalization, but is still far from solving it.










