Uploaded August 2023 | Updated September 2026, 4 days ago
Speaker: Mayee Chen, PhD Student, Stanford University,
Bio: Mayee Chen is a PhD student in the Computer Science department at Stanford University advised by Professor Christopher Ré. She is interested in understanding and improving how models learn from data. Recently, she has focused on problems in data selection, data labeling, and data representations, especially in the setting where there exist multiple input signals or objectives. Previously, she graduated summa cum laude from Princeton University with a concentration in Operations Research and Financial Engineering (ORFE) and a certificate in Applications of Computing, where she worked with Professors Elad Hazan and Miklos Racz.
Abstract: The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good downstream model performance across tasks. We develop a new framework based on a simple hypothesis: just as humans acquire interdependent skills in a deliberate order, language models also follow a natural order when learning a set of skills from their training data. If such an order exists, it can be utilized for improved understanding of LMs and for data-efficient training. Using this intuition, our framework formalizes the notion of a skill and of an ordered set of skills in terms of the associated data. First, using both synthetic and real data, we demonstrate that these ordered skill sets exist, and that their existence enables more advanced skills to be learned with less data when we train on their prerequisite skills. Second, using our proposed framework, we introduce an online data sampling algorithm, Skill-It, over mixtures of skills for both continual pre-training and fine-tuning regimes, where the objective is to efficiently learn multiple skills in the former and an individual skill in the latter. On the LEGO synthetic in the continual pre-training setting, Skill-It obtains 36.5 points higher accuracy than random sampling. On the Natural Instructions dataset in the fine-tuning setting, Skill-It reduces the validation loss on the target skill by 13.6% versus training on data associated with the target skill itself. We apply our skills framework on the recent RedPajama dataset to continually pre-train a 3B-parameter LM, achieving higher accuracy on the LM Evaluation Harness with 1B tokens than the baseline approach of sampling uniformly over data sources with 3B tokens.
Speaker: Mayee Chen, PhD Student, Stanford University,
Bio: Mayee Chen is a PhD student in the Computer Science department at Stanford University advised by Professor Christopher Ré. She is interested in understanding and improving how models learn from data. Recently, she has focused on problems in data selection, data labeling, and data representations, especially in the setting where there exist multiple input signals or objectives. Previously, she graduated summa cum laude from Princeton University with a concentration in Operations Research and Financial Engineering (ORFE) and a certificate in Applications of Computing, where she worked with Professors Elad Hazan and Miklos Racz.
Abstract: The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good downstream model performance across tasks. We develop a new framework based on a simple hypothesis: just as humans acquire interdependent skills in a deliberate order, language models also follow a natural order when learning a set of skills from their training data. If such an order exists, it can be utilized for improved understanding of LMs and for data-efficient training. Using this intuition, our framework formalizes the notion of a skill and of an ordered set of skills in terms of the associated data. First, using both synthetic and real data, we demonstrate that these ordered skill sets exist, and that their existence enables more advanced skills to be learned with less data when we train on their prerequisite skills. Second, using our proposed framework, we introduce an online data sampling algorithm, Skill-It, over mixtures of skills for both continual pre-training and fine-tuning regimes, where the objective is to efficiently learn multiple skills in the former and an individual skill in the latter. On the LEGO synthetic in the continual pre-training setting, Skill-It obtains 36.5 points higher accuracy than random sampling. On the Natural Instructions dataset in the fine-tuning setting, Skill-It reduces the validation loss on the target skill by 13.6% versus training on data associated with the target skill itself. We apply our skills framework on the recent RedPajama dataset to continually pre-train a 3B-parameter LM, achieving higher accuracy on the LM Evaluation Harness with 1B tokens than the baseline approach of sampling uniformly over data sources with 3B tokens.






![Language AI for RNA Virus and RNA Vaccine
Abstract:
Linguistics and biology are two sides of the same coin. This talk features several highly unexpected connections between them which yield efficient algorithms with substantial biological impacts. One such connection (Nature, 2023) is between messenger RNA (mRNA) vaccines and formal language theory. Although widely used in COVID, these vaccines still suffer from instability. But how to design more stable and efficient mRNAs? Here we show a surprising reduction of the mRNA design problem to the classical (1961) concept of “lattice parsing” in speech recognition, which enables efficient search in the exponentially large design space. Experiments on COVID and another virus show that our designs dramatically improves mRNA half-life, protein expression, and in vivo antibody response, compared to the standard method used by Pfizer and Moderna. Another connection (PNAS, 2021) is between COVID variants and multilingual parsing. Here we show that aligning and folding various coronavirus genomes (in order to find conserved structures for drug design) can be viewed as “synchronous parsing” for multiple languages. This enables efficient global prediction of COVID genome structure that matches experimental work.
[1] Nature paper: https://www.nature.com/articles/s41586-023-06127-z
[2] Nature news: https://www.nature.com/articles/d41586-023-01487-y (‘Remarkable’ AI tool designs mRNA vaccines that are more potent and stable)
[3] PNAS paper: https://www.pnas.org/doi/10.1073/pnas.2116269118
Bio:
Liang Huang (PhD, Penn, 2008) is a Professor of Computer Science at Oregon State University, and co-founder of Coderna.ai. Until recently, he was also a Distinguished Scientist at Baidu Research USA. He also worked at Google Research, USC, and City Univ. of New York. He was known for algorithms and theory in computational linguistics, where he received several best paper awards (ACL 2008 Best Paper Award, EMNLP 2016 Best Paper Honorable Mentions, NAACL 2022 Best Demo Paper Award) and delivered keynotes at ACL 2019 and CVPR 2021. But in recent years, he has shifted his attention to applying these natural language algorithms to computational biology, esp. RNA folding and RNA design, with the hope of fighting COVID. This line of linguistics-inspired biology work eventually led to PNAS (2021) and Nature (2023) papers, and is widely covered in the media. Language AI for RNA Virus and RNA Vaccine](https://i.ytimg.com/vi/B-fiTnUkq2A/mqdefault.jpg)



