Uploaded November 2023 | Updated September 2026, 1 day ago
Abstract: Cross-lingual semantic parsing transfers parsing capability from a high-resource language (e.g., English) to low-resource languages with scarce training data. Previous work has primarily considered silver-standard data augmentation or zero-shot methods, however, exploiting few-shot gold data is comparatively unexplored. We propose a new approach to cross-lingual semantic parsing by explicitly minimizing cross-lingual divergence between probabilistic latent variables using Optimal Transport. We demonstrate how this direct guidance improves parsing from natural languages using fewer examples and less training. We evaluate our method on two datasets, MTOP and MultiATIS++SQL, establishing state-of-the-art results under a few-shot cross-lingual regime. Ablation studies further reveal that our method improves performance even without parallel input translations. In addition, we show that our model better captures cross-lingual structure in the latent space to improve semantic representation similarity. Given the improvement in semantic structure, we also consider new areas for improvement in cross-lingual structured alignment and how to approach future challenges in this domain. arxiv.org/abs/2307.04096
Bio: I am a final year PhD student at Edinburgh advised by Mirella Lapata on data-efficient cross-lingual semantic parsing and structure prediction. My research focuses on applying mathematical modeling to overcome data and resource constraints in adapting neural models to languages beyond English. My recent work focuses on optimization strategies for this goal including flatness-seeking optimization robustness and compute-efficient meta-learning. I was also recently awarded Outstanding Paper at ACL 2023 for collaborative work on MT evaluation. I'm excited by methods to improve data, resource and parameter efficiency for both cross-language adaptation and wider modeling improvement. I have previously interned at the Allen Institute for AI and Apple Siri. Prior to my PhD, I completed Masters degrees at University of Edinburgh, University of Cambridge and University College London.
tomsherborne.github.io
Abstract: Cross-lingual semantic parsing transfers parsing capability from a high-resource language (e.g., English) to low-resource languages with scarce training data. Previous work has primarily considered silver-standard data augmentation or zero-shot methods, however, exploiting few-shot gold data is comparatively unexplored. We propose a new approach to cross-lingual semantic parsing by explicitly minimizing cross-lingual divergence between probabilistic latent variables using Optimal Transport. We demonstrate how this direct guidance improves parsing from natural languages using fewer examples and less training. We evaluate our method on two datasets, MTOP and MultiATIS++SQL, establishing state-of-the-art results under a few-shot cross-lingual regime. Ablation studies further reveal that our method improves performance even without parallel input translations. In addition, we show that our model better captures cross-lingual structure in the latent space to improve semantic representation similarity. Given the improvement in semantic structure, we also consider new areas for improvement in cross-lingual structured alignment and how to approach future challenges in this domain. arxiv.org/abs/2307.04096
Bio: I am a final year PhD student at Edinburgh advised by Mirella Lapata on data-efficient cross-lingual semantic parsing and structure prediction. My research focuses on applying mathematical modeling to overcome data and resource constraints in adapting neural models to languages beyond English. My recent work focuses on optimization strategies for this goal including flatness-seeking optimization robustness and compute-efficient meta-learning. I was also recently awarded Outstanding Paper at ACL 2023 for collaborative work on MT evaluation. I'm excited by methods to improve data, resource and parameter efficiency for both cross-language adaptation and wider modeling improvement. I have previously interned at the Allen Institute for AI and Apple Siri. Prior to my PhD, I completed Masters degrees at University of Edinburgh, University of Cambridge and University College London.
tomsherborne.github.io



![From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2
AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy!, but the rich variety of standardized exams has remained a landmark challenge. Even as recently as 2016, the best AI system could achieve merely 59.3% on an 8th Grade science exam.
This talk reports success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exams non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern Natural Language Processing (NLP) methods can result in mastery on this task. While not a full solution to general question-answering (the questions are limited to 8th Grade multiple-choice science) it represents a significant milestone for the field. [ Paper at AI Magazine 41 (4), Winter 2020, https://arxiv.org/pdf/1909.01958.pdf ] From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2](https://i.ytimg.com/vi/CR3aICkhCJM/mqdefault.jpg)






