Adapting to Long Tail Domains: A Case Study in Clinical Information | AI2 @allenai
Adapting to Long Tail Domains: A Case Study in Clinical Information | AI2  @allenai
Uploaded November 2021 | Updated September 2026, 10 hours ago
Adapting to Long Tail Domains: A Case Study in Clinical Information
Aakanksha Naik

Advances in deep learning, especially self-supervised representation learning, have produced models that reach human parity on many benchmark datasets, which cover a variety of natural language understanding tasks. However, benchmark datasets are constructed from naturally occurring texts, and are no exception to Zipf's law, containing a small proportion of highly frequent phenomena and a long tail of less frequent phenomena. Benchmark-driven evaluation and model development favors NLU models that perform well on the head, sidelining domains and phenomena that are underrepresented.

Transfer learning research has improved the generalization abilities of models trained on NLU benchmarks, but an important issue plaguing long tail domains is data scarcity (especially of labeled data). In this talk, I will focus on the question: how far can transfer learning improve performance on long tail domains in low-data regimes? In the first part of the talk, I will discuss likelihood-based instance weighting, an unsupervised adaptation technique and analyze its effectiveness in adapting event extractors to clinical text. In the second part of the talk, I will present a systematic study of the effectiveness of various popular adaptation methods on a range of clinical entity and event extraction tasks in an unsupervised adaptation setting. Through these studies, I will demonstrate that we have an arsenal of adaptation techniques that can be applied to long tail domains, but still some way to go in developing a systematic understanding of their strengths, weaknesses and effectiveness.
Adapting to Long Tail Domains: A Case Study in Clinical Information | AI2Concept Bottleneck Models for Text ClassificationComputing with a Mess | Embodied AI Lecture Series at AI2Imaginative Vision Language ModelsIntroducing FlexOlmo: A New Paradigm for Language Model Training and Data CollaborationGenerative AI & CopyrightAi2 at NVIDIA GTC 2026Mitigating Knowledge Collapse through Epistemic DiversityolmOCR: an open-source tool to extract clean plain text from PDFs!Molmo 2 | Counting objects and actionsBLADE: Benchmarking Language Model Agents for Data-Driven ScienceJust-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTE
Ai2 |

Adapting to Long Tail Domains: A Case Study in Clinical Information | AI2

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER