Uploaded January 2025 | Updated September 2026, 1 day ago
Abstract:
The intersection of language and vision is a fundamental aspect of human cognition. Our ability to interpret visual stimuli is often guided by linguistic constructs, enabling us to describe, understand, and interact with the world around us. This intrinsic connection suggests that language can serve as a powerful prior for creating systems with advanced visual intelligence.
Despite the broad success of existing systems, their applicability to critical domains remains limited, which we hypothesize is due to a lack of appropriate priors. Without the guidance of effective priors, models may
overly rely on in-domain data to form solutions, leading to potential catastrophic failures when deployed in different scenarios.
In this talk, I will introduce two ways of applying language priors to vision systems to enhance their interpretability, robustness, and data efficiency. First, I will show how language models can be used to construct high-performance concept bottlenecks for interpretable image classifiers, which exhibit significantly greater robustness to domain shifts in medical applications. Second, to address data scarcity in embodied AI and vision-language models, I will demonstrate how to use language-guided synthetic data to develop more generalizable embodied agents and multimodal models.
Bio:
Yue Yang is a final-year PhD candidate at the University of Pennsylvania, advised by Chris Callison-Burch and Mark Yatskar. His research interests lie in the intersection of Natural Language Processing and Computer Vision. His recent works aim to leverage the knowledge priors from large language models to improve the interpretability, robustness, and data efficiency of vision systems.
Personal Website: yueyang1996.github.io
Abstract:
The intersection of language and vision is a fundamental aspect of human cognition. Our ability to interpret visual stimuli is often guided by linguistic constructs, enabling us to describe, understand, and interact with the world around us. This intrinsic connection suggests that language can serve as a powerful prior for creating systems with advanced visual intelligence.
Despite the broad success of existing systems, their applicability to critical domains remains limited, which we hypothesize is due to a lack of appropriate priors. Without the guidance of effective priors, models may
overly rely on in-domain data to form solutions, leading to potential catastrophic failures when deployed in different scenarios.
In this talk, I will introduce two ways of applying language priors to vision systems to enhance their interpretability, robustness, and data efficiency. First, I will show how language models can be used to construct high-performance concept bottlenecks for interpretable image classifiers, which exhibit significantly greater robustness to domain shifts in medical applications. Second, to address data scarcity in embodied AI and vision-language models, I will demonstrate how to use language-guided synthetic data to develop more generalizable embodied agents and multimodal models.
Bio:
Yue Yang is a final-year PhD candidate at the University of Pennsylvania, advised by Chris Callison-Burch and Mark Yatskar. His research interests lie in the intersection of Natural Language Processing and Computer Vision. His recent works aim to leverage the knowledge priors from large language models to improve the interpretability, robustness, and data efficiency of vision systems.
Personal Website: yueyang1996.github.io






![Open AI: considering the ethical upsides and downsides of Open AI development
Abstract:
In this talk, I will discuss the ethical upsides and downsides of releasing AI openly.
I will first present our FAccT’22 paper [1], where we interview contributors to an open source Deepfake tool about their sense of responsibility and agency to prevent harm. We show that open source licenses and norms combine with notions of technological inevitability and neutrality to lead contributors to disavow responsibility for harmful ways their tool is used.
I will then broaden to discuss other work examining AI openness, situated in the context of “Open”AI’s U-turn on openness. I will discuss benefits of AI openness, such as supporting open science, and enabling wider scrutiny for harms such as bias, and downsides, such as enabling the proliferation of powerful tools which can be used to harm.
I will then conclude by enumerating and advocating for a variety of “middle ground” approaches to AI openness, including methods of norm setting, ethical licenses, release gating, or hard technical restrictions, before opening up discussion for other ways of tackling this thorny problem.
[1] https://dl.acm.org/doi/abs/10.1145/3531146.3533779
Bio:
David Gray Widder (he/him) studies how people creating “Artificial Intelligence” systems think about the downstream harms their systems make possible. He is a Doctoral Student in the School of Computer Science at Carnegie Mellon University, and previously worked at Intel Labs, Microsoft Research, and NASA’s Jet Propulsion Laboratory. He was born in Tillamook, Oregon, and raised in Berlin and Singapore. He maintains a conceptual-realist artistic practice, advocates against police terror and pervasive surveillance, and enjoys distance running.
https://davidwidder.me/ Open AI: considering the ethical upsides and downsides of Open AI development](https://i.ytimg.com/vi/HZP3kps9TsU/mqdefault.jpg)



