Uploaded August 2022 | Updated September 2026, 1 day ago
Abstract:
The advent of big data promises to revolutionize medicine by making it more personalized and effective, but big data also presents a grand challenge of information overload. For example, tumor sequencing has become routine in cancer treatment, yet interpreting the genomic data requires painstakingly curating knowledge from a vast biomedical literature, which grows by thousands of papers every day. Electronic medical records contain high-definition patient information for speeding up clinical trial recruitment and drug development, but curating such real-world evidence from clinical notes can take hours for a single patient. Natural language processing (NLP) can play a key role in interpreting big data for precision medicine. In particular, machine reading can help unlock knowledge from text by substantially improving curation efficiency. However, standard supervised methods require labeled examples, which are expensive and time-consuming to produce at scale. In this talk, I'll present Project Hanover, where we overcome the annotation bottleneck by combining deep learning with probabilistic logic, by exploiting self-supervision from readily available resources such as ontologies and databases, and by leveraging domain-specific pretraining on unlabeled text. This enables us to extract knowledge from tens of millions of publications, structure real-world data for millions of cancer patients, and apply the extracted knowledge and real-world evidence to support precision oncology.
Bio:
Hoifung Poon is the Senior Director of Biomedical NLP at Microsoft Health Futures and an affiliated professor at the University of Washington Medical School. He leads Project Hanover, with the overarching goal of structuring medical data for precision medicine. He has given tutorials on this topic at top conferences such as the Association for Computational Linguistics (ACL) and the Association for the Advancement of Artificial Intelligence (AAAI). His research spans a wide range of problems in machine learning and natural language processing (NLP), and his prior work has been recognized with Best Paper Awards from premier venues such as the North American Chapter of the Association for Computational Linguistics (NAACL), Empirical Methods in Natural Language Processing (EMNLP), and Uncertainty in AI (UAI). He received his Ph.D. in Computer Science and Engineering from the University of Washington, specializing in machine learning and NLP.
Abstract:
The advent of big data promises to revolutionize medicine by making it more personalized and effective, but big data also presents a grand challenge of information overload. For example, tumor sequencing has become routine in cancer treatment, yet interpreting the genomic data requires painstakingly curating knowledge from a vast biomedical literature, which grows by thousands of papers every day. Electronic medical records contain high-definition patient information for speeding up clinical trial recruitment and drug development, but curating such real-world evidence from clinical notes can take hours for a single patient. Natural language processing (NLP) can play a key role in interpreting big data for precision medicine. In particular, machine reading can help unlock knowledge from text by substantially improving curation efficiency. However, standard supervised methods require labeled examples, which are expensive and time-consuming to produce at scale. In this talk, I'll present Project Hanover, where we overcome the annotation bottleneck by combining deep learning with probabilistic logic, by exploiting self-supervision from readily available resources such as ontologies and databases, and by leveraging domain-specific pretraining on unlabeled text. This enables us to extract knowledge from tens of millions of publications, structure real-world data for millions of cancer patients, and apply the extracted knowledge and real-world evidence to support precision oncology.
Bio:
Hoifung Poon is the Senior Director of Biomedical NLP at Microsoft Health Futures and an affiliated professor at the University of Washington Medical School. He leads Project Hanover, with the overarching goal of structuring medical data for precision medicine. He has given tutorials on this topic at top conferences such as the Association for Computational Linguistics (ACL) and the Association for the Advancement of Artificial Intelligence (AAAI). His research spans a wide range of problems in machine learning and natural language processing (NLP), and his prior work has been recognized with Best Paper Awards from premier venues such as the North American Chapter of the Association for Computational Linguistics (NAACL), Empirical Methods in Natural Language Processing (EMNLP), and Uncertainty in AI (UAI). He received his Ph.D. in Computer Science and Engineering from the University of Washington, specializing in machine learning and NLP.


![Open AI: considering the ethical upsides and downsides of Open AI development
Abstract:
In this talk, I will discuss the ethical upsides and downsides of releasing AI openly.
I will first present our FAccT’22 paper [1], where we interview contributors to an open source Deepfake tool about their sense of responsibility and agency to prevent harm. We show that open source licenses and norms combine with notions of technological inevitability and neutrality to lead contributors to disavow responsibility for harmful ways their tool is used.
I will then broaden to discuss other work examining AI openness, situated in the context of “Open”AI’s U-turn on openness. I will discuss benefits of AI openness, such as supporting open science, and enabling wider scrutiny for harms such as bias, and downsides, such as enabling the proliferation of powerful tools which can be used to harm.
I will then conclude by enumerating and advocating for a variety of “middle ground” approaches to AI openness, including methods of norm setting, ethical licenses, release gating, or hard technical restrictions, before opening up discussion for other ways of tackling this thorny problem.
[1] https://dl.acm.org/doi/abs/10.1145/3531146.3533779
Bio:
David Gray Widder (he/him) studies how people creating “Artificial Intelligence” systems think about the downstream harms their systems make possible. He is a Doctoral Student in the School of Computer Science at Carnegie Mellon University, and previously worked at Intel Labs, Microsoft Research, and NASA’s Jet Propulsion Laboratory. He was born in Tillamook, Oregon, and raised in Berlin and Singapore. He maintains a conceptual-realist artistic practice, advocates against police terror and pervasive surveillance, and enjoys distance running.
https://davidwidder.me/ Open AI: considering the ethical upsides and downsides of Open AI development](https://i.ytimg.com/vi/HZP3kps9TsU/mqdefault.jpg)







