Uploaded September 2024 | Updated September 2026, 3 days ago
Abstract: As large pretrained models underlying generative AI systems have grown larger, inscrutable, and widely-deployed, interest in understanding their nature as emergent rather than engineered systems has grown. I believe to AI science forward, we need to focus on building rigorous observational tools for these systems, which can characterize capabilities unambiguously. At their best, benchmarks and metrics could meet this need, but at present they are often treated as mere leaderboards to chase and only very indirectly measure capabilities of interest. This talk first presents the case for building a dedicated discipline of "model metrology" that treats building evaluations as first-class technical contributions. Then, it covers ongoing research directions in metrology for multilingual and multicultural competence in generative language and image systems. This evaluation work unveils interesting model capabilities, unexpected ethical challenges, new technical problems, and exciting new directions for model development, all by making the previously unmeasureable, measurable.
Bio: Michael Saxon is a 5th-year PhD candidate and NSF Fellow in the NLP Group at the University of California, Santa Barbara. His research sits at intersections within benchmarking, multilinguality/multiculturality, AI ethics, and multimodal systems. His research follows from the idea that ethical issues in generative AI are both important to address on their own and motivate interesting new technical challenges. (https://saxon.me)
Abstract: As large pretrained models underlying generative AI systems have grown larger, inscrutable, and widely-deployed, interest in understanding their nature as emergent rather than engineered systems has grown. I believe to AI science forward, we need to focus on building rigorous observational tools for these systems, which can characterize capabilities unambiguously. At their best, benchmarks and metrics could meet this need, but at present they are often treated as mere leaderboards to chase and only very indirectly measure capabilities of interest. This talk first presents the case for building a dedicated discipline of "model metrology" that treats building evaluations as first-class technical contributions. Then, it covers ongoing research directions in metrology for multilingual and multicultural competence in generative language and image systems. This evaluation work unveils interesting model capabilities, unexpected ethical challenges, new technical problems, and exciting new directions for model development, all by making the previously unmeasureable, measurable.
Bio: Michael Saxon is a 5th-year PhD candidate and NSF Fellow in the NLP Group at the University of California, Santa Barbara. His research sits at intersections within benchmarking, multilinguality/multiculturality, AI ethics, and multimodal systems. His research follows from the idea that ethical issues in generative AI are both important to address on their own and motivate interesting new technical challenges. (https://saxon.me)










