Side Effects May Include Homogenization and Overusing Cliches @allenai
Side Effects May Include Homogenization and Overusing Cliches  @allenai
Uploaded January 2025 | Updated September 2026, 1 day ago
Abstract:
In recent years, the evolution of large language models (LLMs) from fine-tuned, single-purpose tools to more versatile, suggestion models has led to a surge in collaborative writing with model assistance. However, as LLMs are trained to be more general-purpose, side effects from this training emerge. In this talk, I will cover two recent projects where we uncover concerns when different profiles of users interact with LLMs for writing tasks. The first involves more amateur, everyday users who might use LLMs to spruce up their writing. Through a controlled user study where users write with and without model help, we develop a set of diversity metrics and find that, while they might feel that their writing output is of higher quality, this might
come at the cost of producing more homogeneous content. Second, we conduct a qualitative user study to examine the use of LLMs in the workflow of emerging professional writers. We find that authors find model content frustrating, often too literal and cliche to be used in their writing, a side effect of alignment tuning to produce more “safe” output. Finally, I’ll wrap up with initial explorations into mitigating these issues, at training and inference time.

Bio:
Vishakh Padmakumar is a PhD student at the Center for Data Science, New York University (NYU). He is currently working on problems in text generation and human-AI collaboration as part of the Machine Learning for Language (ML2) group. Prior to this, he was a Graduate Research Associate at the NYU Center for Social Media and Politics working on political stance classification and multimodal content sharing in online disinformation campaigns. He has previously obtained a Master's degree in Computer Science from NYU and a Bachelor's degree in Information Technology at the National Institute of Technology - Karnataka.
Side Effects May Include Homogenization and Overusing ClichesGooAQ: Open Question Answering with Diverse Answer Types | AI2Language AI for RNA Virus and RNA VaccineDoing for our robots what nature did for us | Embodied AI Lecture series at AI2Optimal Transport Posterior Alignment for Cross-lingual Semantic ParsingFiguring out how the world works: causality in a world full of real peopleAutomated Hypothesis Validation with Agentic Sequential FalsificationsOpenWebMath: An Open Dataset of High-Quality Mathematical Web TextFrom F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2Meet an AI Using ScholarReliability and interactive debugging for language modelsUsing Asta AutoDiscovery: AI-powered autonomous scientific discovery
Ai2 |

Side Effects May Include Homogenization and Overusing Cliches

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER