Forcefully controlling AI’s ideology doesnt really work... (yet?) @bycloudAI
Forcefully controlling AI’s ideology doesnt really work... (yet?)  @bycloudAI
Uploaded November 2024 | Updated September 2026, 1 week ago
Happy FlexiSpot Black Friday Sale now, Up to 65% OFF! You also have the chance to win free orders during this period. Use my code ''YTE7P50 '' to get EXTRA $50 off on the E7 Plus standing desk.

FlexiSpot E7 Plus standing desk: USA: bit.ly/3OiLQJ5
CAN: bit.ly/3ZiPE3a
FlexiSpot Black Friday: bit.ly/410B95i


In this video, we will explore the latest research blog Anthropic AI has published discussing about feature steering regarding to mitigate social biases. This is their first step of applying their mechanistic interpretability research to alignment.



Check out my patreon for the research maps!
patreon.com/c/bycloud
for existing YouTube members, please DM me on discord


my research paper newsletter
mail.bycloud.ai

Evaluating feature steering: A case study in mitigating social biases
[Blog] anthropic.com/research/evaluating-feature-steering


This video is supported by the kind Patrons & YouTube Members:
🙏Andrew Lescelius, Ben Shaener, Chris LeDoux, Miguilim, Deagan, FiFaŁ, Robert Zawiasa, Marcelo Ferreira, Owen Ingraham, Daddy Wen, Tony Jimenez, Panther Modern, Jake Disco, Demilson Quintao, Penumbraa, Shuhong Chen, Hongbo Men, happi nyuu nyaa, Carol Lo, Mose Sakashita, Miguel, Bandera, Gennaro Schiano, gunwoo, Ravid Freedman, Mert Seftali, Mrityunjay, Richárd Nagyfi, Timo Steiner, Henrik G Sundt, projectAnthony, Brigham Hall, Kyle Hudson, Kalila, Jef Come, Jvari Williams, Tien Tien, BIll Mangrum, owned, Janne Kytölä, SO, Richárd Nagyfi, Hector, Drexon, Claxvii 177th, Inferencer, Michael Brenner, Akkusativ, Oleg Wock, FantomBloth, Thipok Tham, Clayton Ford, Theo, Handenon, Diego Silva, mayssam, Kadhai Pesalam, Tim Schulz, jiye, Anushka, Henrik Sundt

[Discord] discord.gg/NhJZGtH
[Twitter] twitter.com/bycloudai
[Patreon] patreon.com/bycloud

[Profile & Banner Art] twitter.com/pygm7
[Video Editor] @Booga04
Forcefully controlling AI’s ideology doesnt really work... (yet?)OpenAIs Sora Got Some New Results... And Theyre Way Too GoodMetas Reasoning AI Model... Reasons Without Using Words?The Death of RAG? Recursive LM ExplainedControlNet Revolutionized How We Use AI To Generate ImagesAI that can read 150,000 words at onceSo Gaslighting AI Is A Thing Now [ChatGPT]They Found a Way to Steal Frontier LLM’s ReasoningApple Joins The Extended Reality Development [The AI Timeline #4]Anti-AI Art Bill Backfires? How C2PA Really WorksResearchers Are Getting Really Creative Training LLMs [Token Order Prediction]How Distributed Training Will Revive Open Source AI
bycloud |

Forcefully controlling AI’s ideology doesn't really work... (yet?)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER