Uploaded January 2022 | Updated September 2026, 2 weeks ago
#mlnews #convnext #mt3
Your update on what's new in the Machine Learning world!
OUTLINE:
0:00 - Intro
0:15 - ConvNeXt: Return of the Convolutions
2:50 - Investigating Saliency Cropping Algorithms
9:40 - YourTTS: SOTA zero-shot Text-to-Speech
10:40 - MT3: Multi-Track Music Transcription
11:35 - China regulates addictive algorithms
13:00 - A collection of Deep Learning interview questions & solutions
13:35 - Helpful Things
16:05 - AlphaZero explained blog post
16:45 - Ru-DOLPH: HyperModal Text-to-Image-to-Text model
17:45 - Google AI 2021 Review
References:
ConvNeXt: Return of the Convolutions
arxiv.org/abs/2201.03545
github.com/facebookresearch/ConvNeXt
twitter.com/giffmana/status/1481054929573888005
twitter.com/wightmanr/status/1481150080765739009
twitter.com/tanmingxing/status/1481362887272636417
Investigating Saliency Cropping Algorithms
openaccess.thecvf.com/content/WACV2022/papers/Birhane_Auditing_Saliency_Cropping_Algorithms_WACV_2022_paper.pdf
vinayprabhu.github.io/Saliency_Image_Cropping/paper_html/main.html
vinayprabhu.medium.com/on-the-twitter-cropping-controversy-critique-clarifications-and-comments-7ac66154f687
vinayprabhu.github.io/Saliency_Image_Cropping
YourTTS: SOTA zero-shot Text-to-Speech
github.com/coqui-ai/TTS?utm_source=pocket_mylist
arxiv.org/abs/2112.02418?utm_source=pocket_mylist
coqui.ai/?utm_source=pocket_mylist
coqui.ai/blog/tts/yourtts-zero-shot-text-synthesis-low-resource-languages
MT3: Multi-Track Music Transcription
arxiv.org/abs/2111.03017
github.com/magenta/mt3
huggingface.co/spaces/akhaliq/MT3
reddit.com/r/MachineLearning/comments/rtlx0r/r_mt3_multitask_multitrack_music_transcription
China regulates addictive algorithms
technode.com/2022/01/05/china-issues-new-rules-to-regulate-algorithms-targeting-addiction-monopolies-and-overspending
qz.com/2109618/china-reveals-new-algorithm-rules-to-weaken-platforms-control-of-users
A collection of Deep Learning interview questions & solutions
arxiv.org/abs/2201.00650?utm_source=pocket_mylist
arxiv.org/pdf/2201.00650.pdf
Helpful Things
docs.deepchecks.com/en/stable/index.html
github.com/deepchecks/deepchecks
docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html
dagshub.com
dagshub.com/docs/index.html
dagshub.com/blog/launching-dagshub-2-0
bayesiancomputationbook.com/welcome.html
mlcontests.com
github.com/Yard1/ray-skorch
github.com/skorch-dev/skorch
rumbledb.org/?utm_source=pocket_mylist
github.com/DarshanDeshpande/jax-models
github.com/s3prl/s3prl
AlphaZero explained blog post
joshvarty.github.io/AlphaZero/?utm_source=pocket_mylist
Ru-DOLPH: HyperModal Text-to-Image-to-Text model
github.com/sberbank-ai/ru-dolph
colab.research.google.com/drive/1gmTDA13u709OXiAeXWGm7sPixRhEJCga?usp=sharing
Google AI 2021 Review
ai.googleblog.com/2022/01/google-research-themes-from-2021-and.html
Links:
TabNine Code Completion (Referral): bit.ly/tabnine-yannick
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: discord.gg/4H8xxDF
BitChute: bitchute.com/channel/yannic-kilcher
LinkedIn: linkedin.com/in/ykilcher
BiliBili: space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
#mlnews #convnext #mt3
Your update on what's new in the Machine Learning world!
OUTLINE:
0:00 - Intro
0:15 - ConvNeXt: Return of the Convolutions
2:50 - Investigating Saliency Cropping Algorithms
9:40 - YourTTS: SOTA zero-shot Text-to-Speech
10:40 - MT3: Multi-Track Music Transcription
11:35 - China regulates addictive algorithms
13:00 - A collection of Deep Learning interview questions & solutions
13:35 - Helpful Things
16:05 - AlphaZero explained blog post
16:45 - Ru-DOLPH: HyperModal Text-to-Image-to-Text model
17:45 - Google AI 2021 Review
References:
ConvNeXt: Return of the Convolutions
arxiv.org/abs/2201.03545
github.com/facebookresearch/ConvNeXt
twitter.com/giffmana/status/1481054929573888005
twitter.com/wightmanr/status/1481150080765739009
twitter.com/tanmingxing/status/1481362887272636417
Investigating Saliency Cropping Algorithms
openaccess.thecvf.com/content/WACV2022/papers/Birhane_Auditing_Saliency_Cropping_Algorithms_WACV_2022_paper.pdf
vinayprabhu.github.io/Saliency_Image_Cropping/paper_html/main.html
vinayprabhu.medium.com/on-the-twitter-cropping-controversy-critique-clarifications-and-comments-7ac66154f687
vinayprabhu.github.io/Saliency_Image_Cropping
YourTTS: SOTA zero-shot Text-to-Speech
github.com/coqui-ai/TTS?utm_source=pocket_mylist
arxiv.org/abs/2112.02418?utm_source=pocket_mylist
coqui.ai/?utm_source=pocket_mylist
coqui.ai/blog/tts/yourtts-zero-shot-text-synthesis-low-resource-languages
MT3: Multi-Track Music Transcription
arxiv.org/abs/2111.03017
github.com/magenta/mt3
huggingface.co/spaces/akhaliq/MT3
reddit.com/r/MachineLearning/comments/rtlx0r/r_mt3_multitask_multitrack_music_transcription
China regulates addictive algorithms
technode.com/2022/01/05/china-issues-new-rules-to-regulate-algorithms-targeting-addiction-monopolies-and-overspending
qz.com/2109618/china-reveals-new-algorithm-rules-to-weaken-platforms-control-of-users
A collection of Deep Learning interview questions & solutions
arxiv.org/abs/2201.00650?utm_source=pocket_mylist
arxiv.org/pdf/2201.00650.pdf
Helpful Things
docs.deepchecks.com/en/stable/index.html
github.com/deepchecks/deepchecks
docs.deepchecks.com/en/stable/examples/guides/quickstart_in_5_minutes.html
dagshub.com
dagshub.com/docs/index.html
dagshub.com/blog/launching-dagshub-2-0
bayesiancomputationbook.com/welcome.html
mlcontests.com
github.com/Yard1/ray-skorch
github.com/skorch-dev/skorch
rumbledb.org/?utm_source=pocket_mylist
github.com/DarshanDeshpande/jax-models
github.com/s3prl/s3prl
AlphaZero explained blog post
joshvarty.github.io/AlphaZero/?utm_source=pocket_mylist
Ru-DOLPH: HyperModal Text-to-Image-to-Text model
github.com/sberbank-ai/ru-dolph
colab.research.google.com/drive/1gmTDA13u709OXiAeXWGm7sPixRhEJCga?usp=sharing
Google AI 2021 Review
ai.googleblog.com/2022/01/google-research-themes-from-2021-and.html
Links:
TabNine Code Completion (Referral): bit.ly/tabnine-yannick
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: discord.gg/4H8xxDF
BitChute: bitchute.com/channel/yannic-kilcher
LinkedIn: linkedin.com/in/ykilcher
BiliBili: space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

![[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)
Paper: https://arxiv.org/abs/2508.21038
Abstract:
Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following, coding, and more. These new benchmarks push embeddings to work for any query and any notion of relevance that could be given. While prior works have pointed out theoretical limitations of vector embeddings, there is a common assumption that these difficulties are exclusively due to unrealistic queries, and those that are not can be overcome with better training data and larger models. In this work, we demonstrate that we may encounter these theoretical limitations in realistic settings with extremely simple queries. We connect known results in learning theory, showing that the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding. We empirically show that this holds true even if we restrict to k=2, and directly optimize on the test set with free parameterized embeddings. We then create a realistic dataset called LIMIT that stress tests models based on these theoretical results, and observe that even state-of-the-art models fail on this dataset despite the simple nature of the task. Our work shows the limits of embedding models under the existing single vector paradigm and calls for future research to develop methods that can resolve this fundamental limitation.
Authors: Orion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk Lee
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n [Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)](https://i.ytimg.com/vi/zKohTkN0Fyk/mqdefault.jpg)



