Uploaded March 2024 | Updated September 2026, 2 weeks ago
Learn about the key challenges in improving efficiency of LLM serving, and an overview of multiple techniques the team is developing to address this problem. This talk also discusses matformers, a technique to train one model but read-off 100s of smaller models, along with techniques to speed up decoding in LLMs.
Watch more Research@ Bangalore → https://goo.gle/3VSup7d
Google Research Channel → https://goo.gle/GoogleResearch
#GoogleResearch
Learn about the key challenges in improving efficiency of LLM serving, and an overview of multiple techniques the team is developing to address this problem. This talk also discusses matformers, a technique to train one model but read-off 100s of smaller models, along with techniques to speed up decoding in LLMs.
Watch more Research@ Bangalore → https://goo.gle/3VSup7d
Google Research Channel → https://goo.gle/GoogleResearch
#GoogleResearch










