simon alexandersonThis video presents a system that generates both speech and gestures together from TEXT. Note that both modalities were trained on the same speaker, so the synthetic speech and gestures are coherent. The work was presented at IVA 2020.
Authors: Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko and Jonas Beskow.
Generating coherent speech and gesture from textsimon alexanderson2020-11-02 | This video presents a system that generates both speech and gestures together from TEXT. Note that both modalities were trained on the same speaker, so the synthetic speech and gestures are coherent. The work was presented at IVA 2020.
Authors: Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko and Jonas Beskow.
For more information, please see our project page: simonalexanderson.github.io/IVA2020[SIGGRAPH 2023] Listen, denoise, action! Audio-driven motion synthesis with diffusion modelssimon alexanderson2023-05-16 | This video presents our SIGGRAPH 2023 paper on audio-driven motion synthesis using diffusion models. Given audio and (optionally) a desired style, our models generate dancing or full-body gesticulation with top-of-the-line motion quality. The style expression can be made more or less pronounced, and we also demonstrate a new way to blend and transition between styles. The latter innovation uses a new product-of-experts setup, where an ensemble of diffusion models together guide the output in each denoising step.
For more information and links to our paper, code and data, please see our project page at: https://www.speech.kth.se/research/listen-denoise-action/
To try out the system in action, please go to: motorica.ai
Authors: Simon Alexanderson (1,2) Rajmund Nagy (1) Jonas Beskow (1) Gustav Eje Henter (1,2)
(1) KTH Royal Institute of Technology, Stockholm, Sweden (2) Motorica.aiMoGlow: Probabilistic and controllable motion synthesis using normalising flowssimon alexanderson2020-12-03 | This video introduces our SIGGRAPH Asia 2020 paper on animating motion using so-called normalising flows. In particular, we describe a new, deep machine-learning architecture, called MoGlow, for data-driven animation, that: 1) is general and does not make any task-specific assumptions (such as the motion being cyclic) 2) can be controlled interactively using high-level, "weak" input signals, and 3) is probabilistic and can describe (and sample) many different behaviours.
The resulting animation looks highly natural regardless of the task (e.g., animating a human or a dog).
Authors: Gustav Eje Henter*, Simon Alexanderson*, Jonas Beskow KTH Royal Institute of Technology Stockholm, Sweden
*) joint first authorsStyle-Controllable Speech-Driven Gesture Synthesis Using Normalizing Flows - Eurographics 2020simon alexanderson2020-05-27 | This is a video-introduction of our Eurographics 2020 paper titled "Style-Controllable Speech-Driven Gesture Synthesis Using Normalizing Flows". The paper received an honourable mention for the Günter Enderle Award at the conference.