Accelerate Transformer inference with AWS Inferentia 2 @juliensimonfr
Accelerate Transformer inference with AWS Inferentia 2  @juliensimonfr
Uploaded April 2023 | Updated September 2026, 2 weeks ago
In this video, I show you how to accelerate Transformer inference with AWS Inferentia 2, a custom chip designed by AWS.

Starting from a BERT model that I fine-tuned on AWS Trainium (youtu.be/HweP7OYNiIA) , I compile it with the Neuron SDK for Inferentia 1. Then, using an inf2.xlarge instance (1 Inferentia2 chips, 2 Neuron Cores), I show you how to get to 1.3 ms latency at 1,700 inferences per second.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos ⭐️⭐️⭐️

- Amazon EC2 Inf2: aws.amazon.com/ec2/instance-types/inf2
- Deep Learning AMI Neuron PyTorch 1.13.0 (Ubuntu 20.04) 20230405
ami-02f8a3b8fe70e81b9 (64-bit (x86))
- AWS Neuron SDK documentation:
* awsdocs-neuron.readthedocs-hosted.com/en/latest/frameworks/torch/inference.html
* awsdocs-neuron.readthedocs-hosted.com/en/latest/general/arch/neuron-hardware/neuron-core-v2.html#neuroncores-v2-arch
- Code: gitlab.com/juliensimon/huggingface-demos/-/tree/main/inferentia2
Accelerate Transformer inference with AWS Inferentia 2Arcee Orchestra - Build an Agentic Workflow to Augment and Localize YouTube ContentIntroducing the Arcee Model EngineCut your GPT 5 costs by 100xDeep Dive: Teaching Arcee Trinity Mini to Read Medical Research with RLVR and GRPODeep Dive: Advanced distributed training with Hugging Face LLMs and AWS Trainium
Julien Simon |

Accelerate Transformer inference with AWS Inferentia 2

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER