Uploaded April 2024 | Updated September 2026, 2 weeks ago
In this series of three videos, I walk you through the deployment of Hugging Face models on Google Cloud, in three different ways:
- Deployment from the hub model page to Inference endpoints (youtu.be/mlU-2QYx4a0), with the Google Gemma 7B model,
- Deployment from the hub model page to Vertex AI (youtu.be/cBdLw5BnGrk), with the Microsoft Phi-2 2.7B model,
- Deployment directly from within Vertex AI (this video), with the TinyLlama 1.1B model.
Get started at huggingface.co :)
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
In this series of three videos, I walk you through the deployment of Hugging Face models on Google Cloud, in three different ways:
- Deployment from the hub model page to Inference endpoints (youtu.be/mlU-2QYx4a0), with the Google Gemma 7B model,
- Deployment from the hub model page to Vertex AI (youtu.be/cBdLw5BnGrk), with the Microsoft Phi-2 2.7B model,
- Deployment directly from within Vertex AI (this video), with the TinyLlama 1.1B model.
Get started at huggingface.co :)
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️










