Deploy Hugging Face models on Google Cloud: from the hub to Vertex AI @juliensimonfr
Deploy Hugging Face models on Google Cloud: from the hub to Vertex AI  @juliensimonfr
Uploaded April 2024 | Updated September 2026, 2 weeks ago
In this series of three videos, I walk you through the deployment of Hugging Face models on Google Cloud, in three different ways:

- Deployment from the hub model page to Inference endpoints (youtu.be/mlU-2QYx4a0), with the Google Gemma 7B model,
- Deployment from the hub model page to Vertex AI (this video), with the Microsoft Phi-2 2.7B model,
- Deployment directly from within Vertex AI (youtu.be/PFHzfzyY2iY), with the TinyLlama 1.1B model.

Get started at huggingface.co :)

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Deploy Hugging Face models on Google Cloud: from the hub to Vertex AI🚀 Unveiling the Future: Meet Arcee Nova, the Ultimate LLM! 🌟Open Source AI with Hugging Face - Dallas AI  meetup (05/2024)Decoder-only inference: a step-by-step deep diveDeep dive: model merging (part 1)Migrating from OpenAI models to Hugging Face modelsIs Waiting for AI to Be Safe a Smart Strategy or a Missed Opportunity?Deploying SuperNova-Lite on Inferentia2: the best 8B model for $1 an hour!Parameter-efficient fine-tuning with QLoRA and Hugging FaceAWS User Group DubaiInterview BFM Business - Hugging Face (04/2023)What is Vibe Coding? Build based on intent!
Julien Simon |

Deploy Hugging Face models on Google Cloud: from the hub to Vertex AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER