Deploy Hugging Face models on Google Cloud: from the hub to Inference Endpoints @juliensimonfr
Deploy Hugging Face models on Google Cloud: from the hub to Inference Endpoints  @juliensimonfr
Uploaded April 2024 | Updated September 2026, 2 weeks ago
In this series of three videos, I walk you through the deployment of Hugging Face models on Google Cloud, in three different ways:

- Deployment from the hub model page to Inference endpoints (this video), with the Google Gemma 7B model,
- Deployment from the hub model page to Vertex AI (youtu.be/cBdLw5BnGrk), with the Microsoft Phi-2 2.7B model,
- Deployment directly from within Vertex AI (youtu.be/PFHzfzyY2iY), with the TinyLlama 1.1B model.

Get started at huggingface.co :)

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Deploy Hugging Face models on Google Cloud: from the hub to Inference EndpointsSLM in Action: Local Inference with Arcee Nova 72B and OllamaPhi-2 on Intel Meteor Lake - Physics questionVirtuoso Lite and Virtuoso Medium v2: distilling DeepSeek-V3 to 10B & 32BUnlock the Hidden Value of Your Data! 🌟Benchmarking TurboQuant with MLX on Apple SiliconUnlocking the Power of AI: How Chatbots Enhance Communication!Deep dive: model merging, part 2Arcee AI Drops Trinity Nano & Mini MoE Models!Unlock the Secret Goldmine of Data!SLM in Action: Arcee-Scribe, a 7.7B model for creative writingUnlock the True Power of Data! 💰📊
Julien Simon |

Deploy Hugging Face models on Google Cloud: from the hub to Inference Endpoints

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER