Deploy Hugging Face models on Google Cloud: directly from Vertex AI @juliensimonfr
Deploy Hugging Face models on Google Cloud: directly from Vertex AI  @juliensimonfr
Uploaded April 2024 | Updated September 2026, 2 weeks ago
In this series of three videos, I walk you through the deployment of Hugging Face models on Google Cloud, in three different ways:

- Deployment from the hub model page to Inference endpoints (youtu.be/mlU-2QYx4a0), with the Google Gemma 7B model,
- Deployment from the hub model page to Vertex AI (youtu.be/cBdLw5BnGrk), with the Microsoft Phi-2 2.7B model,
- Deployment directly from within Vertex AI (this video), with the TinyLlama 1.1B model.

Get started at huggingface.co :)

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos. Follow me on Medium at julsimon.medium.com or Substack at https://julsimon.substack.com. ⭐️⭐️⭐️
Deploy Hugging Face models on Google Cloud: directly from Vertex AIMy prompt:  Build a street art gallery websiteLLMs from the trenches - Data is how you create a competitive advantage, not modelsArcee Maestro 7B - A small language model that outperforms o1-previewArcee Conductor and Zerve: Bringing Model Routing to AI and Data Science WorkflowsUnlocking the Minds of Business Leaders: What Drives Their AI Decisions?Unleashing the Future of AI: Experience the Lightning Speed of Arcee-Lite!Run performant and cost-effective GenAI Applications with AWS Graviton and Arcee AIHugging Face, the story so farArcee Nova 72B running on my Mac #ai #largelanguagemodels #opensourceThis model costs 15 cents per million tokens!How Synapse Medicine leverages Hugging Face to improve medication safety
Julien Simon |

Deploy Hugging Face models on Google Cloud: directly from Vertex AI

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER