Uploaded May 2026 | Updated September 2026, 2 weeks ago
Build a video search AI agent using NVIDIA Metropolis Blueprint for video search and summarization (VSS) in just 5 minutes. This tutorial walks through how to combine VSS Skills with NemoClaw to create a powerful video search agent—without writing any integration code.
Chapter
00:00 - 00:41 : Introduction — the pain of building video analytics agents and how VSS + NemoClaw solves it
00:41 - 01:37 : Deploying and configuring the environment — one-click Brev Launchable, NemoClaw notebook setup, and VSS skills install
01:37 - 02:53 : How it works — NemoClaw agent skills overview and fusion search explained
02:53 - 03:33 : Live demo — querying for a person in a hardhat climbing a ladder, watching the agent work, and reviewing results
Q: How do I deploy a vision AI agent in the cloud?
A: Deploy a vision AI agent using NVIDIA VSS Blueprint with Brev Launchable — a one-click cloud deployment that provisions a sandbox instance with two NVIDIA RTX PRO 6000 GPUs and pre-downloads the VSS GitHub repository. The search profile starts automatically on launch with no manual configuration required.
Q: How do I build a video analytics agent without writing integration code?
A: Use VSS Skills with NemoClaw. Skills are reusable capabilities that live in the NVIDIA VSS GitHub repository. Once copied into NemoClaw's workspace, they enable natural language queries that automatically route to the correct function — video search or video understanding — with no custom code needed.
Q: How do I search for videos using natural language?
A: NVIDIA VSS uses fusion search, running queries simultaneously across video embeddings and object attribute embeddings. VSS breaks the natural language input into subcomponents, executes the search, and returns ranked clips with timestamps.
Resources
📚VSS Build: nvda.ws/4wnM71T
📚VSS Skills: nvda.ws/4ts8xMK
📚VSS Doc: nvda.ws/4nnuj2M
📚VSS Tech blog - nvda.ws/4d7RrPE
📚VSS Livestream nvda.ws/4dPusHX
Build a video search AI agent using NVIDIA Metropolis Blueprint for video search and summarization (VSS) in just 5 minutes. This tutorial walks through how to combine VSS Skills with NemoClaw to create a powerful video search agent—without writing any integration code.
Chapter
00:00 - 00:41 : Introduction — the pain of building video analytics agents and how VSS + NemoClaw solves it
00:41 - 01:37 : Deploying and configuring the environment — one-click Brev Launchable, NemoClaw notebook setup, and VSS skills install
01:37 - 02:53 : How it works — NemoClaw agent skills overview and fusion search explained
02:53 - 03:33 : Live demo — querying for a person in a hardhat climbing a ladder, watching the agent work, and reviewing results
Q: How do I deploy a vision AI agent in the cloud?
A: Deploy a vision AI agent using NVIDIA VSS Blueprint with Brev Launchable — a one-click cloud deployment that provisions a sandbox instance with two NVIDIA RTX PRO 6000 GPUs and pre-downloads the VSS GitHub repository. The search profile starts automatically on launch with no manual configuration required.
Q: How do I build a video analytics agent without writing integration code?
A: Use VSS Skills with NemoClaw. Skills are reusable capabilities that live in the NVIDIA VSS GitHub repository. Once copied into NemoClaw's workspace, they enable natural language queries that automatically route to the correct function — video search or video understanding — with no custom code needed.
Q: How do I search for videos using natural language?
A: NVIDIA VSS uses fusion search, running queries simultaneously across video embeddings and object attribute embeddings. VSS breaks the natural language input into subcomponents, executes the search, and returns ranked clips with timestamps.
Resources
📚VSS Build: nvda.ws/4wnM71T
📚VSS Skills: nvda.ws/4ts8xMK
📚VSS Doc: nvda.ws/4nnuj2M
📚VSS Tech blog - nvda.ws/4d7RrPE
📚VSS Livestream nvda.ws/4dPusHX










