WAD at CVPR
[CVPR22 WAD] Keynote - Ashok Elluswamy, Tesla
updated
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
03:00 Lesson 1: Scaling as a game of 9s
08:47 Lesson 2: Maps are an extremely powerful sensor
12:27 Lesson 3: Scaling is the law
13:35 Lesson 4: Rare events are merely not frequent yet
18:17 Lesson 5: Building a scalable critic is essential
20:26 The Waymo Flywheel
22:35 Conclusion
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:51 Video World Models
03:22 Current Challenges
04:12 Motion Appearance Decoupling
07:52 Infinite-Length Video Generation
11:22 LayerSync
15:09 SenCache
20:09 Limitations and Future Directions
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:06 The optimistic view
08:24 The pessimistic view
23:28 New Datasets and Applications
26:40 Conclusion
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:04 Evolution of Simulation
06:28 Simulation 2.0: Neural Reconstruction
09:55 Simulation 3.0: Generative World Models
11:26 Real-Time World Model (OmniDreams)
21:30 OmniDreams Results
24:10 OmniDreams as a World-Action Model
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:55 Can Autonomy Generalize?
05:19 End-to-End AI for Driving
09:47 Generative AI for Autonomy
12:45 Could we learn to drive in a World Model?
14:50 End-to-End should extend to Safety and Simulation too
19:07 Platform will reach greater scale than vertically integrated autonomy
23:15 Cost matters
27:29 Wayve Labs
29:24 Conclusion
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro: The Bitter Lesson
01:15 Scaling has Diminishing Returns
03:38 A Hay in a Haystack
07:55 Toward Efficient Scaling: AdaBoost
12:30 Scenario Boosting
16:15 Scenario Boosting as a Game
18:47 Building the Meteor Agent
21:05 The Meteor Agent in Action
25:51 Conclusion
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:14 Tesla Self-Driving
06:35 End-to-end Foundation Model for all of Robotics
07:45 The Case for and Challenges in E2E driving
11:09 Challenge 1: Curse of Dimensionality
15:55 Challenge 2: Interpretability and Safety
18:06 Challenge 3: Evaluation
21:58 Conclusion
Talk given at the CVPR Workshop on Autonomous Driving 2026: https://cvpr2026.wad.vision/.
00:00 Intro
01:10 Design Characteristics of Autonomous Systems
02:57 Evolution of Self-Driving
07:30 "Next Gen" AV2.0
11:47 Waabi World
16:00 Digital Twinning
19:10 The Rise of Generative Models
21:18 How can we prove safety?
28:06 Conclusion
00:00 Intro: What is Argoverse
01:45 (Un)supervised Scene Flow
07:07 Scene Flow Challenge Winners
10:13 Scenario Mining
14:20 Scenario Mining Challenge Winners
22:00 Future Directions
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
00:56 End-to-End Autonomous Driving
05:07 Development Stack of Driving Models
08:00 World Engine (E2E Driving 2.0)
10:28 Nexus: Decoupled Diffusion Sparks Adaptive Scene Generation
14:26 MTGS: Multi-Traversal Gaussian Splatting
17:07 NAVSim v2: Pseudo-Simulation for Autonomous Driving
21:04 Key Challenge: Scalable RL
24:18 Takeaways
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
00:48 Trends in AI-based Robotics
01:58 Trend 1: Probabilistic Robotics
09:53 Trend 2: Deep Networks
12:47 Uncertainty-Aware Panoptic Segmentation (EvPSNet)
14:54 Uncertainty-Aware LiDAR Panoptic Segmentation (EvLPSNet)
16:25 uPLAM: Robust Panoptic Localization & Mapping
19:00 Do We Really Need Maps?
21:04 Trend 3: Foundation Models
22:13 Visual Language Maps for Robot Navigation
24:03 FM-Loc: Using Foundation Models for Vision-Based Navigation
29:26 Outlook
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
01:46 NVIDIA Cosmos Foundation Model
04:32 LiDAR Representation
05:11 LiDAR Tokenization
08:12 Conditions for LiDAR Generation
12:31 Joint Modeling of LiDAR & RGB
16:25 Final Thoughts
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
01:12 Section 1: Deep Dive into NuPlan & Planning Metrics
04:06 Reactive Agents
07:10 State-of-the-Art on NuPlan
08:03 Adopting to City-Specific Behavior
15:55 Motion Understanding
19:03 Utilizing Pseudo-Labels
24:02 Scenario Mining
26:42 CVPR Scenario Mining Challenge
27:29 Conclusion
00:00 Introduction
00:35 What is Argoverse (2)?
02:36 Challenge 1: (Un)supervised Scene Flow
05:50 Scene Flow Leaderboard
06:23 Winning Method (Delta Flow)
08:46 Highlighted Method (Floxels)
10:12 Challenge 2: Scenario Mining
17:32 Scenario Mining Leaderboard
21:03 Challenge 3: Multi-Agent Forecasting
22:50 Multi-Agent Forecasting Leaderboard
23:56 Future Directions
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
00:25 Scalable Simulation for E2E Stacks
02:28 Gaussian Splatting
03:08 DeSiRe-GS: 4D Street Gaussians
06:18 Incomplete Objects from 3DGS
07:28 R3D2: Realistic 3D Asset Insertion via Diffusion
11:36 X-Drive: Cross-Modality Multi-Sensor Data Synthesis
15:30 Feed-Forward 3DGS & DrivingRecon
18:40 PixelGaussian: Generalizable Gaussian from Arbitrary Views
20:00 Onboard and Offboard Models
21:26 S2GO: Streaming Sparse Gaussian Occupancy Prediction
26:12 Neural Simulator
27:25 Conclusion
00:00 Introduction
00:44 About Nexar
02:54 Nexar Open Dataset
06:02 Nexar Crash Prediction Challenge
09:12 Leaderboard
11:22 1st Place
12:19 2nd Place
13:18 3rd Place
15:55 Examples of Difficult Cases
17:10 Where Next?
18:12 Key Takeaways
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
01:48 Waymo One Service
03:30 Challenging Driving Examples
04:28 Waymo's Safety Record
05:54 The Waymo Foundation Model
07:53 Adverse Weather
12:47 Sensor Fusion
14:22 Occlusion Reasoning
19:32 Scene Understanding
23:00 Data
24:30 Depots - Enabling Operational Excellence
25:57 Fast Real-Time Decisions
Talk given at the CVPR Workshop on Autonomous Driving 2025: https://cvpr2025.wad.vision/.
00:00 Introduction
00:40 XPeng's Driving Assistance System
02:26 Capabilities in Mass Production
03:56 SW 3.0 AI Factory
06:45 Foundation Model
10:40 Alignment & CoT
15:22 Supervised Finetuning (SFT)
17:30 Reinforcement Learning
19:35 Summary
21:23 Additional Demos
00:00 Introduction
01:25 Datasets and Benchmark Overview
03:28 2025 Challenges
04:43 End-to-End Driving Challenge
06:30 New Waymo Dataset of Rare Events
09:22 End-to-End Driving Challenge Winners
10:08 1st Place: Poutine
21:50 Perception Dataset
22:41 Motion Dataset
23:36 Interaction Prediction Challenge
26:19 Interaction Prediction Challenge Winners
27:20 1st Place: Parallel ModeSeq
32:00 Sim Agents Challenge
34:32 Sim Agents Challenge Winners
35:05 1st Place: TrajTok
40:05 Waymax Update
40:30 Scenario Generation Challenge
42:07 Scenario Generation Challenge Winners
42:50 1st Place: SimFormer
49:42 Conclusion
00:00 Hongyang Li: End-to-end Autonomous Driving: Past, Current and Onwards
25:50 Wolfram Burgard: Probabilistic and Deep Learning Approaches for Automated Driving
56:00 Laura Leal-Taixe: Repurposing Generative Models for 3D Data
01:13:43 Deva Ramanan: Perception and Simulation for Self-Driving Vehicles
01:42:26 Argoverse Challenges
02:07:13 Wei Zhan: Scalable Neural Simulation for Autonomy
02:35:40 Nexar Challenges
02:54:52 Chen Wu: Solving Real-World Challenges of Large-Scale AV Deployment
03:22:24 Waymo Open Dataset Challenges
04:12:53 Xianming Liu: Scaling up Autonomous Driving via Large Foundation Models
00:00 Introduction
01:43 AV Stack Evolution
03:43 NVIDIA's Generative AI & Drive Thor
05:11 Today's Talk
05:41 Foundation Models in Data Tools
06:47 Visual Retrieval
07:55 Semantic Scenario Clustering
08:54 Foundation Model powered AV Stack (AVFM)
10:40 PARA-Drive
11:10 Hydra-MDP
11:41 Trajeglish
12:21 Video Generation
13:53 Video Modeling
16:31 Video Tokenization
17:48 Open-Loop Trajectory Prediction
18:54 Simulation
20:15 Physics, Dynamics and Editing
22:00 Closed-Loop Testing
23:34 Object Insertion and Handling Large Scenes
25:05 Scene Generation
25:56 fVDB Release
27:28 Conclusion
00:00 Introduction
02:30 Perception Dataset
04:30 Motion Dataset
06:22 Benchmarks
07:25 Waymax Simulator Release
09:18 2024 Challenge Overview
10:47 Motion Prediction Challenge
15:07 1st Place: MTR v3
20:12 Occupancy and Flow Challenge
23:20 1st Place: DOPP
25:11 2nd Place: STNet
27:45 3rd Place: HGNET
29:57 Sim Agents Challenge
33:57 1st Place: BehaviorGPT
38:40 Unranked: MPS
42:42 3D Semantic Segmentation Challenge
45:28 1st Place: Point Transformer V3
49:51 Conclusion
00:00 Introduction
01:52 Waymo's Experience and Service Areas
03:26 Driving Examples
05:30 Construction, Emergency Vehicles and Freeways
09:25 Waymo's Safety Record
10:34 Rare Event Examples
11:21 Long-Tail Handling
12:48 LLM/VLM Reasoning
14:13 Open-Vocabulary Perception using VL Distillation
19:20 MotionLM: Modeling Driving as a Conversation
23:10 MoST: Scene Tokenization for Motion Prediction
29:25 Open Questions in the LLM/VLM Space
31:25 Waymo Rider Experience
32:17 Conclusion
00:00 Introduction
00:49 What is Argoverse (2)
01:15 Sensor Dataset
01:55 LiDAR Dataset
02:24 Motion Forecasting Dataset
02:55 Map Change Dataset
03:57 Argoverse post-Argo
05:07 (Un)supervised Scene Flow Challenge
18:27 End-to-End Forecasting Challenge
32:22 Multi-Agent Motion Forecasting Challenge
41:56 4D Occupancy Forecasting Challenge
53:32 Conclusion
00:00 Introduction
00:59 Autonomy Architectures
03:25 Structured Optimization vs E2E Learned Policies
06:38 Hybrid Architecture
08:34 Early Sensor Fusion
11:03 Learned Depth Completion
12:45 Handling Debris
15:33 The IID Assumption is Brittle
16:37 OOD Detection
20:06 OOD Detection using Foundation Model Embeddings
23:30 Domain Knowledge Helps Create Data
25:06 Directed Scenario Generation
28:41 JAX-Accelerated Reinforcement Learning
30:30 Symbol Grounding & Conclusion
00:00 Introduction
01:07 The Road to Embodied AI
01:57 From Vision to Action
03:43 AV2.0: An End-to-End System
04:57 Promises of End-to-End Systems
06:18 Focus Areas to Create Embodied AI
06:53 Focus Area: Simulation
08:36 Ghost Gym: Neural Simulator
09:28 Introducing PRISM-1
12:09 Launching WayveScenes101
14:07 GAIA: Generative World Model
18:13 Focus Area: Multimodality
19:27 LINGO-1: Offline Video QA
21:21 LINGO-2: Grounded Vision-Language-Action Model
24:32 AV2.0 Reimagining the Driving Experience
25:30 Focus Area: Engineering to Product
28:00 Embodied AI Foundation Models
31:38 Safety by Design
33:05 Conclusion
00:00 Introduction
00:52 The Gap between Industry and Academia
02:02 Driving Datasets
02:51 Driving Simulators
04:41 The MetaDrive Simulator
07:26 ScenarioNet
11:29 TrafficGen
13:30 Closed-Loop Adversarial Training (CAT)
16:18 SimGen Driving Scene Generation
21:04 Mobility Anywhere
22:55 MetaUrban Simulator
25:35 Conclusion
00:00 Introduction
00:30 Benchmarking AVs is hard
02:36 Flaws in Open Loop Benchmarks
05:18 What about Simulation
07:18 Non-reactive Simulation
08:11 No Sensor Simulation
11:54 No Traffic Simulation
13:33 NAVSIM Metrics
16:00 The Predictive Driver Model (PDM) Score
18:10 Comparing PDMS vs Closed-loop Metrics
20:09 OpenScene Release to make E2E Research more Accessible
24:15 OpenScene Benchmark Results
25:54 2024 NAVSIM Challenge
26:35 Limitations
27:56 Conclusion
00:00 Introduction
00:58 Perceiving Humans in 4D
03:17 Improving Robustness in Human Motion Estimation
04:49 Human Mesh Recovery (HMR)
07:46 HMR Examples
11:28 Application for Tracking
13:47 SLAHMR: Decoupling Human and Camera Motion in the Wild
18:36 SLAHMR Examples
21:30 Reconstructing Humans and Scenes
22:24 Reconstructing 3D Humans in Contact
22:55 Reasoning about Humans in Simulation
23:31 Building Realistic Avatars
24:00 Conclusion
00:00 Introduction
00:27 A Simplified Self-Driving Stack
01:05 ViP3D: End-to-End Visual Prediction
02:25 Scalability
04:56 3D Occupancy Prediction
07:28 Auto-Labeling Occupancy Datasets
12:25 The Occ3D and SSCBench Benchmarks
13:47 Handling New Geo-Locations
16:31 VectorMapNet
18:59 Neural Map Priors
22:10 Map Prior Improving Range and Robustness
23:35 Conclusion
00:00 Introduction
01:50 What is Argoverse (2)
7:17 Multi-Agent Motion Forecasting Challenge
20:57 End-to-End Forecasting Challenge
36:26 4D Occupancy Forecasting Challenge
44:38 Self-Supervised Scene Flow Challenge
00:00 Introduction
00:35 First Winner
12:35 Second Winner
00:00 Introduction
06:15 What's New
08:04 Waymax & Perception Object Assets
11:13 New Dataset Format
14:43 Polling
21:47 Pose Estimation Challenge
31:57 Motion Prediction Challenge
42:30 Sim Agents Challenge
54:42 2D Video Panoptic Segmentation
00:00 Introducing Social Forecasting
03:05 Robot Experiment
04:34 Representation Learning: Perception
08:52 Representation Learning: Social Forecasting
16:57 Representation Learning: Planning
19:13 7 Foundational Principles for Autonomous Mobility
19:22 P1: Predictive Coding
21:18 P2: Opposites
23:30 P3: Multimodality
24:09 P4: Focused Learning
25:21 P5: Compositionality and Sparsity
26:41 P6: Causality
28:07 P7: In Vivo Practice
29:11 Summary
00:00 Introduction
02:09 Occupancy Networks Recap
04:04 Generative Modeling of Lanes
06:15 Object Prediction & Properties
08:10 Fleet Auto-Labeling
13:23 Learning a General World Model
18:04 Large Compute
20:02 Q&A
00:00 Introduction
00:55 Long-Tail Problem
02:11 Momenta Flywheel Strategy
02:57 Momenta's Mapless Solution
04:55 DDLD: Data Driven Landmark Detection
10:48 DLP: Deep Learning Planning
17:35 DDPF: Data Driven Pose Fusion
25:25 Summary
00:00 Introduction
01:25 Examples of Distribution Shift
03:48 Takeaways from Benchmark Creation Process
08:05 What Should We Do About Distribution Shift?
08:49 Detecting Distribution Shift
19:51 Improve Predictions under Shift
24:06 Diversify and Disambiguate (DivDis)
26:50 Overall Takeaways
00:00 Introduction
03:12 nuPlan
04:53 Open-Loop and Closed-Loop
05:57 Lesson 1: A route centerline is all you need for ego-forecasting
08:25 Lesson 2: Rule-based beat learned planners in closed-loop
10:06 Lesson 3: Planning & ego-forecasting tasks are misaligned
16:01 Lesson 4: Learned/long-term forecasting doesn't improve closed-loop
17:49 CARLA
22:00 TransFuser
22:58 Lesson 5: Target point conditioned methods learn shortcut when out-of-distribution
25:26 Lesson 6: Average pooling remove spatial info and increases bias
27:11 Lesson 7: Waypoint representations work well as they learn to interpolate
28:48 Summary
0:00 Introduction
0:21 Waymo's Progress
3:36 Adverse Weather
5:21 Navigating Construction Zones
6:49 Handling Dense Traffic
7:35 Safety
10:28 Core Architecture: MotionDiffuser
14:11 Core Architecture: SWFormer
17:37 Long-Tail: Rare Example Mining
21:04 Long-Tail: Open-Set Perception
24:56 Holistic Approach to Handling Adverse Weather
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20
2022-06-20


