Visit pytorch.org for more information
PyTorch
PyTorch is an open source machine learning framework that accelerates the path from research prototyping to production deployment. This video explains the fundamental concepts behind deep learning, and how tools like PyTorch enable developers to build and deploy AI.
Visit pytorch.org for more information
Visit pytorch.org for more information
updated 7 years ago
Visit pytorch.org for more information
He will share practical techniques for training models with multiple loss functions simultaneously and demonstrate how to easily implement these workflows using the TorchJD library. Whether you are optimizing multi-task architectures or managing complex loss tradeoffs, this session will provide concrete steps to streamline your pipeline.
Register for PyTorch Conference North America: https://hubs.la/Q04v4SL60
Register for PyTorch Conference America today: https://hubs.la/Q04v4SL60
Their session will detail how CRCR enables a tiered onboarding model alongside real-time visibility for downstream repositories in HUD.
Join us in San Jose on October 20 and 21 to learn how CRCR streamlines testing workflows across the ecosystem.
Register today: https://hubs.la/Q04v4SL60
At PyTorch Conference North America, Jen Wei will walk through Muon, Dion, and Dion3 and look at what happens when you add the messy details: sharding, communication, schedulers, and everything else that can go wrong.
If you are interested in optimizers and training dynamics, see you at PyTorchCon NA. Get your ticket: https://hubs.la/Q04v4SL60
Bring your questions about the release to our live Q&A. Andrey Talman (Meta), Nikkita Shulga (Thinking Machines Lab), Joe Spisak (Reflection AI), and Chris Gottbrath (Gottbrath Tech, moderator) will share an overview of PyTorch 2.14 and answer community questions about PyTorch and the new capabilities in the release.
Topics will include:
- NVGEMM and CuTeDSL-generated CUTLASS kernels in Inductor
- The new nccl2 backend for PyTorch Distributed
- Fault-tolerant collectives and process-group reconfiguration in c10d
- Native linear algebra and additional Metal kernel improvements on Apple Silicon
- torch.switch and CUDA graph capture for torch.while_loop
- Declarative dynamic shapes with @dynamic_spec
- Experimental torch.compile support for complex-valued tensors
- Expanded ROCm, Intel XPU, and NVIDIA platform support
PyTorch 2.14 includes 2,995 commits from 487 contributors since PyTorch 2.13. The release includes work across compilation, distributed communication, device support, and accelerator platforms. PyTorch Conference North America 2026 takes place October 20–21 in San Jose, with sessions spanning compiler and runtime work, distributed communication, device portability, release engineering, CI, observability, accelerator integration, contributor infrastructure, and more. Explore PyTorch Conference North America 2026.
The system combines a GPU-resident torch tensor index, tensor-based retrieval operations, custom CUDA kernels for attribute filtering, and TorchScript model execution within a unified serving architecture orchestrated through a Rust and tch-rs backend.
Dhritiman will also share how the approach scales to critical use cases like feed and search at LinkedIn.
Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60
View the poster sessions: events.linuxfoundation.org/pytorch-conference-north-america/program/schedule-posters
#PyTorchCon
The session covers practical lessons from training across diverse institutional datasets and building privacy-aware AI systems for medical imaging.
Register for PyTorch Conference North America 2026, October 20–21 in San Jose: https://hubs.la/Q04w5M9L0
Chris Lattner, CEO and Co-founder, Modular and Executive Vice President of Advanced AI Software and Platforms, Qualcomm, will deliver a keynote at PyTorch Conference North America about an open software platform for heterogeneous compute powered by Mojo and MAX.
PyTorch has always been the place where the best models come together, and now there's a way to get those models onto all kinds of hardware.
If you're interested in Al and compute, join us at the PyTorch Conference in San Jose, CA. Register now: https://hubs.la/Q04v4SL60
Register today to join the open source AI community at PyTorch Conference North America in San Jose, October 20–21: https://hubs.la/Q04tBgv_0
PyTorch Conference North America 2026 is where that community comes together in person. Join us to collaborate with open source pioneers, expand your deployment capabilities, and directly contribute to shaping the future of machine learning.
Register for PyTorch Conference North America today: https://hubs.la/Q04tBgv_0
Join us October 20-21: https://hubs.la/Q04v4SL60
Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60
This approach removes redundant prefill computation and creates streamlined inference pipelines that deliver high performance directly on edge hardware.
Join us San Jose this October 20-21 to learn more: https://hubs.la/Q04v4SL60
Join us in San Jose on October 20th to learn practical strategies for optimizing complex AI workloads: https://hubs.la/Q04v4SL60
In this video, Arun Bhandari, PyTorch Foundation Ambassador and leader of PyTorch Community Nepal, shares why he is excited to connect with the community and explore the latest developments in PyTorch.
Register here to join us in San Jose, CA this October 20-21: https://hubs.la/Q04v88dJ0
They will detail end-to-end production workflows, covering graph capture, model splitting, and serving across large-scale platforms like Facebook and Instagram. Attendees will gain technical insights into advanced optimizations for high TPS and low latency, as well as strategies for multi-accelerator support to handle diverse machine learning workloads at scale.
Register now and connect with the open source AI community in San Jose, CA, October 20-21: https://hubs.la/Q04w5M9L0
In his session at PyTorch Conference North America 2026, Avik Chaudhuri from Meta will introduce Pyrefly, a static type checker for Python which enables end-to-end static shape coverage for PyTorch models.
Join us in San Jose, CA, October 20-21 to explore the future of open source AI and the impact of PyTorch Foundation projects like PyTorch, vLLM, DeepSpeed, Ray, Helion, and Safetensors.
Register now: https://hubs.la/Q04w5M9L0
Register for PyTorchCon NA and connect with the global PyTorch community: https://hubs.la/Q04v88dJ0
PyTorch Conference North America returns October 20–21, 2026, in San Jose.
Register: https://hubs.la/Q04w5M9L0
Join Anshu Raina and Peyman Razaghi of AMD at PyTorchCon NA where they show how the Primus Tuning Agent predicts how fast each option will run, ahead of time, from quick hardware benchmarks, so you find the best configuration without running them all, saving thousands of GPU-hours of trial-and-error.
Register for PyTorchCon North America today: https://hubs.la/Q04v4SL60
PyTorch Conference North America (October 20-21, 2026 in San Jose, CA) will be here before you know it. Don't miss out. Register now: https://hubs.la/Q04w5M9L0
PyTorch Conference features in-depth technical talks, hands-on workshops, and candid conversations spanning the full AI stack, from bare metal infrastructure to applications and agent-based systems. The program features keynote sessions from leading voices in AI and practical deep dives on training, inference, applications across the ecosystem, kernel engineering, responsible AI, and more.
PyTorch Conference is where the open source AI community connects, learns, and shapes what comes next. Register now: https://hubs.la/Q04w5M9L0
Join us in San Jose, CA, October 20-21 for PyTorch Conference North America 2026 to explore the future of open source AI and the impact of PyTorch Foundation projects like PyTorch, vLLM, DeepSpeed, Ray, Helion, and Safetensors.
Register now: https://hubs.la/Q04w5M9L0
Register for PyTorchCon NA: https://hubs.la/Q04v88dJ0
“Efficient Pretraining of LLMs in NVFP4,” will cover the recipe and PyTorch tooling that make this possible, including recipe design, kernel choices, and the API surface needed to bring NVFP4 training into native PyTorch workflows.
The recipes are being upstreamed into the PyTorch ecosystem through TorchAO and TorchTitan, with dense linear NVFP4 training available in TorchAO today.
Register for PyTorchCon North America: https://hubs.la/Q04v4SL60
Register for PyTorchCon North America today: https://hubs.la/Q04v4SL60
#PyTorchCon North America returns October 20–21, 2026, in San Jose.
Register by September 4 to save on your conference pass: https://hubs.la/Q04vsCf70
#PyTorchCon North America returns October 20–21, 2026, in San Jose.
Register by September 4 to save on your conference pass: https://hubs.la/Q04vsCf70
#PyTorchCon North America returns October 20–21, 2026, in San Jose.
Register by September 4 to save on your conference pass: https://hubs.la/Q04vsCf70
Register for PyTorchCon North America today: https://hubs.la/Q04v4SL60
Register today: https://hubs.la/Q04v4SL60
On Wednesday, July 22, 2026, at 11 a.m. PT, PyTorch maintainers and contributors will provide a brief overview of the PyTorch 2.13 release and answer questions from the community live.
Topics will include:
- FlexAttention support on Apple Silicon and deterministic backward computation on CUDA
- The CuTeDSL "Native DSL" backend for Inductor
- nn.LinearCrossEntropyLoss for reducing peak GPU memory
- torchcomms for large-cluster training
- FSDP2 communication overlap improvements
- Torch wheel support for Python 3.15 on Linux, including free-threaded 3.15t builds
- Expanded ROCm, Arm, and Intel XPU platform support
The live Q&A will feature expert panelists Alban Desmaison, Andrey Talman, and Piotr Bialecki, with Chris Gottbrath moderating.
PyTorch 2.13 includes 3,328 commits from 526 contributors since PyTorch 2.12.
Register for the live Q&A today!
PyTorchCon EU featured inspiring keynotes, cutting-edge technical sessions, hands-on workshops, and vibrant discussions shaping the future of AI and deep learning.
From major announcements and deep technical insights to powerful community moments and the unforgettable energy of Paris, here are the highlights from the first PyTorch Conference in Europe.
Whether you joined us in person or are catching up now, we hope these moments inspire your next breakthrough.
► View all the session recordings here youtube.com/playlist?list=PL_lsbAsL_o2DZbCiISSjbucDOW0QOxwBF
#PyTorchCon
See you at PyTorch Conference China and PyTorch Conference North America later this year!
Join us on Wednesday, May 20 at 10:00 AM PT for a live Q&A with panelists Andrey Talman, Alban Desmaison, and Joe Spisak, moderated by Chris Gottbrath. The panel will provide a brief overview of the release and answer your questions live. Register today!
Topics include:
-Device-Agnostic Accelerator Graph Capture
-ProcessGroup Support in Custom Ops
-torch.export.save Support for Microscaling Quantization Formats
-Fused Adagrad Optimizer Support
-FlightRecorder Updates
-Multi-GPU and Multi-Node Profiling Improvements
-Updated Backend Selection for torch.linalg.eigh on CUDA
-Expanded CUDA, ROCm, XPU, MPS, and Arm Platform Support
Register today.
Panelists:
Andrey Talman is a Software Engineer at Meta, primarily focused on open source releases for PyTorch and its ecosystem libraries. He works on release management, continuous integration, and process improvements, ensuring high-quality and timely delivery of PyTorch and related projects.
Alban Desmaison is a Research Engineer at Meta and the Lead Core Maintainer of PyTorch.
Joe Spisak is Vice President of Product and Head of Open Source at Reflection AI. He is a PyTorch core maintainer, serves on the PyTorch Foundation Governing Board, and previously worked at Meta.
Moderator:
Chris Gottbrath is a Group Technical Program Manager supporting PyTorch at Meta and Chair of the PyTorch Foundation Marketing Committee.
Whether you're new to open source or an experienced contributor, the Docathon offers opportunities to improve tutorials, guides, examples, and website content. Many issues are beginner-friendly, with tasks labeled by skill level and support available through the PyTorch Discord.
Contributors can make an immediate impact while learning more about PyTorch modules, tutorials, and workflows.
Schedule
May 5
Kickoff and Q&A, 10:00 AM PT
May 6 to May 15
Submissions and feedback
May 16 to May 18
Final reviews
May 20
Winner announcements
RSVP to receive updates, participation details, and event instructions.
WideEP—wide expert parallelism fails not because experts are expensive, but because routing ignores where state already lives. In PyTorch LLM serving with vLLM, WideEP fans tokens across many experts while KV caches accumulate unevenly across data-parallel replicas. When routing is unaware of KV placement and per-replica load, requests land on replicas that cannot reuse cache or make progress efficiently and latency spikes as expert fan-out grows.
The fix is not reshaping expert parallelism, but making routing data-parallel aware using signals vLLM already exposes. In this talk, we show how llm-d extends its router to leverage KV-cache locality and load awareness when routing WideEP flows. Rather than treating replicas as interchangeable, the router prefers replicas with warm KV state and available capacity, aligning routing decisions with vLLM’s execution reality and reducing cache fragmentation.
This session walks through how KV-aware, data-parallel routing changes WideEP inference in practice: which signals matter, how routing behavior evolves, and where the gains come from. Attendees leave with a clear mental model for when KV- and load-aware routing unlocks higher throughput.
Deep learning has revolutionised imaging, a foundation of science and healthcare. DeepInverse is the PyTorch library for solving imaging problems, unifying deep learning methods (e.g. diffusion models), physics (medical, optics) and modern tooling. In this talk, we’ll show how the PyTorch community can get involved in this exciting yet accessible application of open-source AI.
AI methods in imaging must model the imaging physics, leading to interesting engineering problems e.g. efficient differentiable ops, physics-informed losses. We’ll show notebooks on real use-cases: accelerating brain MRI, reducing radiation in CT scans, imaging black holes.
PyTorch enthusiasts at any level/background can contribute - from training infra for scientific data to high-level generative modelling frameworks - their AI engineering skills can directly impact imaging across multiple fields.
DeepInverse is supported by a growing international user community and proudly rooted in Paris. We’ve joined the PyTorch Ecosystem and received the Prix Science Ouverte in 2024. We’re excited to join the PyTorch Conf to celebrate the vibrant French developer community!
The Hugging Face transformers library is built on pure PyTorch and can be succinctly described as a model-definition framework. It provides an unified, familiar, clear and concise interface to multiple machine learning architectures across modalities.
Serving and inference optimizations are not its focus.
However, transformers model definitions become the de-facto reference implementations multiple other projects use. This includes training libraries, fast deployment engines such as vLLM and SGLang, and on-device libraries like MLX and llama.cpp.
This session describes the path towards increasingly simpler downstream integration of transformers models into inference and deployment libraries, and how transformers and PyTorch core features enable the ecosystem to enjoy newly-released models as soon as they are released.
We'll go through the journey towards easier modeling, which implies easier downstream porting and adaptation. The end-game is pure interoperability, where no code changes are required! This is now possible with vLLM and SGLang, and we'll show how. We'll end up discussing our ideas on upcoming interop features with MLX and llama.cpp.
Rapid advances in AI have expanded the range of capabilities required for successful real-world deployment. Understanding where we are in this multi-dimensional frontier is essential for accelerating innovation through effective quality assurance. Rigorous evaluation is increasingly difficult to scale as development requires testing many checkpoints across numerous benchmarks. Model comparison is further complicated by limited transparency of reported results. This talk explores challenges, best practices, and open-source tools that elevate evaluation to a core component of LLM development, delivering continuous signals across the model lifecycle.
We discuss principles for standardizing evaluation methods and improving consistency through practical patterns and anti-patterns, and examples of integrating the science of evaluation directly into model development. Using Nemo-Evaluator, an open-source scalable evaluation tool, we demonstrate modular architectures that enable transparent, reproducible measurement. Finally, we show how Nemo-Evaluator supports reproducible evaluation for the Nemotron model family, helping enable one of the most open development processes in modern AI.
Passive acoustic monitoring is a powerful tool for wildlife conservation, but deploying deep learning models in remote rainforest environments introduces strict constraints on power, memory, and compute. In this talk, we present an end-to-end PyTorch-based pipeline for detecting and analyzing the endangered three-wattled bellbird using embedded deep learning systems.
We cover the full lifecycle from audio preprocessing and model training in PyTorch to optimization and deployment on resource-constrained embedded devices. Topics include model architectures for sparse bioacoustic event detection, handling extreme class imbalance, model compression and quantization, and practical trade-offs between accuracy, latency, and power consumption.
The session emphasizes real-world lessons learned deploying machine learning at the edge, where unreliable connectivity, noisy signals, and limited hardware define success more than benchmark metrics. Attendees will gain practical patterns for building and deploying PyTorch models for embedded and edge AI applications with real environmental impact.
Production LLM serving faces a critical trade-off: while continuous batching maximizes throughput, it often sacrifices SLAs due to Head-of-Line (HoL) blocking. When long-context requests hijack the engine, tail latencies spike. Without fine-grained preemption, guaranteeing priority or fairness remains nearly impossible.
We propose a solution: Chunked Decoding. By treating a fixed number of tokens as a "time slice," we bring 50 years of OS scheduling wisdom to inference. This technique decouples generation from completion, enabling a preemptive multitasking environment for LLMs.
In this talk, we present a sidecar implementation for PyTorch-based servers (like vLLM) that orchestrates decoding in manageable chunks. This allows the system to pause, hold, or swap requests mid-stream without discarding the KV cache. We will share early evaluation results, discussing how varying chunk sizes impact priority handling and tail latency. Attendees will learn how a sidecar approach enables sophisticated scheduling while keeping the core engine lean—offering a blueprint for integrating preemptive scheduling into the next generation of model servers.
CUDA streams are a widely-used method for parallelizing GPU computation on NVIDIA GPUs. They have long been requested by our users and enable multiple key capabilities - overlapping communication and compute kernels, training on multiple batches in parallel and parallelizing kernels, all of which are needed for achieving SOTA training performance. Another key capability is activation offloading - this can be applied to any model to prevent OOMs by asynchronously storing activations in cpu memory until they are needed by the model.
Before this work, torch.compile previously would graph break on CUDA stream contexts, which can be costly for models that utilize streams. Although workarounds exist (e.g. wrapping stream manipulation into custom ops), these solutions add complexity and create friction in the user experience. By enabling seamless CUDA stream support in PT2, we allow our users to leverage the familiar eager APIs for stream assignment and synchronization directly within torch.compile. This not only simplifies the workflow but also ensures that models using custom streaming patterns can run efficiently out-of-the-box without manual intervention or code restructuring.
ExecuTorch extends PyTorch's reach to the most resource-constrained devices: microcontrollers, DSPs, and specialized neural processing units powering always-on sensors, wearables, and embedded systems. In this talk, we'll share the current state and roadmap for running ExecuTorch on platforms where every kilobyte of memory and milliwatt of power matters.
What you'll learn:
- How ExecuTorch's design enables deployment from ultra-low-power MCUs to DSP and NPU accelerators, all from a single PyTorch workflow
- The state of backend support for Cadence DSPs, ARM Ethos-U and Cortex-M
- Practical considerations for deploying models with sub-megabyte footprints and milliwatt power budgets
- Case studies spanning always-on audio, embedded vision, and TinyML applications
With the advent of geospatial foundation models, unexplored use cases are emerging that require well-curated datasets. Currently, no standardised approach exists for creating such AI-ready geospatial datasets. In this session, we introduce TerraKit: a comprehensive open-source Python library for retrieving, and processing geospatial data, that seamlessly integrates with upstream geospatial model training libraries such as TorchGeo or TerraTorch.
From raster/vector annotations, TerraKit will match, download, process, align and split the requested data source (e.g., EarthData, CDSE, Planetary Computer) based on user specifications provided by a simple configuration file. TerraKit also supports spatial train/val splits and exports datasets in standard formats such as TACO datasets. TerraKit streamlines the pipeline from raw EO data to AI-ready datasets, accelerating the development of custom geospatial applications, and ensuring query and processing pipelines are reproducible. By lowering the barrier to entry, a wider community of TorchGeo and TerraTorch users are empowered to leverage foundation models for Earth observation.
Uncertainty quantification is becoming more and more important as neural networks are used for increasingly critical tasks. Bayesian neural networks (BNNs) inherently provide a measure of their own uncertainty, but can be either hard to implement or inflexible if one uses common frameworks. In this session I discuss how to efficiently implement BNNs using Variational Inference within PyTorch and present torch_blue, a light-weight open source library that implements these methods with the goal of being easy to pick up, yet flexible enough for research on BNNs.
Renewable energy is clean — but it’s also inherently variable. Solar PV generation can change dramatically within minutes due to cloud cover and weather conditions, making accurate short-term forecasts essential for grid stability, energy trading, and smart-home optimisation.
Open Climate Fix builds open and high-impact forecasting tools to accelerate the transition to a low-carbon energy system. One of these projects is Open Quartz Solar Forecast: an open-source model that uses public PV generation data, site metadata, and numerical weather prediction variables to forecast solar power for any location.
In this talk, I’ll present a real case study from my Google Summer of Code project where I implemented and trained a Temporal Fusion Transformer for multi-horizon solar forecasting. I’ll cover the practical engineering challenges behind making transformer forecasting work in Python: building continuous training windows, aligning weather forecast steps with observations, separating static vs time-varying features, and stabilising training using PyTorch Forecasting and PyTorch Lightning.
Attendees will leave with reusable patterns for real-world time-series forecasting pipelines.
PyTorch's success depends on more than users—it needs engineers who understand what's inside. Engineers who can debug framework issues, optimize at the systems level, contribute upstream, and build what comes next. But ML education today produces practitioners who call APIs without understanding them. They train models without knowing why Adam needs 3× the memory of SGD, or what happens when they call loss.backward().
TinyTorch is a 20-module open-source curriculum that closes this gap. Students construct PyTorch's core components—tensors, autograd, optimizers, CNNs, transformers—in pure Python, building a complete framework where every operation is code they wrote. By the final module, they don't just use PyTorch; they understand how to build it.
The curriculum uses progressive disclosure, systems-first profiling from Module 01, and build-to-validate milestones—recreating ML breakthroughs from Perceptron (1958) through Transformers (2017), culminating in MLPerf-style benchmarking.
TinyTorch is how we grow the next generation of PyTorch contributors and the engineers who will build what comes after.
Open source: mlsysbook.ai/tinytorch
Distributed neural network training frameworks typically optimize for specific architectures while minimizing communication overhead. Transformer layers can be efficiently parallelized, but other operations such as convolutions often remain inefficient. This creates bottlenecks for complex model architectures.
Moreover, existing tensor parallelism strategies typically replicate input data across all processes, creating redundant I/O that scales poorly with input size. In applications with heavy I/O demands-weather forecasting, medical imaging, or video processing-unsharded input data creates additional data-loading bottlenecks that could benefit from parallelization.
Jigsaw is a PyTorch library that shards both model weights and input data across parallel processes. It maintains a PyTorch-like interface while parallelizing activations, convolutions, linear layers, and attention through a distributed matrix multiplication backend. We demonstrate the usability of Jigsaw across a wide range of model architectures and shows performance when scaling multi-billion-parameter models sharded across up to 8 processes and compares the scalability to DDP, FSDP, and Megatron-LM approaches.


