Uploaded April 2026 | Updated September 2026, 2 weeks ago
Large language models for code are transforming how developers create, maintain, and understand software—but building these models from the ground up requires the right tools and know-how. In this talk, we’ll explore how to leverage the NVIDIA NeMo framework and NVIDIA’s accelerated computing infrastructure to train state-of-the-art (SOTA) coding LLMs capable of handling multiple programming languages.
Gain practical insights into dataset preparation, multilingual training challenges, and effective strategies for balancing diverse code sources. We’ll cover workflow design for large-scale distributed training, discussing optimization techniques and best practices for scaling across GPU clusters. The session will also highlight real-world applications of multilingual coding LLMs—from powering intelligent developer assistants to managing complex enterprise codebases.
Miguel Martinez | Sr. Applied Deep Learning Researcher | NVIDIA
Meriem Bendris | Sr. Deep Learning Data Scientist | NVIDIA
Oleg Sudakov | Solutions Architect | NVIDIA
Key Takeaways:
A practical understanding of how to use NVIDIA NeMo and NVIDIA infrastructure to train SOTA coding LLMs
Insights into dataset preparation, multilingual training challenges, and strategies to overcome them
Knowledge of optimization techniques and best practices for scaling training across large compute environments
Awareness of real-world applications for multilingual coding LLMs, from developer assistance to enterprise codebase management
A clear roadmap for experimenting with and deploying advanced coding LLMs using NVIDIA’s technology stack
Industry: All Industries
Topic: Agentic AI / Generative AI - Code / Software Generation
Technical Level: Technical - Intermediate
Intended Audience: Data Scientist
NVIDIA Technology: DGX Platform, TensorRT, NeMo, Triton, NVIDIA NIM, NVIDIA AI Enterprise, DGX Cloud, Nemotron
#nvidiagtc
Large language models for code are transforming how developers create, maintain, and understand software—but building these models from the ground up requires the right tools and know-how. In this talk, we’ll explore how to leverage the NVIDIA NeMo framework and NVIDIA’s accelerated computing infrastructure to train state-of-the-art (SOTA) coding LLMs capable of handling multiple programming languages.
Gain practical insights into dataset preparation, multilingual training challenges, and effective strategies for balancing diverse code sources. We’ll cover workflow design for large-scale distributed training, discussing optimization techniques and best practices for scaling across GPU clusters. The session will also highlight real-world applications of multilingual coding LLMs—from powering intelligent developer assistants to managing complex enterprise codebases.
Miguel Martinez | Sr. Applied Deep Learning Researcher | NVIDIA
Meriem Bendris | Sr. Deep Learning Data Scientist | NVIDIA
Oleg Sudakov | Solutions Architect | NVIDIA
Key Takeaways:
A practical understanding of how to use NVIDIA NeMo and NVIDIA infrastructure to train SOTA coding LLMs
Insights into dataset preparation, multilingual training challenges, and strategies to overcome them
Knowledge of optimization techniques and best practices for scaling training across large compute environments
Awareness of real-world applications for multilingual coding LLMs, from developer assistance to enterprise codebase management
A clear roadmap for experimenting with and deploying advanced coding LLMs using NVIDIA’s technology stack
Industry: All Industries
Topic: Agentic AI / Generative AI - Code / Software Generation
Technical Level: Technical - Intermediate
Intended Audience: Data Scientist
NVIDIA Technology: DGX Platform, TensorRT, NeMo, Triton, NVIDIA NIM, NVIDIA AI Enterprise, DGX Cloud, Nemotron
#nvidiagtc










