We are seeking a Distinguished Engineer to lead AI Resiliency at NVIDIA!
Join NVIDIA and help push the boundaries of AI. In this role, you will architect, design, and develop world-class software resiliency features for training ground breaking AI models on the largest AI superclusters in the world. Leading a team of cross-functional experts, you will drive and shape our end-to-end AI software stack, ensuring seamless training of frontier models on industry-leading frameworks like PyTorch and JAX/XLA, with near-zero downtime. Your optimizations will span from algorithmic innovations to robust software architecture, with a significant impact on NVIDIA’s most critical customers. This highly visible role demands exceptional technical expertise and leadership across organizations, with direct exposure to NVIDIA's senior leadership.
What You'll Be Doing:
Define a scalable software architecture to enable single-job resilient training on hundreds of thousands of GPUs with minimal downtime.
Design and deliver modular, resilient software features to support large-scale AI training for our top customers.
Innovate and evolve resilient architecture designs to achieve stringent uptime requirements (downtime < 1%), through solutions like in-memory check-pointing, in-process restart, and anomaly/SDC detection.
Collaborate closely with internal partners, spearheading successful project execution and communicating regular progress updates to senior leadership.
What We Need to See:A Master’s or Ph.D. in Computer Science, Electrical or Computer Engineering from a top-tier university, or equivalent experience.
15+ years of experience in software architecture or related fields, with a deep understanding of AI-optimized systems.
Excellent and proven ability to collaborate and communicate effectively across multiple engineering teams.
At least 5 years of hands-on experience in software development on high-complexity projects involving HPC or AI.
Ways to Stand Out from the Crowd:Proven experience with large-scale AI supercomputing applications, particularly in the training phase.
5+ years of experience with using and contributing to modern AI frameworks like PyTorch and JAX/XLA, specifically for large-scale training workloads.
A strong passion for designing system architectures tailored for AI, covering CPU, GPU, memory, storage, and networking.
Hands-on involvement in the entire lifecycle—from design to deployment—of large-scale High-Performance Computing (HPC) systems.
Experience in implementing HPC software development best practices in large-scale systems.
NVIDIA continues to expand its presence in the Datacenter space, and our team plays a pivotal role in enhancing the value of our rapidly growing datacenter deployments. We also drive a data-driven approach to hardware design and system software development. You will collaborate with a wide array of teams across NVIDIA, including deep learning research, CUDA kernel and framework development, and silicon architecture. NVIDIA is widely recognized as one of the technology industry’s most desirable employers, with some of the most dedicated and innovative minds working with us. If you’re creative, driven, and autonomous, we want to hear from you!
The base salary range is 308,000 USD - 471,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.