NVIDIA
2396 open jobs
View all jobs →
Principal Networking AI Systems Architect
We are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.
What you'll be doing:
- Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning.
- Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML.
- Drive the integration of AI capabilities into system architecture and engineering workflows.
- Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
- Translate system behavior, dependencies, data, and operational constraints into formulated research problems.
What we need to see:
- Ph.D in electrical engineering, machine-learning, computer-science or another relevant field.
- 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
- Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures.
- Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field.
- Deep knowledge of AI/ML and networking-hardware/system-architecture.
- Excellent ability to convey and communicate data-based insights to stakeholders and management.
- Experience demonstrating an excellent track of collaboration with hands-on teams.
Ways to stand out from the crowd:
- Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments.
- Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms.
- Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization.
- High energy and a positive, proactive and curious approach.