NVIDIA
Principal Networking AI Systems Architect
NVIDIA logo NVIDIA 2396 open jobs

Principal Networking AI Systems Architect

Tel Aviv, Israel · Senior (5+ years) · Full Time ·

We are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.

What you'll be doing:

  • Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning.
  • Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML.
  • Drive the integration of AI capabilities into system architecture and engineering workflows.
  • Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
  • Translate system behavior, dependencies, data, and operational constraints into formulated research problems.

What we need to see:

  • Ph.D in electrical engineering, machine-learning, computer-science or another relevant field.
  • 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
  • Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures.
  • Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field.
  • Deep knowledge of AI/ML and networking-hardware/system-architecture.
  • Excellent ability to convey and communicate data-based insights to stakeholders and management.
  • Experience demonstrating an excellent track of collaboration with hands-on teams.

Ways to stand out from the crowd:

  • Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments.
  • Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms.
  • Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization.
  • High energy and a positive, proactive and curious approach.