NVIDIA Logo

NVIDIA

Senior System Software Engineer, Software Defined Networking

Posted 9 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Design, develop, operate, and optimize large-scale software-defined networking solutions for NVIDIA AI Cloud environments. Responsibilities include building OVS/OVN control and data plane software, Kubernetes networking services, observability systems, CI/CD pipelines, GitOps integrations, and secure gRPC/REST services. The role owns production reliability, incident response, performance tuning, and upstream open-source contributions while collaborating with SRE, DevOps, and network engineering teams.
The summary above was generated by AI

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response.
 

What you'll be doing:

  • Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow)

  • Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes

  • Drive upstream contributions to OVN-Kubernetes and related open-source projects; Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis

  • Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments

  • Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs

  • Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs

  • Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure; Drive reliability through incident management, resource monitoring, and performance tuning

  • Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational tooling

What we need to see:

  • BS/MS in Computer Science or related technical field, or a comparable blend of education and relevant experience

  • 5+ years of proven experience in software development for large-scale distributed environments

  • Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols

  • Strong programming skills in C and Go; advanced scripting in Bash and Python

  • Deep knowledge of Kubernetes, practical experience deploying and supporting CNIs (OVN-Kubernetes)

  • Hands-on experience with Infrastructure-as-Code and deployment tools (Ansible, Terraform, ArgoCD, Flux)

  • Experience designing and operating complex, multi-stage CI/CD pipelines

  • Hands-on experience developing secure, high-performance services using gRPC and REST with TLS and strong authentication

  • Strong knowledge of datacenter routing, switching, and Linux host/VM networking

Ways to stand out from the crowd:

  • Contributions to open-source projects (especially OVS, OVN, OVN-Kubernetes, or other Kubernetes networking projects)

  • Experience with hardware acceleration (GPU, DPU or equivalent experience) for networking

  • Practical experience with major cloud providers (AWS, Azure, GCP) and hybrid/multi-cloud deployments

  • SRE/DevOps top-level expertise — on-call, incident management, operations focused on service reliability targets, production ownership

  • Experience with observability platforms and tools (Prometheus, Grafana, Jaeger, OpenTelemetry, ELK)

Similar Jobs

31 Minutes Ago
Remote
Gujarat, IND
Senior level
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Lead technicians to install, maintain, calibrate, and troubleshoot semiconductor lab and manufacturing equipment. Ensure ISO/IEC 17025 calibration compliance, manage equipment tracking, drive RCA/RCCA and preventive actions, coordinate vendors and cross-functional teams, support audits, and provide training while maintaining safety, ESD, and cleanroom standards.
Top Skills: 3D X-RayArtificial IntelligenceBend TesterCleanroomCsamEquipment Tracking SystemEsdFibFtirHast ChamberIso/Iec 17025KaizenLean Six SigmaLinuxReflow OvenSemShock TesterSoak ChamberTemp CycleTemperature ChamberTesterThbWindowsX-Section
31 Minutes Ago
Remote
Gujarat, IND
Senior level
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Lead operation, maintenance, troubleshooting, and continuous improvement of facility gas generation, storage, and distribution systems (LN2, N2, CO2, O2, argon, helium, hydrogen, CDA). Ensure gas purity, availability, safety compliance (EHS, LOTO, PTW), perform preventive and corrective maintenance, support commissioning/expansion, maintain documentation, and use PLC/BMS/SCADA and CMMS/SAP tools. Proficiency with GenAI/Copilot tools is expected.
Top Skills: Agent BotsBmsCmmsCo2 Supply SystemsCopilotCopilot LibraryCryogenic Tank & Vaporizer SystemsGas Cabinets & VmbsGas Detection SystemsGen Ai ToolsNitrogen Generation SystemsOxygen Distribution SystemsP&IdPlcSAPScadaSpecialty Gas Panels & Manifolds
An Hour Ago
Remote or Hybrid
India
Entry level
Entry level
Financial Services
This is a test requisition with no actual job responsibilities, qualifications, technologies, compensation details, or application requirements provided.

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account