NVIDIA Logo

NVIDIA

Senior Systems Software Engineer - NV Cloud Functions

Posted 3 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Design and ship Java, Go, and Rust services for NVIDIA Cloud Functions, a distributed platform routing AI workloads across GPU fleets. Improve performance, reliability, scalability, and cloud-native build and release processes. Collaborate across NVIDIA technologies, contribute to an open-source project, triage community issues, review pull requests, and write documentation. The role requires expertise in systems programming, distributed architecture, Kubernetes, containerization, Linux internals, scripting, and continuous integration.
The summary above was generated by AI

Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA Cloud Functions (NVCF) is an open-source platform that links workloads to GPUs. It lets teams deploy, manage, and serve GPU-accelerated, containerized applications across regions and clusters worldwide. The platform routes inference, streaming, and batch jobs across decentralized GPU clusters. This allows endpoints to scale repeatably, whether hosted on-prem or in the cloud.

We are seeking a Senior Systems Software Engineer to join our team. You will focus on improving the performance, reliability, and scaling behavior of a system that routes AI workloads onto distributed GPU fleets. You will work on a polyglot platform that is now fully open source, with both control plane and edge deployments. The work suits someone with deep experience in systems performance, distributed systems, and Kubernetes-based runtimes. We are looking for engineers who want to learn and grow. Expect to be challenged in an environment with rapidly shifting priorities, where insight, focus, and execution are key.

What you will be doing:

  • You'll be working in a distributed team that explores innovative ways to make GPU- and DPU-accelerated applications easier to develop, deploy, and monitor on the latest and greatest NVIDIA hardware.

  • Design and ship services in Java, Go, and Rust, building in the open on a public repository where your commits, design proposals, and reviews are transparent to the community.

  • Work on automating and optimizing build, test, integration, and release processes for cloud native.

  • Partner with engineering teams across NVIDIA so the platform integrates with adjacent NVIDIA technologies, including the KAI Scheduler, NVIDIA NIM, Grove, and Dynamo.

  • Help steward an open-source project. You will triage community issues and pull requests and write docs contributors can build on.

What we need to see:
  • Bachelor’s or Master’s Degree in Computer Science or equivalent experience

  • 3+ years of hands-on software engineering.

  • Expert-level knowledge in a systems programming language (Go, C, Rust) and a proven understanding of Data Structures, Algorithms, and Distributed Software Architecture

  • Strong understanding of container orchestration systems (Kubernetes) and container technologies with hands-on automation experience in continuous integration frameworks like GitLab & ArgoCD.

  • Expertise in a scripting language (Bash, Python) and knowledge and experience working with the system internals of Unix/Unix-like kernels such as Linux.

  • Understanding of performance, security, and reliability in complex distributed systems.

Ways to stand out from the crowd:
  • Background with pub-sub models and message queues

  • Experience optimizing for high-throughput network paths, with a working understanding of unary versus streaming and bidirectional protocols across HTTP/2 and gRPC.

  • Experience with developing Kubernetes Custom Resources and Operators deployed in Cloud Service Providers

With competitive salaries and a generous benefits package, NVIDIA is considered one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the industry working for us. Due to unprecedented growth, our exclusive engineering teams are expanding rapidly. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!

Similar Jobs

3 Hours Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads complex, cross-functional scientific learning projects for Medical Affairs. Develops curricula, e-learning, and blended learning programs; partners with medical, scientific, agency, and learning systems teams; manages multiple projects, timelines, budgets, and priorities; evaluates emerging e-learning technologies; and updates existing learning resources across therapeutic areas.
Top Skills: Articulate StorylineDigital Learning TechnologyE-Learning Authoring ToolsLearning Management Systems
5 Hours Ago
Remote
Gujarat, IND
Senior level
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Lead technicians to install, maintain, calibrate, and troubleshoot semiconductor lab and manufacturing equipment. Ensure ISO/IEC 17025 calibration compliance, manage equipment tracking, drive RCA/RCCA and preventive actions, coordinate vendors and cross-functional teams, support audits, and provide training while maintaining safety, ESD, and cleanroom standards.
Top Skills: 3D X-RayArtificial IntelligenceBend TesterCleanroomCsamEquipment Tracking SystemEsdFibFtirHast ChamberIso/Iec 17025KaizenLean Six SigmaLinuxReflow OvenSemShock TesterSoak ChamberTemp CycleTemperature ChamberTesterThbWindowsX-Section
5 Hours Ago
Remote
Gujarat, IND
Senior level
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Lead operation, maintenance, troubleshooting, and continuous improvement of facility gas generation, storage, and distribution systems (LN2, N2, CO2, O2, argon, helium, hydrogen, CDA). Ensure gas purity, availability, safety compliance (EHS, LOTO, PTW), perform preventive and corrective maintenance, support commissioning/expansion, maintain documentation, and use PLC/BMS/SCADA and CMMS/SAP tools. Proficiency with GenAI/Copilot tools is expected.
Top Skills: Agent BotsBmsCmmsCo2 Supply SystemsCopilotCopilot LibraryCryogenic Tank & Vaporizer SystemsGas Cabinets & VmbsGas Detection SystemsGen Ai ToolsNitrogen Generation SystemsOxygen Distribution SystemsP&IdPlcSAPScadaSpecialty Gas Panels & Manifolds

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account