Jobs›System Software

Senior System Software Engineer, Software Defined Networking

NVIDIA · Work from home
Work from home
PayPay not listed
WhereWork from home
TypeFull timeSenior
Posted13 Sep23 days ago, via Himalayas
Skills they list12 named
Software Defined NetworkingSystems Software EngineeringCloud Infrastructure EngineeringNetwork EngineerSRESenior Network Software EngineerSenior Systems Software EngineerSenior Network Systems EngineerSenior Infrastructure Software EngineerSenior Network Infrastructure EngineerSenior Network EngineerSenior IP Network Engineer
About this job

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response.

What you'll be doing

Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow)

Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes

Drive upstream contributions to OVN-Kubernetes and related open-source projects; Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis

Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments

Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs

Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs

Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure; Drive reliability through incident management, resource monitoring, and performance tuning

Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational tooling

What we need to see

BS/MS in Computer Science or related technical field, or a comparable blend of education and relevant experience

5+ years of proven experience in software development for large-scale distributed environments

Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols

Strong programming skills in C and Go; advanced scripting in Bash and Python

Deep knowledge of Kubernetes, practical experience deploying and supporting CNIs (OVN-Kubernetes)

Hands-on experience with Infrastructure-as-Code and deployment tools (Ansible, Terraform, ArgoCD, Flux)

Experience designing and operating complex, multi-stage CI/CD pipelines

Hands-on experience developing secure, high-performance services using gRPC and REST with TLS and strong authentication

Strong knowledge of datacenter routing, switching, and Linux host/VM networking

Ways to stand out from the crowd

Contributions to open-source projects (especially OVS, OVN, OVN-Kubernetes, or other Kubernetes networking projects)

Experience with hardware acceleration (GPU, DPU or equivalent experience) for networking

Practical experience with major cloud providers (AWS, Azure, GCP) and hybrid/multi-cloud deployments

SRE/DevOps top-level expertise — on-call, incident management, operations focused on service reliability targets, production ownership

Experience with observability platforms and tools (Prometheus, Grafana, Jaeger, OpenTelemetry, ELK)

Originally posted on Himalayas

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on Himalayas
Opens himalayas.app in a new tab