Senior Systems Software Engineer - NV Cloud Functions

JobgetherIndiaOn-siteFull-timeSenior, 5–8 yearsListed 1 day ago

Apply now

About this role

Accountabilities:

- Design, develop, and ship scalable services using Java, Go, and Rust for a distributed GPU workload platform.

- Improve the performance, reliability, scalability, and operational behavior of systems that route AI workloads across distributed GPU fleets.

- Build and optimize cloud-native services capable of supporting inference, streaming, and batch workloads across diverse infrastructure environments.

- Develop and maintain control-plane and edge components for a distributed, open-source platform.

- Automate and optimize build, testing, integration, deployment, and release processes for cloud-native software.

- Work with Kubernetes and container technologies to develop reliable and scalable workload orchestration capabilities.

- Collaborate with engineering teams across the organization to integrate the platform with related scheduling, inference, orchestration, and AI infrastructure technologies.

- Investigate complex systems and distributed infrastructure challenges, balancing performance, security, reliability, and scalability requirements.

- Contribute to a public open-source repository through code, design proposals, technical reviews, documentation, and issue resolution.

- Triage community issues and pull requests while helping maintain a high-quality contributor experience.

- Explore new approaches for making GPU- and DPU-accelerated applications easier to develop, deploy, monitor, and operate.

- Contribute to continuous improvements in engineering processes, platform architecture, and production reliability.

Requirements:

- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related discipline, or equivalent practical experience.

- 3+ years of hands-on professional software engineering experience.

- Expert-level proficiency in at least one systems programming language, such as Go, C, or Rust.

- Strong understanding of data structures, algorithms, distributed software architecture, and systems engineering principles.

- Hands-on experience with Kubernetes, container orchestration, and container technologies.

- Experience automating software delivery and infrastructure workflows using continuous integration and deployment frameworks such as GitLab and ArgoCD.

- Strong scripting skills in Bash, Python, or a comparable scripting language.

- Solid understanding of Linux or other Unix-like operating systems and familiarity with system and kernel internals.

- Strong understanding of performance, security, reliability, and scalability considerations in complex distributed systems.

- Ability to work effectively in a distributed engineering environment with shifting priorities and technically complex challenges.

- Strong analytical and problem-solving skills, with the ability to investigate issues and translate technical findings into practical solutions.

- Clear technical communication skills and an interest in contributing to open-source projects.

- Experience with publish-subscribe architectures and message queues is an advantage.

- Experience optimizing high-throughput network paths and familiarity with unary, streaming, and bidirectional communication over HTTP/2 and gRPC is preferred.

- Experience developing Kubernetes Custom Resources and Operators for deployment in cloud service provider environments is a plus.

Benefits:

- Competitive salary aligned with experience, skills, and the local market.

- Generous benefits package designed to support employees' professional and personal needs.

- Opportunity to work on advanced AI infrastructure and distributed GPU computing technologies.

- Exposure to large-scale cloud-native systems, Kubernetes, systems programming, and high-performance computing.

- Opportunity to contribute to open-source software and collaborate with a global technical community.

- Work alongside experienced engineers on complex infrastructure and performance challenges.

- Distributed and collaborative work environment with opportunities for continuous technical learning and growth.

- Exposure to emerging AI, GPU, DPU, and cloud technologies.

- Opportunities to contribute to projects with broad impact across AI and accelerated computing.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1