About this role
About the Role
This is an early-team platform engineering role at a seed-stage AI infrastructure startup building an open-source, GitOps-native distributed operating system on top of Kubernetes. You will work closely with the founding team to design and evolve core platform features, making multi-cluster management intuitive for developers running demanding AI workloads. The work spans the full infrastructure stack, from bare metal to production-grade Kubernetes.
What You'll Do
- Build and extend core platform features in Go, including custom Kubernetes operators and controllers.
- Design and implement GitOps workflows with ArgoCD to make continuous deployment seamless.
- Develop infrastructure-as-code patterns using Terraform and Helm to provision and manage clusters.
- Work on distributed storage solutions using Ceph and WEKA for high-performance, scalable cluster storage.
- Create observability and monitoring systems with Prometheus and Grafana to surface cluster health and performance.
- Build and optimize container networking with Cilium for network security and observability.
- Design and implement federated Kubernetes architectures for multi-cluster management.
- Build automation tooling that reduces operational overhead for developers running production workloads.
What We're Looking For
- 2 to 5+ years of software development experience, with a strong systems or infrastructure focus.
- Hands-on proficiency in Go, including writing Go in a Kubernetes environment.
- Production Kubernetes experience: managing clusters at meaningful scale, writing operators and controllers, working with CRDs.
- Experience with distributed storage solutions such as Ceph or WEKA.
- Experience designing federated Kubernetes architectures for multi-cluster management.
- Familiarity with GitOps workflows and ArgoCD.
- Experience with infrastructure-as-code tools such as Terraform and Helm; Ansible or Kubespray is a plus.
- Background with container networking, CNI plugins, or service mesh technologies such as Istio or Linkerd.
- Experience with cloud platforms (AWS, GCP, or Azure) and their managed Kubernetes offerings.
- Familiarity with observability tooling such as Prometheus and Grafana.
- Experience with GPU infrastructure or bare metal environments is a strong plus.
- Contributions to Kubernetes ecosystem tooling or similar open-source infrastructure projects are a bonus.
Compensation & Benefits
Salary range: $180,000 to $210,000 USD annually. Visa sponsorship is available.
Location
On-site in San Francisco, California . Candidates must be able to work in person full-time.