Platform Engineer

VerdaSanta Clara, Palo Alto, CaliforniaOn-siteFull-timeMid level, 2–5 yearsListed 7 hours ago

Apply now

About this role

At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.

We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join Verda while it’s still being built - not once it’s finished.

Why Verda

- Cash and equity compensation along with local benefits.
- 40+ nationalities, with 6 different ones on the management team.
- A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.

Practicalities

- Work mode: Remote
- Level: Mid / Senior
- Employment type: Full time and permanent

Your responsibilities

- Build, operate, and improve self-hosted Kubernetes platforms across deployment, automation, scaling, upgrades, and production operations
- Maintain platform reliability through on-call participation, incident response, and troubleshooting
- Manage Kubernetes infrastructure including networking, observability, storage, and platform services
- Develop infrastructure automation using Ansible, GitOps, and CI/CD workflows
- Collaborate with engineering teams to improve developer experience, deployment processes, and operational efficiency
- Operate and troubleshoot Linux systems, container platforms, and distributed infrastructure
- Implement Kubernetes networking and security best practices, including cluster hardening and access controls
- Leverage AI-assisted engineering tools to improve automation, operations, and troubleshooting
- Contribute to platform standards, documentation, and operational best practices

Your key competencies

- 3-7 years operating production Kubernetes environments
- Strong experience managing self-hosted Kubernetes clusters end-to-end
- Hands-on experience with Kubernetes operations, upgrades, scaling, and troubleshooting
- Experience with on-call rotations and production incident management
- Strong background in infrastructure automation (Ansible), GitOps, and CI/CD
- Solid Linux administration and debugging skills
- Experience with containerized and cloud-native infrastructure
- Understanding of Kubernetes networking, CNI plugins, and platform security
- Scripting or programming experience (Python, Bash, Go, or similar)
- Comfortable using AI-powered tooling to improve engineering workflows
- Strong collaboration and communication skills

Nice to have

- Experience with Rancher and Cilium
- Familiarity with Kubernetes security policies and cluster hardening
- Experience with GPU workloads, AI/ML infrastructure, or NVIDIA container runtimes
- Experience with observability tooling such as Prometheus, Grafana, and Loki
- Familiarity with Infrastructure as Code and Kubernetes ecosystem tools (ArgoCD, Helm, Kustomize, operators)
- Knowledge of distributed systems, storage, and high-availability environments

What's next

We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.

Please submit your application through our Careers page. We don't accept applications sent by email.