Senior Storage Engineer

VerdaSanta Clara, Palo Alto, CaliforniaOn-siteFull-timeSenior, 5–8 yearsListed 7 hours ago

Apply now

About this role

At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.

We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join Verda while it’s still being built - not once it’s finished.

Why Verda

- Cash and equity compensation along with local benefits.
- 40+ nationalities, with 6 different ones on the management team.
- A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.

Practicalities

- Work mode: Remote
- Level: Mid / Senior
- Employment type: Full time and permanent

About the role

The company is constructing a European cloud computing platform powered by renewable energy. As Senior Storage Engineer, you would lead the architecture and operations of Ceph clusters handling petabyte-scale customer data, serving demanding GPU workloads across object, block, and file storage systems.

Your responsibilities

- Direct design, deployment, and operation of large-scale production Ceph clusters
- Establish technical direction for storage infrastructure including architecture and operational standards
- Launch a managed Object Storage product with access keys, bucket management, and SLOs
- Scale storage systems to petabytes or hundreds of petabytes
- Manage Ceph across CephFS, RBD, and RADOS Gateway interfaces
- Mentor engineers through code review, design collaboration, and shared on-call responsibilities
- Enhance observability, automation, and operational tooling
- Diagnose complex performance issues in distributed storage environments
- Collaborate with compute and networking teams on GPU cluster integration
- Oversee capacity planning, infrastructure upgrades, and lifecycle management
- Engage in production operations and incident response leadership

Your key competencies

- Extensive Ceph expertise including deployment, operations, troubleshooting, and performance optimization
- Demonstrated experience operating Ceph at multi-petabyte production scale
- Proven technical leadership managing storage infrastructure end-to-end
- Engineering mentorship experience
- Advanced Linux systems/DevOps skills, baremetal and internals knowledge
- Solid networking expertise with distributed systems debugging capability
- Mission-critical production infrastructure operations background
- Infrastructure automation development experience
- Strong collaboration and cross-team communication abilities

Nice to have

- RADOS Gateway production experience as managed Object Storage
- CephFS production environment familiarity
- Ansible proficiency
- Expertise with monitoring tools like Prometheus, Grafana, Loki, or similar
- Cloud provider storage operations background
- Experience with 10-100 PB scale storage systems
- Understanding of GPU/AI workload storage patterns

What's next

We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.

Please submit your application through our Careers page. We don't accept applications sent by email.