About this role
We’re looking for an experienced HPC infrastructure engineer to lead operations on what is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.
You may be a good fit if:
You love anime and the anime aesthetic.
This probably the only lab in the world where you will get to combine your expertise on anime and HPC systems training state-of-the-art anime models.
You’re familiar with the modern HPC software landscape…
Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)
…and hardware landscape too:
We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.
You're comfortable working on small, fast-paced teams
We currently have a very tiny research team, and you’ll be working alongside some of the best AI researchers in the world, on the very best anime image model in the world.
We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Visa sponsorships are available.