About this role
About the Team
The ByteDance infrastructure storage team provides storage services for all ByteDance products and builds the Volcano Engine public cloud storage product portfolio. In the AI era, the team powers Doubao large models and the Ark large model platform. With the enterprise service experience of stable storage for hundreds of EB of data on tens of thousands of servers, the team serves external customers through a rich storage product matrix, spanning pooled storage, object storage, block storage, file storage, log service, message queue, and big data storage (HDFS). The Infrastructure Storage US team is at the forefront of the team's global expansion, responsible for storage system development, operations, and stability for global businesses and customers.
We are looking for talented engineers to join the Infrastructure Storage US team. In this role, you will design, develop, and optimize the foundational infrastructure of our large-scale distributed storage systems (object storage, block storage, file storage, pooled storage, HDFS, etc.), improve system stability, performance, and cost efficiency, and ensure reliable support for business applications.
Responsibilities
- Design, develop, and optimize the modules of distributed storage systems (pooled storage, big data file storage, message queue, cache storage, object storage, block storage, file storage, and log service), and research cutting-edge storage technologies;
- Design and enhance modules of the distributed storage system in terms of stability, functionality, performance, and cost based on business requirements;
- Drive project execution according to project schedules, write detailed design documents, and own module implementation, performance tuning, and functional testing;
- Provide timely technical support for online applications, identify potential requirements and optimization opportunities from production issues, and continuously optimize the system;
- Participate in the deployment, on-call operations, and site reliability engineering for the storage US team; develop deep expertise in system design, identify and resolve bottlenecks to enhance cluster stability, elasticity, and cost competitiveness.
Minimum Qualifications
- Bachelor's degree or above in Computer Science or a related field;
- Strong coding skills in C++, with a strong commitment to code quality;
- Solid understanding of distributed systems principles; Hands-on experience with open-source distributed storage systems such as HDFS, Ceph, and Kafka, or equivalent in-house distributed storage systems;
- Able to think independently, proactively identify problems, and possess systematic problem analysis and solving capabilities;
- Strong aptitude for quickly ramping up in unfamiliar technical domains, along with excellent teamwork and communication skills.
Preferred Qualifications
- Experience with storage tech stacks, including NVMe, SPDK/DPDK, RocksDB, and metadata-related technologies;
- Hands-on experience designing and building large-scale distributed systems, including high IOPS, high throughput, stability, and degraded-mode handling;
- Experience with AI storage, such as KV cache storage and optimization, multi-tier caching, and model distribution and loading.
- Experience with Linux performance tuning is preferred;