About this role
You will take on the following responsibilities:
- Design and implement GPU-accelerated kernels for financial computation workloads
- Optimize GPU code for throughput, latency, and memory efficiency across current and next-generation hardware (Blackwell, Rubin)
- Develop procedures for precision management (FP8/FP4 training and inference) in financial applications
- Profile and optimize GPU workloads using NVIDIA tooling (Nsight Systems, Nsight Compute)
- Build reusable GPU libraries and abstractions that modeling teams can use without requiring deep CUDA expertise
- Evaluate and integrate GPU-accelerated libraries (RAPIDS, CUTLASS, cuBLAS, TensorRT) for financial use cases
You should possess the following qualifications:
- BS or MS in Science, Technology, Engineering or Math
- Minimum 1 year of experience required; 4-10 years of experience preferred
- Expert-level CUDA programming: kernel development, memory management, stream and graph optimization
- Deep understanding of GPU architecture: SM structure, warp scheduling, memory hierarchy (registers, shared memory, L1/L2, HBM)
- Experience with performance profiling and optimization of GPU workloads
- Strong C++ and Python skills, as well as familiarity with mixed-precision computation and numerical stability
- Track record of delivering meaningful speedups on real workloads (not just benchmarks)
Preferred experience:
- Background in HPC, scientific computing, or computational finance
- Experience with multi-GPU and multi-node GPU programming (NCCL, MPI)
- Familiarity with GPU-accelerated data processing frameworks (RAPIDS, cuDF)