Research Engineer - [Seed Model - Infra - Training Performance & ML Compilation (Torch Compile)]
ByteDance
- ๐บ๐ธ United States
- On-site
- 1 month ago
- AI
- PyTorch
- C++
- Python
- CUDA
- Triton
1 month ago
The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.
Responsibilities
- Optimize training performance for large-scale foundation models through compiler-level techniques, including graph optimization, operator fusion, and kernel generation.
- Develop and extend ML compilation capabilities based on the PyTorch compilation stack (e.g. FX, Dynamo, Inductor) to improve training efficiency across heterogeneous GPU platforms.
- Design and optimize high-performance GPU kernels for training workloads.
- Conduct performance profiling and analysis of large-scale training jobs; identify and resolve bottlenecks in collaboration with research and infrastructure teams.
Minimum Qualification(s)
- Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.
- Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.
- Hands-on experience in training-side performance optimization for deep learning workloads.
- Hands-on experience writing and optimizing GPU kernels (e.g., CUDA, Triton).
- Experience with the PyTorch compilation stack, meeting at least one of the following: Direct experience using, debugging, or extending Inductor or FX.
- Proficiency in Triton kernel development.
- Solid experience with PyTorch computation graph work (graph optimization, graph capture, operator fusion).
Preferred Qualification(s)
- Experience with TorchDynamo or bytecode-level program transformation.
- Experience with Triton compiler internals or other ML compiler backends (e.g., MLIR, LLVM).
- Contributions to related open-source projects (e.g., PyTorch, Triton, FlashAttention).
- Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).
Research Engineer - [Seed Model - Infra - Training Performance & ML Compilation (Torch Compile)] ยท ByteDance