Research Scientist Graduate (DPU & AI Infra) - 2027 Start (PhD)
- ๐บ๐ธ United States
- On-site
- Entry level
- 1 month ago
- AI
- AI/ML
- C++
- Rust
- Linux
- eBPF
- FPGA
- CUDA
About the Team
The ByteDance DPU (Data Processing Unit) team builds foundational cloud and AI computing infrastructure for ByteDance and Volcano Engine. Our mission is to advance the architecture, development, and research of next-generation software-hardware co-design technologies across compute, networking, and storage for cloud and AI computing. Our technology stack spans
- Cloud virtualization, hypervisors, and operating systems
- High-performance networking, including DPDK and RDMA
- High-speed interconnects, virtual switching, and network offload
- Distributed storage and I/O acceleration
- Orchestration and scheduling for AI/ML workloads
We work at the intersection of systems research, distributed infrastructure, and hardware acceleration. Our technologies operate at cloud scale and help shape the next generation of cloud and AI computing platforms.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
- Conduct research and development in DPU-based cloud acceleration and large-scale AI infrastructure.
- Explore software-hardware co-design opportunities for AI/ML infrastructure, leveraging DPUs, GPUs, and custom hardware to optimize distributed training and inference.
- Develop new techniques for accelerating distributed AI training and inference, including communication, data movement, resource management, and memory or cache systems.
- Perform end-to-end performance analysis and optimization across hardware, device drivers, operating-system kernels, communication libraries, and user-space runtimes.
- Build prototypes and evaluate proposed designs using representative cloud and AI workloads at scale.
- Contribute to architecture design, technical proposals, and long-term research directions.
Minimum Qualifications
- Ph.D. in related fields with research training and publications.
- Proficiency in C/C++ or Rust, including systems-level development and debugging.
- Familiar with Linux systems development experience.
- Solid understanding of operating systems, computer architecture, networking, or distributed systems.
- Background in at least one of: software-hardware co-design, computer architecture, distributed storage systems, high-performance networking, or AI/ML systems.
Preferred Qualifications
- Experience with software-hardware co-design (networking, storage, or distributed compute).
- Hands-on experience with network virtualization (OVS, SR-IOV, eBPF).
- Familiarity with DPDK and high-performance user-space networking.
- Bonus points for hardware acceleration experience, FPGA/ASIC/GPU/CUDA
- Bonus points for experience with NCCL Collectives along with AI communication patterns and parallelization techniques
- Proven experience designing and building AI/ML infrastructure related but not limited to inference kv cache system, data preprocessing system.
Research Scientist Graduate (DPU & AI Infra) - 2027 Start (PhD) ยท ByteDance