Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ByteDance logo

Research Engineer - [Seed Model - Infra - LLM/VLM Inference Optimization (Kernel & Compiler)]

ByteDance
  • ๐Ÿ‡บ๐Ÿ‡ธ United States
  • On-site
  • 6 months ago
  • AI
  • CUDA
  • Triton
  • C++
  • Python
  • vLLM
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

Responsibilities

  • Design, implement, and optimize high-performance GPU kernels for large-scale LLM/VLM inference workloads, including attention, GEMM, and other compute- and memory-intensive operators.
  • Develop and tune inference kernels in CUDA and Triton, and drive end-to-end performance optimization of production inference systems at scale.
  • Conduct in-depth performance analysis and profiling to identify bottlenecks across the inference stack, from kernel level to serving level.
  • Collaborate with research and infrastructure teams to land kernel- and compiler-level optimizations in production inference systems.

Minimum Qualifications

  • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.
  • Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.
  • Hands-on experience in LLM/VLM inference optimization with demonstrated impact on latency, throughput, or serving cost.
  • Hands-on experience writing and optimizing GPU kernels in CUDA and/or Triton.
  • Deep understanding of GPU architecture (memory hierarchy, occupancy, instruction throughput) with solid optimization experience.

Preferred Qualifications

  • Experience with ML compiler internals (e.g., Triton, MLIR, LLVM).
  • Contributions to related open-source projects (e.g., Triton, vLLM, SGLang, FlashAttention, CUTLASS).
  • Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).

Research Engineer - [Seed Model - Infra - LLM/VLM Inference Optimization (Kernel & Compiler)] ยท ByteDance

Auto apply with Likeremote