LLM Engineer (Optimization)
- AI
- vLLM
- TensorRT
- ONNX
- CUDA
- Triton
- Machine Learning
- PyTorch
- Python
- C++
- Kubernetes
- AI/ML
About the Team & Mission
ยLLM Engineer (Optimization)๋ ๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ(LLM)์ ์ถ๋ก ์ฑ๋ฅ(Inference Performance)์ ๊ทน๋ํํ์ฌ ์ค์ ์๋น์ค ํ๊ฒฝ์ ์ต์ ํ๋ AI ์์คํ ์ ๊ฐ๋ฐํฉ๋๋ค.
์๋ฒ(GPU Cluster)๋ถํฐ Edge ๋ฐ On-device ํ๊ฒฝ๊น์ง ๋ค์ํ ํ๋์จ์ด์์ ์ต๊ณ ์ ์ฑ๋ฅ๊ณผ ํจ์จ์ ๋ฌ์ฑํ ์ ์๋๋ก Inference Engine, Runtime, Compiler ๋ฐ Model Optimization ๊ธฐ์ ์ ์ฐ๊ตฌยท๊ฐ๋ฐํฉ๋๋ค. ์ต์ LLM Serving ๊ธฐ์ ๊ณผ GPU/Accelerator ์ต์ ํ๋ฅผ ํ์ฉํ์ฌ ๊ณ ์ฑ๋ฅยท์ ์ง์ฐยท์ ๋น์ฉ AI ์๋น์ค๋ฅผ ๊ตฌํํ๋ ํต์ฌ ์ญํ ์ ์ํํฉ๋๋ค.
ยResponsibilities
LLM Inference Optimization
๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ(LLM)์ ์ถ๋ก ์ฑ๋ฅ(Latency, Throughput, Memory Efficiency)์ ์ต์ ํํฉ๋๋ค.
๋ค์ํ ๋ชจ๋ธ ๊ตฌ์กฐ ๋ฐ ์ถ๋ก ํ๊ฒฝ์ ๋ง๋ ์ต์ ํ ๊ธฐ๋ฒ์ ์ฐ๊ตฌํ๊ณ ์ ์ฉํฉ๋๋ค.
Long Context, Multi-turn Conversation ๋ฑ ์ค์ ์๋น์ค ํ๊ฒฝ์์์ ์ฑ๋ฅ์ ๊ฐ์ ํฉ๋๋ค.
Inference Engine ๋ฐ Runtime ๊ฐ๋ฐ
GPU ๋ฐ Accelerator ๊ธฐ๋ฐ LLM Inference Engine์ ๊ฐ๋ฐํ๊ณ ์ต์ ํํฉ๋๋ค.
vLLM, TensorRT-LLM, SGLang, llama.cpp, ONNX Runtime, MLX ๋ฑ ์ต์ Inference Framework๋ฅผ ํ์ฉํ๊ฑฐ๋ ๊ฐ์ ํฉ๋๋ค.
Speculative Decoding, Prefill-Decode Disaggregation ๋ฑ ์ต์ Serving ๊ธฐ์ ์ ์ ์ฉํฉ๋๋ค.
Model Compression ๋ฐ Compiler Optimization
Quantization(MXFP8, NVFP4, AWQ, GPTQ ๋ฑ) Pruning, Distillation ๋ฑ ๋ชจ๋ธ ๊ฒฝ๋ํ ๊ธฐ๋ฒ์ ์ฐ๊ตฌํ๊ณ ์ ์ฉํฉ๋๋ค.
CUDA, Triton, TensorRT, TVM, MLIR ๋ฑ Compiler ๋ฐ Kernel Optimization์ ํ์ฉํ์ฌ ์ถ๋ก ์ฑ๋ฅ์ ํฅ์ํฉ๋๋ค.
GPU Memory ๋ฐ Kernel ํจ์จ์ ๋ถ์ํ๊ณ ์ต์ ํํฉ๋๋ค.
Edge AI ๋ฐ On-device Optimization
Mobile, Embedded, Edge Device ํ๊ฒฝ์์ LLM์ ํจ์จ์ ์ผ๋ก ์คํํ๊ธฐ ์ํ ์ต์ ํ ๊ธฐ์ ์ ๊ฐ๋ฐํฉ๋๋ค.
CPU, GPU, NPU ๋ฑ ๋ค์ํ ํ๋์จ์ด ์ํคํ ์ฒ์ ๋ง๋ ์ต์ ํ ์ ๋ต์ ์๋ฆฝํฉ๋๋ค.
์ ํ๋ ๋ฆฌ์์ค ํ๊ฒฝ์์๋ ๋์ ์ฑ๋ฅ๊ณผ ๋ฎ์ ์ ๋ ฅ ์๋น๋ฅผ ๋ฌ์ฑํ ์ ์๋๋ก ์ต์ ํํฉ๋๋ค.
Performance Analysis ๋ฐ Benchmark
๋ค์ํ ํ๋์จ์ด ๋ฐ Inference Backend์ ์ฑ๋ฅ์ ๋ถ์ํ๊ณ Benchmark๋ฅผ ์ํํฉ๋๋ค.
NVIDIA Nsight Systems, Nsight Compute, py-spy ๋ฑ์ Profiling Tool์ ํ์ฉํ์ฌ GPU ๋ฐ Runtime ๋ณ๋ชฉ์ ๋ถ์ํ๊ณ ์ง์์ ์ผ๋ก ์ถ๋ก ์ฑ๋ฅ์ ๊ฐ์ ํฉ๋๋ค.
์ฃผ์ด์ง Prompt ๋ฐ ์๋น์ค ์๊ตฌ์ฌํญ์ ๋ง์ถฐ ๋ชจ๋ธ ํ์ง, ์ถ๋ก ์ฑ๋ฅ(Latency/Throughput), ๋น์ฉ ๊ฐ์ Trade-off๋ฅผ ์ต์ ํํ๊ณ , ์ด์ ์ ํฉํ Serving Architecture๋ฅผ ์ค๊ณํฉ๋๋ค.
Qualifications
ILLM, Machine Learning Infrastructure ๋๋ Inference Optimization ๊ด๋ จ ๊ฒฝ๋ ฅ 3๋ ์ด์
LLM Inference Engine ๋๋ AI Runtime ๊ฐ๋ฐ ๊ฒฝํ
GPU Architecture, CUDA Programming ๋๋ ๋ณ๋ ฌ ์ปดํจํ ์ ๋ํ ์ดํด
Quantization, Model Compression, Compiler Optimization ๋ฑ ๋ชจ๋ธ ์ต์ ํ ๊ธฐ์ ์ ๋ํ ์ดํด
PyTorch, ONNX, TensorRT ๋ฑ ๋ฅ๋ฌ๋ ํ๋ ์์ํฌ ํ์ฉ ๊ฒฝํ
Python ๋ฐ C/C++ ์ค ํ๋ ์ด์์ ์ธ์ด์ ๋ฅ์ํ๋ฉฐ, ์ํํธ์จ์ด ์์ง๋์ด๋ง ์ญ๋์ ๋ณด์ ํ์ ๋ถ
์ฑ๋ฅ ๋ถ์ ๋ฐ ๋ฌธ์ ํด๊ฒฐ ๋ฅ๋ ฅ๊ณผ ํ์ ์ญ๋์ ๊ฐ์ถ์ ๋ถ
Preferred Qualifications
vLLM, TensorRT-LLM, SGLang, llama.cpp, MLX, ONNX Runtime ๋ฑ LLM Serving Framework ํ์ฉ ๋๋ ๊ฐ๋ฐ ๊ฒฝํ
CUDA, Triton Kernel ๋๋ Custom Operator ๊ฐ๋ฐ ๊ฒฝํ
Speculative Decoding, Prefill-Decode Disaggregation, KV Cache Compression, Expert Parallelism(MoE) ๋ฑ ์ต์ LLM Inference Optimization ๋ฐ Serving ๊ธฐ์ ๊ฒฝํ
NVIDIA GPU(H100, B200 ๋ฑ), AMD GPU ๋๋ ๋ค์ํ AI Accelerator ์ต์ ํ ๊ฒฝํ
Edge AI ๋ฐ On-device LLM ์ต์ ํ ๊ฒฝํ(Orin, Thor, Qualcomm, Apple Silicon ๋ฑ)
Kubernetes ๊ธฐ๋ฐ AI Serving ๋๋ ๋๊ท๋ชจ GPU Cluster ์ด์ ๊ฒฝํ
LLM Fine-tuning ๋ฐ Distributed Training์ ๋ํ ์ดํด
AI Systems, ML Systems ๋๋ LLM Infrastructure ๊ด๋ จ ์คํ์์ค ๊ธฐ์ฌ ๊ฒฝํ
AI/ML Systems ๋ถ์ผ์ ์ฐ์ ํํ(NeurIPS, ICML, ICLR, MLSys, OSDI, NSDI, ASPLOS ๋ฑ) ๋ ผ๋ฌธ ๋ฐํ ๋๋ ์ด์ ์คํ๋ ์ฐ๊ตฌ ๊ฒฝํ
Interview Process
์๋ฅ ์ ํ
์ฝ๋ฉยท๊ณผ์ ํ ์คํธ
1์ฐจ ๋ฉด์ (ํ์, 1์๊ฐ ๋ด์ธ)
2์ฐจ ๋ฉด์ (๋๋ฉด ํน์ ํ์, 3์๊ฐ ๋ด์ธ)
์ฒ์ฐ ํ์ยท์ ์ฌ
Additional Information
์ ํ ์ ์ฐจ๋ ์ผ์ ๋ฐ ์งํ ์ํฉ์ ๋ฐ๋ผ ์ผ๋ถ ๋ณ๊ฒฝ๋ ์ ์์ผ๋ฉฐ, ๊ฐ ์ ํ ๊ฒฐ๊ณผ๋ ๋ฑ๋กํ์ ์ด๋ฉ์ผ๋ก ๊ฐ๋ณ ์๋ด๋๋ฆฝ๋๋ค.
์ง์์ ์ ์ถ ์ ์ฃผ๋ฏผ๋ฑ๋ก๋ฒํธ, ๊ฐ์กฑ๊ด๊ณ, ํผ์ธ ์ฌ๋ถ, ์ฐ๋ด, ์ฌ์ง, ์ ์ฒด์กฐ๊ฑด, ์ถ์ ์ง์ญ ๋ฑ ์ฑ์ฉ์ ์ฐจ๋ฒ์ ์๊ตฌ ๊ธ์ง๋ ์ ๋ณด๋ ์ ์ธ ๋ถํ๋๋ฆฝ๋๋ค.
์ง์์ ์ ์ ์ค ์ค๋ฅ๊ฐ ๋ฐ์ํ๊ฑฐ๋ ๊ธฐํ ๋ฌธ์ ์ฌํญ์ด ์์ ๊ฒฝ์ฐ, recruit@42dot.ai๋ก ๋ฌธ์ํด ์ฃผ์๊ธฐ ๋ฐ๋๋๋ค.
๊ตญ๊ฐ๋ณดํ๋์์ ๋ฐ ์ทจ์ ๋ณดํธ ๋์์๋ ๊ด๊ณ๋ฒ๋ น์ ๋ฐ๋ผ ์ฐ๋ํฉ๋๋ค.
์ฅ์ ์ธ ๊ณ ์ฉ ์ด์ง ๋ฐ ์ง์ ์ฌํ๋ฒ์ ๋ฐ๋ผ ์ฅ์ ์ธ ๋ฑ๋ก์ฆ ์์ง์๋ฅผ ์ฐ๋ํฉ๋๋ค.
42dot์ ์๋ขฐํ์ง ์์ ์์นํ์ ์ด๋ ฅ์๋ฅผ ๋ฐ์ง ์์ผ๋ฉฐ, ์์ฒญํ์ง ์์ ์ด๋ ฅ์์ ๋ํด ์์๋ฃ๋ฅผ ์ง๋ถํ์ง ์์ต๋๋ค.
์ง์์ ๋ด์ฉ ์ค ํ์ ์ฌ์ค์ด ๋ฐ๊ฒฌ๋ ๊ฒฝ์ฐ, ์ ์ฌ๊ฐ ์ทจ์๋ ์ ์์ต๋๋ค.
์ธํฐ๋ทฐ ํ๋ก์ธ์ค ์ข ๋ฃ ํ ์ง์์์ ๋์ํ์ ํํ์กฐํ๊ฐ ์งํ๋ ์ ์์ต๋๋ค.
3๊ฐ์์ ์์ต๊ธฐ๊ฐ์ด ์ ์ฉ๋ ์ ์์ต๋๋ค.
LLM Engineer (Optimization) ยท 42dot