
Cloud Inference Engineer
- AI
- CUDA
- Rust
- vLLM
- TensorRT
- PyTorch
About Luminal
Making AI run fast on any hardware.
Tech
**Luminal** uses a search based approach to generate, tune, and verify GPU kernels so engineers do not have to hand write CUDA. ### Search based approach * Express computations in a small IR, then generate candidate kernels via equality saturation rewrite rules (tiling, unrolling, vectorization, memory layout). * Guide exploration with cost models and bandit style search to find the fastest valid kernels for a target GPU. * Compile and benchmark candidates on real hardware, enforce correctness with property tests and equivalence checks, and keep strict shape and dtype constraints. * Cache, version, and reuse the best kernels across models and deployments with full reproducibility. ### Tech stack * **Compiler and runtime:** Rust and egglog based compiler generating GPU kernels. Using a lightweight IR with e-graph style rewrites to search and benchmark kernels. * **Backends:** CUDA and Metal in production today. Other backends in progress.
The role
# Qualifications* CUDA + GPU inference optimization* vLLM, SGLang, or TensorRT-LLM experience* KV caching, paged attention, batching, token streaming, etc.* Distributed compute (with GPUs is a super plus)* No degree required# CompanyLuminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.# RoleFounding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.Day to day responsibilities:* Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.* Conducting model performance reviews* Improve scheduler, batcher, autoscaling; profile latency, cost, utilization* Sometimes write kernels and, yes, occasional tasteful shitposting
Skills
- Torch/PyTorch
- CUDA
Cloud Inference Engineer Β· Luminal