Subscribe to the latest remote jobs:

AI Researcher — Inference Optimization

🌏 Worldwide

CUDA

Python

Machine Learning

Design

AI Researcher — Inference Optimization

from 🌏 Worldwide

Role Overview

We are seeking anAI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection ofmodel architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments.

Key Responsibilities

  • Research and develop techniques tooptimize inference performance for large neural networks.

  • Improvelatency, throughput, memory efficiency, and cost per inference.

  • Design and evaluatemodel-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications).

  • Implementsystems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).

  • Benchmark inference workloads across hardware accelerators.

  • Collaborate with engineering teams todeploy optimized inference pipelines.

  • Translate research insights intoproduction-ready improvements.

Required Qualifications

  • Strong background inmachine learning, deep learning, or AI systems.

  • Hands-on experience optimizing inference forlarge-scale models.

  • Proficiency inPython and modern ML frameworks (e.g., PyTorch).

  • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).

  • Ability to design experiments and communicate results clearly.

Preferred / Nice-to-Have Qualifications

  • Experience deployingproduction inference systems at scale.

  • Familiarity withdistributed and multi-GPU inference.

  • Experience contributing toopen-source ML or inference frameworks.

  • Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields.

  • Experience working close to hardware (CUDA, ROCm, profiling tools).

What Success Looks Like

  • Measurable gains inlatency, throughput, and cost efficiency.

  • Optimized inference systems running reliably in production.

  • Research ideas successfully translated into deployable systems.

  • Clear benchmarks and documentation that inform product decisions.

Relevant Research Areas (Bonus)

  • Long-context inference optimization

  • Speculative decoding

  • KV-cache compression and paging

  • Efficient decoding strategies

  • Hardware-aware inference design

by @maxrusakovic