Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TI

AI Systems & Infrastructure Engineer

Triune Infomatics
  • ๐Ÿ‡บ๐Ÿ‡ธ United States
  • On-site
  • 1 month ago
  • AI
  • Node.js
  • PCIe
  • CI/CD
  • PyTorch
  • TensorFlow
  • Jupyter
  • vLLM
  • Ollama
  • Nim
  • Triton
  • InfiniBand
  • Load Balancing
  • CUDA
  • Kubernetes
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Role: AI Systems & Infrastructure Engineer
Location: San Jose, CA (Hybrid)
Duration: 2+ months


Summary: We are seeking an expert AI Systems & Infrastructure Engineer to design, build, and optimize the end-to-end environment for large-scale AI models. You will bridge the gap between GPU kernels and production serving, focusing on distributed inference orchestration, hardware-software co-optimization, and the elimination of system-level bottlenecks to ensure maximum throughput and minimum latency.

Key Responsibilities:
  • End-to-End Infrastructure: Architect the full lifecycle from GPU resource allocation to high-performance serving layers.
  • Distributed Orchestration: Implement Tensor and Pipeline Parallelism to deploy massive models across multi-GPU/multi-node clusters.
  • System Optimization: Maximize hardware utilization via advanced AI memory management (KV cache, PagedAttention, quantization).
  • Profiling & Benchmarking: Systematically identify AI bottlenecks (NVLink, PCIe, HBM bandwidth) and establish rigorous TPS/TTFT benchmarking suites.
  • Deployment & Testing: Build AI-specific CI/CD pipelines for automated performance gating, model validation, and regression testing.
  • Co-Design: Align model architectures with target hardware constraints to optimize underlying inference engines.

Technical Qualifications:
  1. Frameworks & Acceleration
  • Core: Expert PyTorch and TensorFlow; proficient in Jupyter/Colab for profiling.
  • Serving Engines: Deep expertise in vLLM, SGLang, Ollama, and Nvidia NIM.
  • Acceleration: Advanced knowledge of Nvidia Dynamo (TorchDynamo) and graph-compilation for execution path optimization.
  1. Systems & Infrastructure
  • Distributed Compute: Proficiency in multi-node communication (NCCL, MPI) and distributed inference strategies.
  • Memory & Precision: Expert in GPU memory layouts, quantization (INT8, FP8, NF4), and PagedAttention.
  • Profiling Tools: Mastery of NVIDIA Nsight, PyTorch Profiler, and Triton for kernel and memory access analysis.
  • Deployment: Experience building AI-centric CI/CD pipelines with automated performance gating and canary deployments.
  1. Hardware & Architecture
  • GPU Architecture: Deep understanding of H100/A100 internals (SMs, Tensor Cores, HBM, NVLink/InfiniBand).
  • Bottleneck Analysis: Ability to diagnose and resolve compute-bound, memory-bound, and I/O-bound workloads.
  • Validation: Experience in stress testing, load balancing, and failover validation for distributed AI nodes.

Preferred Qualifications:
  • Custom CUDA kernel development or Triton optimization.
  • Kubernetes (K8s) for GPU orchestration and scheduling.
  • Experience with distributed training (DeepSpeed, Megatron-LM).
  • OS-level knowledge of memory paging and asynchronous I/O.

Soft Skills:
  • Systems Thinking: Ability to map computational operations directly to physical hardware.
  • Analytical Rigor: Data-driven approach to tuning based on profiles rather than intuition.
  • Collaboration: Ability to translate researcher requirements into concrete infrastructure specs.

AI Systems & Infrastructure Engineer ยท Triune Infomatics

Auto apply with Likeremote