Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
AA

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Anyone Ai
πŸ‡¦πŸ‡· Argentina | πŸ‡§πŸ‡· Brazil | πŸ‡¨πŸ‡± Chile | πŸ‡¨πŸ‡΄ Colombia | πŸ‡ͺπŸ‡¨ Ecuador | πŸ‡ͺπŸ‡Έ Spain | πŸ‡²πŸ‡½ Mexico | πŸ‡΅πŸ‡Ή Portugal | πŸ‡ΊπŸ‡Ύ Uruguay
Remote
1 day ago
$65 / hour
  • AI
  • CUDA
  • Triton
  • AWS
  • JAX
  • RLHF
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Anyone AI is recruiting experiencedGPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such asCUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at leasttwo of the following:

    • CUDA

    • Triton

    • NKI / AWS Neuron

    • Pallas / JAX

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such asNsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

    • Memory bandwidth

    • Compute throughput

    • GPU occupancy

    • Shared memory

    • Register pressure

    • Memory coalescing

    • Bank conflicts

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

Nice to Have

  • Experience across bothNVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware,optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance Β· Anyone Ai

Auto apply with Likeremote