Subscribe to the latest remote jobs:

Machine Learning Engineer — AI Architecture Research

🌏 Worldwide

Machine Learning

Design

Machine Learning Engineer — AI Architecture Research

from 🌏 Worldwide

About the Role

We’re looking for aMachine Learning Engineer focused on AI architecture research to help design, prototype, and validate next-generation model architectures. You’ll work at the intersection of research and production — turning new ideas into scalable, real-world systems.

This role is ideal for someone who enjoysquestioning architectural assumptions, experimenting with novel model designs, and pushing beyond standard Transformer-style approaches.

What You’ll Work On

  • Research and developnew neural network architectures (e.g. alternatives or extensions to Transformers, recurrent / hybrid models, long-context systems)

  • Design and runarchitecture-level experiments (scaling laws, memory mechanisms, compute trade-offs)

  • Prototype models end-to-end — from research code to training-ready implementations

  • Collaborate with inference and systems engineers to ensure architectures aredeployable and efficient

  • Analyze model behavior, failure modes, and inductive biases

  • Read, reproduce, and extend cutting-edge research papers

  • Contribute to internal research notes, benchmarks, and open-source efforts (where applicable)

What We’re Looking For

  • Strong background inmachine learning fundamentals and deep learning

  • Hands-on experience implementingmodel architectures from scratch

  • Solid understanding of:

    • Attention mechanisms, RNNs, state-space models, or hybrid architectures

    • Training dynamics, scaling behavior, and optimization

    • Memory, latency, and compute constraints at the model level

  • Comfortable working inPyTorch or JAX

  • Ability to move fluidly between theory, experimentation, and engineering

  • Clear communicator who can explain architectural trade-offs

Nice to Have

  • Experience withnon-Transformer architectures (RNN variants, SSMs, long-context models)

  • Background inresearch-driven startups or open-source ML projects

  • Experience with large-scale training or custom training loops

  • Publications, preprints, or notable research contributions

  • Familiarity with inference optimization and deployment constraints

Why Join

  • Work oncore model architecture, not just fine-tuning

  • Direct influence on the technical direction of a Series-A company

  • Small, high-caliber team with fast feedback loops

  • Opportunity to ship research into production

  • Competitive compensation + meaningful equity

by @maxrusakovic