Senior Machine Learning Engineer
- 🇺🇸 United States
- On-site
- Senior
- 3 weeks ago
- $175,000 – $308,500
- Machine Learning
- AI
- Python
- Rust
- PyTorch
- JAX
- TensorFlow
- Incident Response
- C++
- Parquet
- Ray
- Unity Catalog
- Polars
- DuckDB
- Docker
- Kubernetes
Summary
Join a team at the forefront of ML infrastructure and generative AI, where data and model workflows come together to enable the next generation of intelligent experiences on Apple products and services. We build robust systems that connect scalable data pipelines with advanced ML workflows, accelerating the development of real-world AI applications. Our work spans the full ML lifecycle, from experimentation to deployment, and you’ll play a key role in shaping how AI models are built, optimized, and scaled. We develop a platform for ML data and features that powers advanced GenAI applications. This includes embeddings (generation, evaluation, ANN search, multimodal support), AI Ops, efficient inference, and a modern feature platform designed to streamline experimentation and drive innovation. We’re looking for engineers and researchers passionate about generative models, data-centric ML, and intelligent systems across diverse real-world use cases. With the autonomy to experiment, the scale to make an impact, and the support to take ideas from prototype to production, you’ll work alongside a world-class team to build intelligent, flexible systems that make ML development faster, more reliable, and more creative.
Description
The Apple AI Platform team gives Apple's ML engineers and researchers the data systems and large-scale compute they need to build and ship models at Apple's bar for quality and privacy. Our team owns the data layer that large-scale model training depends on: ingestion, versioning, lineage, and governance on the way in, and high-throughput data loading into the training fleet on the way out. As a Senior ML Infrastructure Engineer, you will own and ship major components of that platform.
Responsibilities
- Build and own high-throughput data access and loading primitives that feed Apple's largest GPU and TPU fleets, keeping training compute-bound rather than I/O-bound.
- Develop and evolve the Python and Rust data libraries and SDKs that ML engineers depend on to access, transform, and load model-ready datasets across every stage of model development.
- Build and operate distributed data pipelines across Spark, Daft, and Rust-based systems for ingestion, transformation, and large-scale data preparation.
- Contribute to the platform behind Apple's largest model builds: ingestion, immutable versioning, lineage, and governance across structured, unstructured, and multimodal data at petabyte scale, so every model run is reproducible from a versioned dataset.
- Make dataset access a first-class concern in the model development loop through tight integration with the data-loading layers of PyTorch, JAX, and TensorFlow.
- Partner with research and product teams to onboard new data sources and unblock fast iteration on the datasets powering foundation-model and GenAI workloads.
- Diagnose, fix, and automate away complex issues across the stack, from ingestion pipelines to dataset APIs to framework integrations, to maximize uptime and throughput.
Minimum Qualifications
- * 5+ years of work experience in machine learning infrastructure, distributed data systems, or a related field.
- * 5+ years of experience building and shipping data or ML infrastructure and systems in production.
- * Experience operating production systems: monitoring, observability, on-call, and incident response.
- * Strong programming skills in Python, plus a systems language for performance-critical work (Rust strongly preferred; C++ or Go acceptable).
- * Familiarity with columnar and lakehouse formats (Parquet, Iceberg, Delta, or Lance) and the trade-offs between them.
- * Hands-on performance engineering for I/O-bound workloads: Arrow, zero-copy, memory mapping, async I/O, and high-throughput object storage access patterns.
- * Working knowledge of the end-to-end ML workflow and how training and inference workloads consume data, enough to design data systems that serve them well.
- Familiarity with modern ML and generative techniques (transformers, diffusion, retrieval-augmented generation, fine-tuning) at the level needed to design for those consumers, not necessarily to build the models yourself.
- Experience designing and operating scalable, highly available systems that prioritize ease of use for the engineers who depend on them.
- Strong collaboration and written and verbal communication skills.
- B.S., M.S., or Ph.D. in Computer Science, Computer Engineering, or equivalent practical experience.
Preferred Qualifications
- * Deep experience with the data-loading and dataset-access layer of a modern ML framework (PyTorch, JAX, or TensorFlow).
- * Distributed data-loading frameworks for ML: Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset.
- * Experience feeding data to GPU or TPU fleets at scale and keeping them saturated.
- * Data lineage and governance systems: DataHub, OpenLineage, Unity Catalog, or equivalent.
- * Contributions to or operational experience with Spark, Daft, Polars, or DuckDB internals.
- * Containerization and orchestration (Docker, Kubernetes).
Senior Machine Learning Engineer · Apple