
Senior Machine Learning Engineer
- Machine Learning
- MLOps
- AI
- Devops
- gRPC
- Vector Search
- Docker
- Kubernetes
- Secrets Management
- CI/CD
- AI/ML
- Python
- Java
- REST API
- Microservices
- MLflow
- Kubeflow
- Databricks
- LLM APIs
- RAG
- LangChain
- LangGraph
- Semantic Kernel
- AWS
- Azure
- GCP
- Git
- RBAC
- OAuth
- OIDC
- NoSQL
- IaC
- Terraform
Career Category
EngineeringJob Description
We are seeking aSenior Machine Learning Platform Engineer to design, build and scale enterprise-grade machine-learning and generative-AI platform capabilities.
This role sits at the intersection ofplatform engineering, software engineering, MLOps and GenAI engineering. Rather than primarily developing individual machine-learning models or business-specific AI applications, you will build theshared platform capabilities, APIs, automation and engineering patterns that enable teams to develop, evaluate, deploy, govern and operate AI solutions at scale.
You will work closely with data scientists, ML engineers, application teams, DevOps, Security, Compliance and Product teams to create a secure, reliable and frictionless AI developer experience. The role combines hands-on engineering with technical leadership, helping define platform standards, reusable patterns and architecture for enterprise AI systems.
Roles & Responsibilities
- Design and buildreusable ML and GenAI platform capabilities that support model development, experimentation, evaluation, deployment and production operations across multiple teams and use cases.
- Buildself-service platform services, APIs and automation that abstract infrastructure complexity and enable developers to provision and consume AI capabilities consistently.
- Develop and maintainMLOps capabilities including experiment tracking, model and prompt registries, evaluation frameworks, deployment workflows and automated promotion across environments.
- Buildmodel-serving and inference capabilities that support classical ML models, deep-learning models and LLMs through scalable REST, gRPC or event-driven interfaces.
- Develop platform capabilities forGenAI and agentic systems, including model access, prompt management, embeddings, vector search, Retrieval-Augmented Generation, tool integration and agent frameworks.
- Engineer integrations with majorcloud-based AI and data platforms, using APIs, SDKs and managed services to provide reusable enterprise capabilities.
- Build and maintaincontainerized platform services using Docker and Kubernetes, including deployment patterns, scaling strategies, service configuration and lifecycle management.
- Design and implementplatform APIs, SDKs, templates and shared libraries that establish standardized development patterns and reduce duplication across engineering teams.
- Implement comprehensiveobservability and operational monitoring, including logs, metrics, distributed tracing, service health, model/LLM usage, latency, errors and operational dashboards.
- ImplementAI evaluation and quality-management capabilities, including automated evaluation pipelines, regression testing, model comparison and release-quality gates.
- Buildsecurity and governance controls into platform capabilities, including authentication, authorization, secrets management, data access controls, auditability, lineage and responsible-AI controls.
- Design platform mechanisms forusage metering, cost visibility and optimization, enabling teams to understand infrastructure and AI-service consumption.
- Engineer platform services forscalability, reliability and resilience, including retries, asynchronous processing, concurrency controls, fault tolerance and graceful failure handling.
- Develop and improveCI/CD pipelines for AI platform services, reusable components and model-based workloads, including automated testing, artifact management and environment promotion.
- Evaluate new AI, ML and cloud technologies and determine when they should be introduced asshared platform capabilities versus application-specific solutions.
- Partner with data scientists and ML engineers to identify recurring development and operational challenges and convert them intoreusable platform patterns and services.
- Provide technical guidance on architecture, scalability, performance, security and production-readiness for ML and GenAI workloads.
- Participate in architecture reviews, code reviews, incident resolution and production troubleshooting across application, platform and infrastructure layers.
- Create and maintaintechnical designs, architecture decision records, development standards, operational runbooks and platform documentation.
- Help define the longer-termtechnical roadmap and engineering standards for ML and GenAI platform capabilities.
Must-Have Skills
- 3โ5 years of experience in machine learning engineering, ML platform engineering, MLOps, backend engineering, cloud engineering or enterprise AI systems.
- Strongsoftware-engineering fundamentals with experience building production-grade distributed services, APIs and reusable libraries.
- Strong programming skills inPython; experience with Java or another enterprise programming language is preferred.
- Experience designing and developingREST APIs, microservices and backend platform services.
- Hands-on experience withDocker and Kubernetes or equivalent containerization and orchestration technologies.
- Strong understanding ofMLOps and GenAIOps concepts, including experiment tracking, model lifecycle management, evaluation, deployment, monitoring and reproducibility.
- Experience with platforms and technologies such asMLflow, Kubeflow, SageMaker, Databricks or equivalent ML platforms.
- Experience building or integratingmodel-serving infrastructure for ML models and/or LLM-based applications.
- Strong understanding of modernGenAI architecture patterns, including LLM APIs, prompt management, embeddings, vector databases, RAG and agent-based systems.
- Experience with GenAI or agent frameworks such asLangChain, LangGraph, Semantic Kernel or equivalent frameworks.
- Experience integrating cloud-basedAI/ML SaaS and PaaS services and building abstractions around those services for broader developer consumption.
- Experience with at least one major cloud platform such asAWS, Azure or GCP.
- Familiarity withCI/CD, Git-based development, automated testing and release-management practices.
- Understanding ofobservability and distributed-system concepts, including logging, metrics, tracing, retries, timeouts and failure handling.
- Understanding ofenterprise security patterns, including authentication, authorization, RBAC, OAuth/OIDC, secrets management and encryption.
- Familiarity withdata-governance and responsible-AI concepts, including lineage, explainability, access controls, model evaluation and bias monitoring.
- Experience working with relational and NoSQL databases and structured or unstructured data systems.
- Ability to designreusable platform abstractions rather than solving requirements through one-off application implementations.
- Ability to evaluate architectural options and communicate trade-offs acrossperformance, scalability, reliability, security and cost.
- Strong collaboration and communication skills, with the ability to work acrossdata science, engineering, infrastructure, security and product teams.
Preferred Skills
- Experience buildinginternal developer platforms, ML platforms or self-service engineering platforms.
- Experience withInfrastructure as Code such as Terraform.
- Experience designingmulti-tenant platform services with resource isolation, quotas, governance and usage attribution.
- Experience withfeature stores, model registries, evaluation platforms, AI gateways or centralized model-serving architectures.
- Familiarity with event-driven architectures, message queues and asynchronous processing.
- Experience defining technical standards, architecture patterns and reusable engineering frameworks across multiple teams.
Senior Machine Learning Engineer ยท Amgen Inc.