
Machine Learning Operations (MLOps) Engineer
- Machine Learning
- MLOps
- AI
- CI/CD
- Vertex AI
- BigQuery
- IaC
- Devops
- Python
- GCP
- Docker
- Kubernetes
- Terraform
- Airflow
- Composer
- SQL
- Gemini
- Risk Management
Machine Learning Operations (MLOps) Engineer
Position Summary
The MLOps Engineer builds and operates the platform and pipelines that take machine learning and AI models from development into reliable production use. This role owns model deployment, automation, monitoring, and lifecycle management, and serves as the ongoing owner of the enterprise sales forecasting models. Beyond operating the platform, the MLOps Engineer is expected to work inside the models themselves, diagnosing accuracy degradation and applying data science judgment to retune features, algorithms, and model logic so that forecast performance continues to deliver measurable business value. The MLOps Engineer partners closely with Data Science, Data Engineering, and Security teams to bridge experimentation and production, and is accountable for the operational health of models once they are live.
Key Responsibilities
·        Design, build, and maintain CI/CD pipelines for model training, validation, deployment, and rollback.
·        Productionize models developed by data scientists and vendor partners, converting experimental code into scalable, maintainable, and reproducible pipelines.
·        Own ongoing maintenance and performance of the enterprise sales forecasting models, including accuracy tracking, periodic recalibration, and continuous improvement.
·        Diagnose forecast accuracy degradation and apply data science techniques to improve results, including feature engineering, algorithm selection, hyperparameter tuning, and adjustment of model logic.
·        Evaluate forecast performance using appropriate accuracy measures such as WMAPE, and report results and improvement trends to business stakeholders.
·        Translate business context, including promotions, seasonality, café openings and closures, and operational disruptions, into model features and adjustments so forecasts remain useful for planning.
·        Partner with Finance, Operations, and Supply Chain to understand how forecasts are consumed and ensure model changes deliver measurable business benefit.
·        Deploy and manage models on Vertex AI, including training jobs, endpoints, batch prediction, and pipeline orchestration.
·        Operationalize LLM and generative AI workloads, including in-warehouse model invocation from BigQuery, prompt and model version management, and output validation.
·        Implement model monitoring for accuracy, drift, data quality, latency, and failure, with automated alerting rather than manual review.
·        Build automated retraining workflows with clear promotion criteria and approval gates before a new model version reaches production.
·        Establish model registry, versioning, and lineage practices so that every production prediction can be traced to a model version, code commit, and training dataset.
·        Manage feature pipelines and ensure consistency between training and serving data to prevent training-serving skew.
·        Monitor and optimize compute and inference cost, including model selection, batch versus real-time tradeoffs, and capacity utilization.
·        Define and enforce infrastructure-as-code standards for ML environments, including environment promotion from development through production.
·        Partner with Security and Privacy teams to ensure models and training data meet data classification, access control, and retention requirements.
·        Support model governance and audit requirements by documenting model behavior, validation results, and approval history.
·        Troubleshoot production model and pipeline failures, drive root cause analysis, and implement preventive controls.
·        Collaborate with Data Engineering to define upstream data requirements and ensure pipeline reliability for model inputs.
·        Contribute to architecture reviews and present ML platform designs and operational readiness for approval.
·        Mentor engineers and data scientists on production standards, deployment practices, and operational discipline.
Qualifications
Education
·        Bachelor's degree in Computer Science, Engineering, Data Science, Analytics, or a related field.
Experience
·        3+ years of experience in MLOps, Machine Learning Engineering, Data Engineering, or DevOps supporting production ML workloads.
·        Demonstrated experience deploying and operating machine learning models in production, and maintaining and improving them once live.
·        Hands-on data science capability, including the ability to independently modify model features, algorithms, and hyperparameters rather than only deploying models built by others.
·        Experience owning and improving time series or demand forecasting models in production, including diagnosing and correcting accuracy degradation.
·        Working knowledge of forecasting techniques and accuracy measures such as WMAPE, and the ability to explain tradeoffs between model options.
·        Proficiency with common ML and statistical libraries used for model development and evaluation.
·        Strong proficiency in Python, including packaging, testing, and writing production-quality code.
·        Hands-on experience with Google Cloud Platform, particularly Vertex AI and BigQuery.
·        Experience building CI/CD pipelines and applying software engineering practices to ML workflows.
·        Experience with containerization and orchestration (Docker, Kubernetes, or similar).
·        Experience with infrastructure-as-code tooling (Terraform or similar).
·        Working knowledge of workflow orchestration tools (Airflow, Cloud Composer, Vertex AI Pipelines, or similar).
·        Strong SQL proficiency and understanding of data pipeline and warehouse concepts.
·        Experience implementing model monitoring, drift detection, and automated alerting.
·        Understanding of model evaluation metrics and the ability to reason about model performance in production.
·        Structured problem-solving skills and persistence in driving production issues to root cause.
·        Excellent written and verbal communication skills, including the ability to explain operational risk to technical and non-technical stakeholders.
Bonus Skills & Preferred Expertise
·        Hands-on experience operationalizing LLMs, including Gemini via Vertex AI or in-warehouse invocation from BigQuery.
·        Experience with LLM-specific operational concerns such as prompt versioning, output evaluation, guardrails, and token cost management.
·        Experience with sales or demand forecasting in a multi-unit retail or restaurant environment, including café or store level granularity.
·        Experience benchmarking competing forecasting approaches and recommending a model based on measured accuracy and business impact.
·        Experience with feature store implementation and management.
·        Experience managing models delivered by external vendors or systems integrators through to production handover.
·        Familiarity with responsible AI practices, model risk management, or AI governance frameworks.
·        Restaurant, retail, or multi-unit operations data experience.
Working Conditions
Hybrid work environment with onsite presence required in Newton MA/St Louis, MO.
Competitive Pay $127,461 to $155,477 annually.
The actual pay offered will be determined by multiple factors, including but not limited to the candidate’s relevant experience, job-related knowledge, skills, and geographical location. Individual compensation decisions are dependent upon the facts and circumstances of each position and candidate.
Saint Louis Support CenterMachine Learning Operations (MLOps) Engineer · Panera, LLC