Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
WQ

Architect - Machine Learning (MLOps Specialist)

wd1:quantiphi:careers_at_quantiphi
๐Ÿ‡ฎ๐Ÿ‡ณ India
On-site
2 months ago
  • Excel
  • Machine Learning
  • MLOps
  • EKS
  • ECS
  • API Gateway
  • CI/CD
  • MLflow
  • Kubeflow
  • Airflow
  • Step Functions
  • SQL
  • Databricks
  • EMR
  • OpenAI
  • Prometheus
  • Grafana
  • Python
  • IAM
  • Secrets Management
  • RAG
  • AI
  • Devops
  • IaC
  • Terraform
  • Helm
  • CloudWatch
  • OpenTelemetry
  • AI/ML
  • Kubernetes
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.


If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Role : Architect - Machine Learning

Experience: 7-14 Years

Location: Mumbai/Bangalore

Must have skills & Qualifications:

  • 8+ years working in ML/AI engineering or MLOps roles with strong architecture exposure.

  • Strong expertise inAWS cloud-native ML stack, including: EKS (primary), ECS, Lambda, API Gateway, CI/CD (CodeBuild/CodePipeline or equivalent)

  • Hands-on experience with at least one major MLOps toolset and awareness of alternatives: MLflow, Kubeflow, SageMaker Pipelines, Airflow, BentoML, KServe, Seldon

  • Deep understanding ofmodel lifecycle management (training โ†’ registry โ†’ deployment โ†’ monitoring).

  • Experience implementing or supportingLLMOps pipelines, including: prompt versioning, evaluation metrics, automation frameworks

  • Deep understanding ofML lifecycle: data ingestion, feature engineering, training, evaluation, model packaging, CI/CD, drift detection, monitoring, and governance.

  • Strong experience withAWS SageMaker (Training, Processing, Batch Transform, Pipelines, Feature Store, Model Registry, Model Monitor).

  • Experience implementingML CI/CD pipelines including automated training, testing, validation, model promotion, and endpoint deployment.

  • Ability to builddynamic and versioned pipelines using SageMaker Pipelines, Step Functions, or Kubeflow.

  • Strong SQL and data transformation experience usingSnowflake, Databricks, Spark, or EMR.

  • Experience withfeature engineering pipelines andFeature Store management (SageMaker or Feast).

  • Understanding oflineage tracking: training data snapshot, feature versions, code versioning, metadata tracking, reproducibility.

  • Hands-on experience withBedrock,OpenAI,Anthropic, orLlama models.

  • Experience withCloudWatch,SageMaker Model Monitor,Prometheus/Grafana, orDatadog.

  • Strong foundation in Python and cloud-native development patterns.

  • Solid understanding of security best practices, IAM, secrets management, and artifact governance.

Good to have skills:

  • Experience with vector databases, RAG pipelines, or multi-agent AI systems.

  • Exposure to DevOps and infrastructure-as-code (Terraform, Helm, CDK).

  • Hands-on understanding of model drift detection, A/B testing, canary rollouts, and blue-green deployments.

  • Familiarity with Observability stacks (Prometheus, Grafana, CloudWatch, OpenTelemetry).

  • Knowledge ofLakehouse (Delta/Iceberg/Hudi) architecture.

  • Ability to translate business goals into scalable AI/ML platform designs.

  • Strong communication and cross-team collaboration skills.

  • Ability to guide engineering teams through technical uncertainty and design choices.

Key Responsibilities:

  • Architect and implement the MLOps strategy for the EVOKE Phase-2 programme, ensuring alignment with the project proposal and delivery roadmap.

  • Design and ownenterprise-grade ML/LLM pipelines covering model training, validation, deployment, versioning, monitoring, and CI/CD automation.

  • Buildcontainer-oriented ML platforms (EKS-first) while evaluating alternative orchestration tools with similar capabilities (Kubeflow, SageMaker, MLflow, Airflow, etc.).

  • Implement hybridMLOps + LLMOps workflows, including prompt/version governance, evaluation frameworks, and monitoring for LLM-based systems.

  • Serve as a technical authority across multiple internal and customer projects, not limited to EVOKE, contributing architectural patterns, best practices, and reusable frameworks.

  • Enableobservability, monitoring, drift detection, lineage tracking, and auditability across ML/LLM systems.

  • Collaborate with cross-functional teams โ€” data engineering, platform, DevOps, and client stakeholders โ€” to deliver production-ready ML solutions.

  • Ensure all solutions adhere tosecurity, governance, and compliance expectations, particularly around handling cloud services, Kubernetes workloads, and MLOps tools.

  • Conduct architecture reviews, troubleshoot complex ML system issues, and guide teams through implementation across cloud-native ML platforms.

  • Mentor engineers and provide guidance on modern MLOps tools, platform capabilities, and best practices.

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Architect - Machine Learning (MLOps Specialist) ยท wd1:quantiphi:careers_at_quantiphi

Auto apply with Likeremote