Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Machine Learning Ops Engineer

ResMed Pty Ltd
🇮🇳 India
On-site
1 week ago
  • AI
  • Machine Learning
  • AWS
  • Kubernetes
  • Terraform
  • CI/CD
  • AI/ML
  • IAM
  • Prometheus
  • Loki
  • Grafana
  • Datadog
  • EKS
  • EC2
  • VPC
  • RDS
  • EMR
  • Athena
  • Airflow
  • Python
  • SQL
  • GitHub
  • GitHub Actions
  • Jenkins
  • LangChain
  • LangGraph
  • CrewAI
  • AutoGen
  • Semantic Kernel
  • vLLM
  • Ray
  • RAG
  • OpenSearch
  • pgvector
  • Pinecone
  • Weaviate
  • LangSmith
  • OpenTelemetry
  • MCP
  • Model Context Protocol
  • Bedrock
  • Node.js
  • Kubeflow
  • MLflow
  • Snowflake
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Global Technology Solutions (GTS) at ResMed is a division dedicated to creating innovative, scalable, and secure platforms and services for patients, providers, and people across ResMed. The primary goal of GTS is to accelerate well-being and growth by transforming the core, enabling patient, people, and partner outcomes, and building future-ready operations.

The strategy of GTS focuses on aligning goals and promoting collaboration across all organizational areas. This includes fostering shared ownership, developing flexible platforms that can easily scale to meet global demands, and implementing global standards for key processes to ensure efficiency and consistency.

About the role 

ResMed’s AI platform powers dozens of data scientists and a growing set of GenAI / Agentic AI products that touch patients, clinicians, and providers worldwide. We run onAWS and Kubernetes, provisioned withTerraform, and shipped through modern CI/CD. 


We are looking forAI/ML Platform Engineer whose core isKubernetes, AWS, Terraform,AIand platform observability — someone who can design, build, andoperate the platform end-to-end and instrument it so nothing is a mystery in production. You should also bring anAI working mindset: curious about how ML and agentic workloads run on the platform, comfortable partnering with data scientists and GenAI teams, and eager to grow the platform towardLLMOps and Agentic AI as those workloads scale. 
 

Whatyou’ll do 

  • Design, build, andoperate theAI/ML platform on AWS + Kubernetes — clusters, networking, IAM, storage, cost, and reliability. 

  • Provision and evolve infrastructure withTerraform; treat infra as code with real review and rollback. 

  • OwnCI/CD for data pipelines, ML models, and AI applications — from repo to production with confidence. 

  • Stand up and evolve theplatform observability stack — Prometheus, Loki, Grafana / Datadog — for metrics, logs, traces, dashboards, alerting, and SLOs. 

  • Automate whatshouldn’t be manual: environment provisioning, golden-path pipelines, self-serve tooling for data scientists. 

  • Partner with product, data science, and GenAI teams to make their workloads first-class on the platform — model serving, evaluation, cost/latency controls, and safe rollout. 

  • Run POCs to pull promising tech into the platform without accumulating debt. 

  • Participate in code review, mentoring, and process improvement; raise the engineering bar. 


Whatwe’re looking for 

Must-have 

  • 3+ years of engineering experience in a complex, technical environment. 

  • Deep, hands-onKubernetes in production. 

  • Hands-onAWS — comfortable with 3+ of: EKS, Lambda, EC2, S3, IAM, Networking (VPC, ALB/NLB), RDS, EMR, Glue, Athena, Batch, SageMaker, MWAA/Airflow. 

  • Working command ofTerraform — modules, state, reviews, drift. 

  • Platform observability experience:Prometheus, Loki, Grafana and/orDatadog — metrics, logs, dashboards, alerting, SLOs. 

  • StrongproductionPython (and SQL for data work). 

  • Experience buildingCI/CD pipelines and APIs end-to-end — GitHub / GitHub Actions,CodePipeline or Jenkins. 

  • Hands-onworking experience with an AI/ML platform in production — data science tooling, model lifecycle, feature / inference infrastructure, and self-serve enablement for DS and GenAI teams. 

  • Deploying AI agents / LLM workloads on Kubernetes — containerizing agent workloads, autoscaling (HPA/KEDA), GPU scheduling where needed, secure egress for tool calls, secrets and rate-limit management, and running long-lived / stateful sessions safely. 

  • Exposure to the modern AI / Agentic AI stack is required — working familiarity with at least a few of: an agent framework (LangChain /LangGraph /CrewAI /AutoGen / Strands / Semantic Kernel /PydanticAI), LLM serving (vLLM,KServe, Ray Serve, TGI), a RAG / vector-store setup (OpenSearch,pgvector, Pinecone,Weaviate), LLM observability (Langfuse,LangSmith,Arize Phoenix,OpenTelemetry GenAI), and MCP (Model Context Protocol) for tool integration. 


Nice-to-have — AI / Agentic AI skills 

  • AI / Agent frameworks:LangChain,LangGraph, Strands, or similar. 

  • Running AI agents on Kubernetes: containerizing agent workloads, autoscaling (HPA/KEDA), stateful sessions, long-running tasks/jobs, secure egress for tool calls,secrets and rate-limit management. 

  • Managed agent platforms: AWSBedrockAgentCore, Bedrock Agents / Knowledge Bases, SageMaker. 

  • MCP (Model Context Protocol): authoring or hosting MCP servers/clients, exposing internal tools/data safely to agents. 

  • LLM/agent observability:Langfuse,LangSmith,Arize orOpenTelemetry GenAI — traces, evaluations, token / cost / latency tracking. 

  • LLM serving on Kubernetes:vLLM,KServe, Ray Serve, TGI; GPU node pools and scheduling. 

  • RAG stack: vector stores (OpenSearch,pgvector, Pinecone), embeddings pipelines, retrieval evaluation. 

  • Guardrails & safety: Bedrock Guardrails, prompt-injection defenses, PII redaction. 

  • ML platform tooling: Kubeflow,MLflow, or comparable. 

  • Snowflake and modern data stack experience. 


Why join 

A supportive, senior team with real problems and real users. Freedom to design and influence. Global collaboration and open exchange of ideas. And the chance to build a platform whose output shows up — directly — in better sleep, better breathing, and better health for millions of people. 

ResMed’s AI platform powers dozens of data scientists and a growing set of GenAI / Agentic AI products that touch patients, clinicians, and providers worldwide. We run onAWS and Kubernetes, provisioned withTerraform, and shipped through modern CI/CD. 

Joining us is more than saying “yes” to making the world a healthier place. It’s discovering a career that’s challenging, supportive and inspiring. Where a culture driven by excellence helps you not only meet your goals, but also create new ones. We focus on creating a diverse and inclusive culture, encouraging individual expression in the workplace and thrive on the innovative ideas this generates. If this sounds like the workplace for you, apply now!We commit to respond to every applicant.

 

Machine Learning Ops Engineer · ResMed Pty Ltd

Auto apply with Likeremote