Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
AP

Sr Platform Engineer

ATech Placement
  • 🇺🇸 United States
  • Remote
  • Senior
  • 15 hours ago
  • Devops
  • CI/CD
  • IaC
  • Disaster Recovery
  • AI
  • Claude Code
  • Cursor
  • MCP
  • GitLab CI/CD
  • AWS
  • CloudFormation
  • Terraform
  • ECS
  • Fargate
  • Aurora
  • CloudFront
  • API Gateway
  • SQS
  • SNS
  • WAF
  • IAM
  • CloudWatch
  • PagerDuty
  • SOC2
  • Secrets Management
  • GitLab CI
  • Docker
  • Kubernetes
  • Python
  • Bash
  • PowerShell
  • Java
  • Angular
  • MySQL
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Platform Engineering Lead / DevOps Lead

We are seeking a hands-on Platform Engineering Lead to build and lead a modern cloud platform function focused on continuous delivery, infrastructure automation, reliability, and observability.

This is a true player/coach role. You will set technical direction and grow the platform engineering team over time, but you will also personally build the CI/CD pipelines, infrastructure-as-code templates, monitoring, and automation that support production systems.

The goal is to enable multiple production deployments per day with automated testing, automated rollback/failback, proactive monitoring, self-healing infrastructure, and tested disaster recovery.

AI-assisted engineering is also an important part of the role. You will use tools such as Claude Code, Cursor, AI agents, agent skills, and MCP servers to accelerate infrastructure and automation work while maintaining ownership of the quality and reliability of what is deployed.

What You’ll Do

  • Build and own the path from code commit through production deployment, enabling multiple production releases per day.

  • Design automated testing gates and automatic rollback/failback when deployments fail.

  • Own GitLab CI/CD, including pipeline architecture, build reliability, and automated deployments across AWS environments.

  • Manage the AWS environment through Infrastructure as Code (IaC) using CloudFormation and/or Terraform, eliminating manual console changes.

  • Work hands-on with AWS services including ECS, Fargate, Aurora, S3, CloudFront, API Gateway, Cognito, SQS/SNS, Lambda, WAF, IAM, and CloudWatch.

  • Build and maintain scalable QA, demo, UAT, and production environments, including automated environment provisioning.

  • Implement self-healing capabilities including health checks, auto-scaling, automated restarts, and recovery mechanisms.

  • Own monitoring, observability, alerting, and on-call infrastructure, including metrics, logs, traces, alarms, and PagerDuty.

  • Establish monitoring that identifies production issues before they become customer-impacting incidents.

  • Own infrastructure disaster recovery, including documented recovery targets and quarterly DR testing.

  • Partner with IT on AWS account structure, IAM, networking, cloud cost optimization, right-sizing, and monthly spend reporting.

  • Support infrastructure-related security and compliance requirements, including SOC 2 evidence, secrets management, encryption at rest, vulnerability remediation, and patching.

  • Use Claude Code and other AI coding agents to build, test, review, and accelerate pipelines, templates, scripts, and infrastructure automation.

  • Set technical direction, conduct code reviews, mentor engineers, and eventually help interview and build out the platform engineering team.

  • Remain highly hands-on with engineering and automation as the team grows.

What We’re Looking For

  • 7+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or similar infrastructure-focused roles.

  • Proven experience owning a production cloud environment.

  • Demonstrated experience taking an engineering organization to continuous delivery, including multiple production deployments per day.

  • Strong experience building automated deployment pipelines with test gates and automated rollback/failback.

  • Deep hands-on AWS experience, ideally including ECS/Fargate, Aurora, S3, CloudFront, API Gateway, Cognito, Lambda, SQS/SNS, WAF, IAM, and CloudWatch.

  • Strong production experience with Infrastructure as Code, preferably CloudFormation and/or Terraform.

  • Significant ownership of CI/CD pipelines at scale; GitLab CI is strongly preferred.

  • Production experience with Docker and container orchestration using ECS/Fargate or Kubernetes.

  • Experience building monitoring, alerting, observability, and on-call processes from the ground up and tuning alerts to reduce unnecessary paging.

  • Scripting experience with Python, Bash, and/or PowerShell.

  • Ability to understand Java and Angular build tooling well enough to troubleshoot CI/CD and deployment issues.

  • Hands-on experience using AI coding agents such as Claude Code or Cursor for infrastructure, DevOps, or automation work. Candidates should be able to explain what they built faster using AI, where the AI-generated solution failed or needed correction, and how they validated the final result.

  • Experience leading a small engineering team or serving as a technical lead, including setting direction, reviewing work, and mentoring engineers.

  • Strong understanding of cloud security practices including least-privilege IAM, secrets management, patching, encryption, and compliance evidence.

  • Bachelor's degree in Computer Science or a related technical field.

Nice to Have

  • PagerDuty experience

  • Aurora MySQL administration or operations

  • Database migrations across multiple customer environments

  • Encryption-at-rest implementations across large database estates

  • Experience building MCP servers, agent skills, or agent-based engineering workflows

  • Experience supporting SOC 2 or similar compliance environments

Sr Platform Engineer · ATech Placement

Auto apply with Likeremote