Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
M

Production Engineering Manager

Meta
๐Ÿ‡ฌ๐Ÿ‡ง United Kingdom
On-site
Manager or above
4 weeks ago
  • Incident Response
  • AI
  • System Design
  • Incident Management
  • Python
  • C++
  • Bash
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems โ€” from capacity planning and performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.

Responsibilities

  • Manage a team of production engineers delivering on reliability, scalability, and operational efficiency across multiple interdependent production systems
  • Drive roadmap creation for infrastructure reliability initiatives, capacity planning, and automation efforts, increasing team scope as AI-driven productivity improves throughput
  • Lead adoption of AI-augmented engineering workflows across the team, sharing learnings and best practices with the broader production engineering organization
  • Contribute hands-on to technical work including code, system design reviews, and incident response, using AI tooling to expand personal and team reach across disciplines
  • Partner cross-functionally with software engineering, data science, and product teams to unblock dependencies and ensure smooth execution of infrastructure and reliability projects
  • Proactively identify and resolve sources of operational toil โ€” including on-call load, alerting gaps, and technical debt โ€” and implement automation to increase team efficiency and scope
  • Set clear goals and expectations for individual team members, provide timely and actionable feedback, and actively develop engineers' skills including proficiency with AI-augmented workflows
  • Establish and monitor service-level objectives, reliability metrics, and engineering efficiency indicators to maintain high engineering craft and product quality
  • Communicate production system health, incident learnings, and infrastructure strategy effectively to engineering leadership and cross-functional stakeholders
  • Recruit, onboard, and retain production engineering talent, ensuring the team structure minimizes single points of failure and supports sustainable growth

Minimum qualifications

  • 4+ years of experience in production engineering, site reliability engineering, or systems software engineering
  • 2+ years of experience managing production engineering or infrastructure engineering teams
  • Experience driving reliability and scalability improvements for large-scale distributed systems, including incident management, capacity planning, and performance optimization
  • Experience coding and debugging in at least one systems or scripting language (such as Python, C++, Go, or Bash) and contributing technically alongside a team
  • Experience setting team goals, managing execution against roadmaps, and communicating infrastructure strategy to technical and non-technical stakeholders Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Track record of cross-functional collaboration with software engineering and data science teams to co-own reliability outcomes for consumer-facing or infrastructure services
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience leading adoption of AI-assisted tooling or automation frameworks within an engineering team to expand operational scope and reduce toil
  • Experience managing on-call rotations, defining service-level objectives, and implementing observability and alerting improvements at scale
  • Familiarity with container orchestration, service mesh architectures, or large-scale deployment pipelines in a production environment
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)

Production Engineering Manager ยท Meta

Auto apply with Likeremote