
Cloud Operations Engineer
- 🇺🇸 United States
- Remote
- 4 hours ago
- Incident Management
- Devops
- Incident Response
- CI/CD
- IaC
- triage
- IIS
- .NET
- Azure
- Azure Government
- Configuration Management
- Disaster Recovery
- PowerShell
- Python
- Bash
- Windows Server
- Azure SQL
- Performance Testing
- Load Balancing
- Datadog
- Elastic
- Terraform
- Ansible
- Azure DevOps
- Cloudflare
About the role
We are seeking a Cloud Operations Engineer to join the team. While the position supports broader cloud operations responsibilities, its primary focus is Site Reliability Engineering (SRE) for JustFOIA, our SaaS platform used by government agencies to receive, track, and fulfill public records requests. This hands-on role is responsible for improving JustFOIA’s reliability, availability, performance, and operational maturity. The role will use observability, automation, and disciplined incident-management practices to identify problems early, restore service quickly, prevent recurring failures, and reduce repetitive operational work. This position will work closely with the JustFOIA development team while contributing to deployment automation, infrastructure improvements, and the platform’s evolving DevOps practices. The role may also support other MCCi SaaS solutions as the Cloud Operations team’s responsibilities evolve.
Â
The anticipated focus of the position is approximately:
·  70% Site Reliability Engineering: Monitoring, incident response, root cause analysis, reliability improvements, automation, and operational ownership
·  20% DevOps: CI/CD, Infrastructure as Code (IaC), and deployment improvements
·  10% Architecture: Design collaboration, reliability guidance, and scalability recommendations
What you'll do
Monitoring and Observability
- Build and maintain monitoring, alerting, dashboards, logging, and operational reporting across application, database, and infrastructure tiers.
- Proactively analyze logs, metrics, traces, and service-health indicators to identify emerging and recurring issues.
- Improve operational visibility while increasing alert quality and reducing noise.
Reliability and Incident Management
- Improve platform reliability, availability, performance, scalability, and maintainability.
- Participate in the on-call rotation and lead incident triage, troubleshooting, escalation, and service restoration.
- Isolate production issues across application, IIS/.NET, SQL Server, network, and cloud-platform components.
- Lead blameless post-incident reviews, complete root cause analyses, and drive corrective actions to completion.
Cloud Operations and Automation
- Operate and improve Azure and Azure Government environments.
- Develop automation that reduces repetitive work, human error, and configuration inconsistency.
- Maintain and enhance Infrastructure-as-Code, configuration management, deployment, and operational tooling.
- Support CI/CD pipelines and application release processes.
Performance and Resilience
- Identify performance bottlenecks, capacity risks, and scalability limitations.
- Help define and track service-level indicators and objectives, error budgets, or equivalent reliability measures.
- Support backup validation, disaster recovery, failover testing, capacity planning, and operational readiness reviews.
Security and Operational Excellence
- Support vulnerability remediation, operational security controls, and compliance initiatives.
- Create and maintain runbooks, troubleshooting guides, recovery procedures, and operational documentation.
- Partner with development teams to improve application reliability, observability, and supportability.
What you'll need to be successful
Required
- Six or more years of overall IT experience.
- Three or more years performing site reliability engineering for production SaaS workloads in a public cloud, with direct responsibility for their availability and performance. Adjacent experience in cloud operations or DevOps alone does not meet this requirement.
- Demonstrated ownership of the full reliability cycle: instrumenting a service, defining what healthy looks like, responding to incidents, completing root cause analysis, and delivering the engineering work that prevented recurrence.
- Experience supporting customer-facing, business-critical applications.
- Strong troubleshooting, communication, collaboration, and technical documentation skills.
- Candidates must have experience performing Site Reliability Engineering functions for a SaaS application, including hands-on production experience in each of the following:
- Defining and operating against reliability measures such as SLIs, SLOs, and error budgets
- Building observability for a production service: metrics, centralized logging, distributed tracing, dashboards, and application performance monitoring
- Designing actionable alerting and reducing alert noise
- Leading incident response, root cause analysis, and blameless post-incident reviews, and driving the resulting reliability improvements to completion
- Writing automation that eliminates recurring operational work, in PowerShell, Python, Bash, or a comparable language
- Infrastructure as Code and configuration management
- Operating workloads in Microsoft Azure or another major public cloud
Preferred Qualifications
- Windows Server and IIS/.NET application support in production
- SQL Server monitoring and performance troubleshooting, on Azure SQL and SQL Server on virtual machines
- CI/CD pipelines and deployment automation
- Performance testing and capacity planning
- Network traffic management, load balancing, web application firewalls, and content delivery
- Backup, replication, failover, disaster recovery, and recovery testing
- Containers and orchestration technologies
- The Cloud Operations team currently uses Datadog, Azure Application Insights, Elastic, Terraform, Ansible, PowerShell, Azure DevOps, and Cloudflare. Experience with these tools is a plus, but we care more about how candidates have applied comparable technologies than about familiarity with any particular product.
What you can expect from us
We really ARE more than a company! We have a passion for growth and hitting goals, but we want to do it together as a team and enjoy the ride as we go.
WE ARE APPROACHABLE.
Want to talk to a member of management or leadership? Walk up and say hi.
WE TRUST YOU.
We don't have a million rules because we believe you will embrace our Culture Code and make good decisions.
WE EMBRACE TECHNOLOGY.
We love it, sell it, support it, use it, and need it.
WE DRESS COMFORTABLY.
We need you to work hard, not wear a tie.
WE ARE KIND.
Being kind, forgiving, empathetic, and respectful can change your life and everyone around you.
WE ENJOY OUR TEAM.
Use your webcam, collaborate, build relationships, and recognize others for doing a great job. You work a lot; make the most of your time here.
WE VIRTUALIZE EVERYTHING.
We strive for 100% inclusion of our remote teammates.
MCCi is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
About Us:
MCCi is a premier provider of intelligent business process automation solutions and a trusted partner to over 2,000 organizations across state and local government, education, and select commercial entities throughout North America.Through our JustFOIA and JustMeetings platforms, and as the world's largest Laserfiche solution provider, we prioritize client delight. Our solutions suite and commitment to client delight enable us to positively impact the lives of our clients and those they serve.
Cloud Operations Engineer · MCCi