C
Datadog SME
CirrusLabs
- ๐บ๐ธ United States
- On-site
- 10 months ago
- Datadog
- Python
- Bash
- PowerShell
- Microservices
- Java
- .NET
- Node.js
- Oracle
- ServiceNow
- CI/CD
- triage
- SQL
- PL/SQL
- Jenkins
- Git
- Azure
- AWS
- IaC
- Terraform
- Ansible
10 months ago
You have an entrepreneurial spirit. You enjoy working as a part of well-knit teams. You value the team over the individual. You welcome diversity at work and within the greater community. You aren't afraid to take risks. You appreciate a growth path with your leadership team that journeys how you can grow inside and outside of the organization. You thrive upon continuing education programs that your company sponsors to strengthen your skills and for you to become a thought leader ahead of the industry curve.
You are excited about creating change because your skills can help the greater good of every customer, industry and community. We are hiring a talented<Job Title>to join our team. If you're excited to be part of a winning team, CirrusLabs (http://www.cirruslabs.io) is a great place to grow your career.
Role Title: Datadog SME
Location: Norfolk, VA / Richmond, VA / Atlanta, GA / Texas / Remote
Job Description
Monitoring Strategy & Implementation Design and deploy Datadog dashboards for both application and database domains.
Configure alerting logic using Datadog monitors, composite alerts, and anomaly detection.
Develop custom scripts (Python, Bash, PowerShell) to support dynamic alerting and data enrichment.
Application Monitoring
Integrate Datadog APM with microservices (Java, .NET, Node.js) and front-end layers.
- Monitor service health, latency, error rates, and throughput.
- Enable distributed tracing and log correlation for root cause analysis.
- Configure synthetic tests and real-user monitoring (RUM) for critical endpoints.
Database Monitoring (Oracle & SQL Server)
Set up Datadog integrations for Oracle and SQL Server to track:
Query performance, blocking sessions, deadlocks.
Connection pool usage, replication lag, backup status.
Tablespace utilization, I/O latency, and cache hit ratios.
Automate alerting for threshold breaches and unusual patterns using scripting.
Scripting & Automation
- Build reusable scripts to:
- Generate dynamic dashboards based on metadata.
- Auto-adjust alert thresholds based on historical trends.
- Integrate Datadog with ServiceNow for incident creation.
- Maintain version-controlled script repositories and CI/CD pipelines for observability assets.
Stakeholder Collaboration
- Engage with application and database teams to gather monitoring requirements.
- Conduct workshops and KT sessions for support teams on dashboard usage and alert triage.
- Partner with SREs to align monitoring with SLIs/SLOs and reliability goals.
Reporting & Governance
- Generate weekly/monthly reliability reports using Datadog analytics.
- Maintain SOPs, runbooks, and RCA documentation.
- Ensure compliance with enterprise monitoring standards and audit readiness.
SkillSet
5+ years in cloud monitoring and reliability engineering.
Proven expertise in Datadog (APM, Infrastructure, Logs, Monitors, Dashboards).
Hands-on experience with Oracle and SQL Server monitoring.
Strong SQL and PL/SQL skills for Oracle and SQL Server.
Proficiency in scripting (Python, Bash, PowerShell) for automation
Experience in configuring complex alerting logic using Datadog monitors.
Understanding of application architectures (Java, .NET, Node.js).
Familiarity with ITSM tools (ServiceNow), CI/CD (Jenkins, Git), and cloud platforms (Azure, AWS).
Excellent communication and documentation capabilities.
Nice-to-Have Skills:
Datadog certification.
Knowledge of healthcare domain and compliance standards.
Experience with CI/CD tools and infrastructure as code (Terraform, Ansible).
Datadog SME ยท CirrusLabs