Tools Engineer
- SolarWinds
- npm
- Node.js
- Configuration Management
- AIOps
- Machine Learning
- Incident Management
- AI
- REST API
- Splunk
- SNMP
- PowerShell
- Python
- AWS
- Azure
- GCP
- Windows
- Linux
- SQL
- TCP/IP
- DNS
- DHCP
Overview:
Observability & Enterprise Monitoring Engineer with specialized expertise in SolarWinds platform administration and broader multi-tool observability ecosystems. Working knowledge of OpenText NNMi will be an added advantage. This role will be responsible for the end-to-end administration, optimization, integration, and operational maintenance of enterprise-scale implementation of monitoring solutions (SolarWinds). Responsible for ensuring platform health, automate alert workflows, manage hybrid/cloud monitoring integrations, and collaborate closely with cross-functional infrastructure teams to maintain high availability and performance.
Roles & Responsibilities:
-
Platform Administration & Lifecycle Management (SolarWinds)
-
Core Module Management: Administer and optimize SolarWinds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem.
-
Upgrades & Maintenance: Perform routine and major version updates across platform components; monitor platform health using Active Diagnostics and My Deployment health checks.
-
Polling Infrastructure: Manage, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance across enterprise environments.
-
Database & Backup Operations: Perform operational tasks on the underlying MS SQL Database, manage, schedule, and verify configuration and database backup jobs.
-
Network & Device Observability Operations
-
Discovery & Asset Management: Execute network discoveries, manage node onboarding/offboarding, assign Universal Device Pollers (UnDP), and maintain custom custom attributes and group hierarchies.
-
Configuration Management (NCM): Build and maintain NCM command templates, automate daily startup/running config backups, archive config files, and remediate compliance/transfer failures.
-
Topology & Visualization: Create and maintain dynamic, accurate network topology maps using Network Atlas and modern visual canvases based on operational requirements.
-
Alerting, Dashboarding & ITSM Integration
-
Signal Optimization: Design, tune, and maintain custom Alert Triggers, Actions, and Thresholds to eliminate alert noise and drive actionable alerting.
-
Ticketing & Automation: Configure bi-directional ITSM/ticketing integrations to enable automatic ticket creation, routing, and lifecycle tracking.
-
Reporting & Visibility: Build custom operational and executive Dashboards, Views, and Reports tailored to stakeholder requirements.
-
Incident Support: Monitor alert channels for operational anomalies, troubleshoot lingering telemetry issues, and collaborate with domain teams to drive root cause resolution.
-
AIOps Operations
-
Leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management.
-
Collaborate with cross-functional teams to integrate AI-driven event correlation models and automated remediation workflows into the central monitoring platform.
-
Integration, Vendor Coordination
-
Manage relationships and support escalations with platform vendors.
-
Work on REST API integrations across applications/tools as per requirements.
-
Operational Troubleshooting & Diagnostics
-
Perform deep-dive troubleshooting and root-cause analysis for platform-level performance degradations, engine polling failures, and monitoring agent corruptions.
-
Utilize Active Diagnostics and system telemetry to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion issues.
Required Skills:
-
Multi-tool expertise (SolarWinds, OpenText NNMi, Splunk, etc.)
-
Protocol & Telemetry Knowledge: In-depth understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, NetFlow/sFlow, and Observability (Metrics, Logs, Traces).
-
Automation & API Integration: Good to have skills in PowerShell/Python, and API-driven automation for monitoring workflows.
-
AIOps & Intelligent Automation: Basic understanding of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks.
-
Cloud & Hybrid Observability: Hands-on experience extending platform monitoring into AWS, Azure, or GCP environments.
-
Infrastructure Knowledge
-
System Administration: Intermediate knowledge of Windows and Linux administration.
-
Database: Understanding of SQL/Database concepts and standard query execution.
-
Networking: Good understanding of networking concepts including TCP/IP, DNS, DHCP, Routing and Switching.
-
ITSM: Experience in ITSM processes and operational support.
Tools Engineer ยท Vantage Point Consulting