EI
Systems/Software Engineer III
eTeam Inc.
- πΊπΈ United States
- Remote
- Senior
- 1 day ago
- $65.00 β $71.60 / hour
- Cordova
- Devops
- Linux
- Slurm
- Azure Cloud
- Azure
- Terraform
- Ansible
- IaC
- GitHub
- Change Management
- ServiceNow
- Okta
- Active Directory
- LDAP
- Splunk
- Datadog
- Confluence
- Linux OS
- Ubuntu
- Identity Management
- NFS
- Git
- Artifactory
- YAML
- Python
- Perl
- Bash
- Jira
1 day ago
Title:Systems/Software Engineer III
Location: Rancho Cordova, CA OR REMOTE (PST & MST States)
Duration: 12 Months
Position Overview
The client is seeking a Senior DevOps Engineer to provide a 12-month contingent engagement supporting our High Performance Computing (HPC) and Electronic Design Automation (EDA) cloud infrastructure team. This role will work directly within the IT Datacenter (ITDC) organization and is expected to operate independently at a senior level with minimal ramp-up time. The ideal candidate brings strong hands-on experience with Linux HPC environments, infrastructure automation, SLURM workload management, and the Azure cloud environment.
Engagement Details
- Position Title: Senior DevOps Engineer β HPC / EDA / SLURM / Azure
- Engagement Type: Contingent Worker (CW)
- Duration: 12 Months
- Location: Remote / Hybrid (Rancho Cordova)
- Department: IT Datacenter Infrastructure (ITDC)
Key Responsibilities
- HPC / EDA Platform Operations
- Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning using Terraform and Ansible.
- Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers.
- Coordinate cross-functionally with EDA/TD NAND teams, storage teams, and IDAM to deliver coordinated platform changes.
- Administer Azure EDA user environment utilizing Thinlinc (VNC).
- Automation & Infrastructure as Code
- Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration.
- Ensure multi-version Ansible playbook compatibility across SLES 15.
- Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories.
- Drive production environment changes through change management workflows using ServiceNow.
- Identity & Access Management
- Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, and SSSD for Linux/HPC environments.
- Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains.
- Validate authentication methods and access behavior across HPC compute and storage environments.
- Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation.
- Monitoring, Logging & Operational Readiness
- Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs.
- Investigate and remediate operational issues in production Linux services (VNC, AutoFS, Datadog, etc.).
- Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence.
Required Skills & Qualifications
- Core Technical Skills
- HPC / EDA Platforms: SLURM, HPC compute/storage administration, EDA infrastructure, datacenter migrations.
- Linux / OS: SLES 15, Ubuntu.
- Provisioning / Automation: Ansible (playbooks, roles, multi-version).
- Identity / Auth: SSSD, Okta, Active Directory, LDAP, cross-domain identity management.
- Storage / Filesystems: NetApp SVM, NFS, AutoFS, RootSquash, storage tier design, IOPS/capacity planning.
- DevOps / Source Control: Git, GitHub, Artifactory, inner-source repository management.
- Monitoring / Logging: Splunk integration, Datadog, operational script hardening, log management.
- Scripting / Languages: Ansible (YAML), Python, Perl (debugging), Bash.
- ITSM / Documentation: ServiceNow (change requests), MOP authoring, Confluence, Jira, technical diagramming.
- Experience Requirements
- 5 years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role.
- Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers.
- Demonstrated experience supporting EDA or scientific computing environments.
- Strong Ansible & Terraform automation skills with production-grade playbook and role development.
- Demonstrated usage and understanding of the Azure cloud compute environment.
- Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, Okta).
- Experience with NetApp or comparable enterprise storage platforms in HPC contexts.
- Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams).
- Strong written and verbal communication skills; capable of coordinating across multiple teams.
Preferred Qualifications
- Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise environment.
- Experience migrating configuration artifacts and binaries to Artifactory.
- Background in semiconductor, storage, or high-tech manufacturing IT environments.
Systems/Software Engineer III Β· eTeam Inc.