Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
M

Technicians (Lab Technicians) - Datacenter Technician 1

Mindlance
🇺🇸 United States
On-site
Mid level
1 day ago
  • Shopify Liquid
  • AI
  • Fabric
  • plumbing
  • LOTO
  • Change Management
  • InfiniBand
  • ASHRAE
  • HPE
  • ITIL
  • ServiceNow
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
CONTRACT WORKER — JOB DESCRIPTION
Liquid-Cooled Datacenter Infrastructure Technician
Position Type: Contract Worker (CW) Department: Datacenter Operations
Experience Required: 4–5 Years (Minimum) Work Schedule: Onsite, 5 days a week Plus weekend support
Environment: High-Density AI Datacenter Clearance: Must pass background check
POSITION SUMMARY We are seeking an experienced and highly motivated Contract Worker to support the installation, maintenance, and operational support of next-generation liquid-cooled AI rack infrastructure in a high-density datacenter environment. The ideal candidate will have direct hands-on experience with Direct Liquid Cooling (DLC) systems, high-density compute platforms — including Instinct GPU compute trays and EPYC server nodes — and the physical plant systems (CDUs, manifolds, busbar, networking fabric) required to sustain racks operating at up to 246 kW per rack.
This role supports Helios-class rack-scale AI infrastructure: open-standard double-wide (ORW) racks housing compute trays (4× Instinct MI455X GPUs + 1× EPYC Venice CPU each) and switch trays interconnected via a UALink / Ethernet fabric. The technician performs hands-on tray-level service, cooling system operations, firmware cycling, and physical cabling to sustain continuous availability of production AI training and inference workloads. KEY RESPONSIBILITIES Liquid Cooling Operations
  • Operate, monitor, and maintain Coolant Distribution Units (CDUs) including flow rate management (target: ~385 L/min per rack system), temperature setpoints, and alarm response.
  • Install, replace, and pressure-test blind-mate quick-disconnect (QD) connectors on compute and switch trays without interrupting adjacent rack cooling circuits.
  • Perform leak detection, containment, and remediation in accordance with site EHS protocols.
  • Maintain secondary loop plumbing (supply/return manifolds, flexible hose assemblies, isolation valves); coordinate with facilities for primary loop support.
  • Support rack-level thermal assessments; contribute to coolant quality monitoring (pH, conductivity, biocide levels).
High-Density Rack & Server Hardware
  • Install and service Helios-class compute trays (~77 kg each; 576 differential connections) using approved lift equipment and torque specifications.
  • Install and service switch trays (1,728 connections; ~310 kg insertion force) using lever-assist handles per OEM guidance.
  • Replace field-replaceable units (FRUs): Instinct MI450/MI455X GPU modules, HBM4 memory, EPYC CPU+DIMM assemblies, NVMe drives, power modules, and fan trays.
  • Execute rack-level hot-swap operations on busbar and power shelf components following LOTO and high-voltage DC (HVDC) safety procedures.
  • Cable-manage high-speed optical interconnects (QSFP-DD / OSFP) and DAC/Client assemblies for UALink, Ethernet, and out-of-band management fabrics.
Networking & Fabric Infrastructure
  • Patch and manage fiber/copper connections within the Broadcom Tomahawk 6-based switch fabric supporting single-hop, multi-plane GPU-to-GPU interconnect.
  • Validate optical power budgets and verify high-speed links following tray moves, adds, or changes.
  • Coordinate with network engineering for fabric topology changes and UALink / UEC configuration validation.
Monitoring, Diagnostics & Firmware
  • Use BMC/Redfish/IPMI tools to monitor GPU and CPU health, review system event logs, and identify hardware faults proactively.
  • Execute firmware update workflows (BIOS/UEFI, BMC, NIC, GPU firmware) following approved change management processes.
  • Run ROCm-based diagnostic utilities (rocm-smi, GPU benchmarks) to validate compute readiness post-maintenance.
  • Interface with DCIM and Client platforms (e.g., Schneider EcoStruxure, Motivair CDU telemetry) to track PUE, coolant flow, and rack-level power draw (up to 246 kW/rack).
Documentation & Compliance
  • Maintain accurate asset inventory, maintenance logs, and incident records in the CMDB/ticketing system.
  • Follow EHS, physical security, and datacenter access procedures at all times.
  • Contribute to runbooks, SOP updates, and lessons-learned documentation after significant incidents or deployments.
  • Participate in on-call rotation for after-hours critical hardware incidents.
REQUIRED SKILLS & QUALIFICATIONS — SKILLS MATRIX
Skill Area Details Status
Liquid Cooling Systems Direct Liquid Cooling (DLC), CDU operation, rear-door HX, manifold & distribution plumbing Required
Blind-Mate Quick-Disconnects Installation, replacement, and leak-testing of QD connectors under operational rack conditions Required
High-Density Rack Platforms OCP Open Rack / ORW form factor ( Helios-class); double-wide rack mechanics up to 246 kW Required
GPU Compute Infrastructure Instinct MI450/MI455X or NVIDIA equivalent; OAM/SXM mezzanine hardware; HBM4 awareness Required
EPYC Server Hardware EPYC (SP5/SP7) platforms; DIMM installation; BIOS/UEFI familiarity Required
Electrical Safety (HV DC) Hot-plug busbar systems, 48V/400V DC bus awareness, LOTO procedures Required
Fiber & High-Speed Networking QSFP-DD/OSFP optics, DAC/Client cable mgmt, fabric interconnect (UALink, InfiniBand, RoCE) Required
Physical Infrastructure Rack installation, torque specs, cable management, PDU work, grounding Required
BMC / IPMI / Redfish Out-of-band management, firmware updates, power cycling, event log review Required
ROCm Software Stack Basic familiarity: rocm-smi, GPU health checks, diagnostic tooling Preferred
OCP / Open Standards OCP Open Rack, UALink, UEC familiarity; open telemetry standards Preferred
DCIM / Client Tools Schneider EcoStruxure, Motivair, Vertiv, or equivalent DCIM / facility monitoring Preferred
Certifications BICSI RCDD/DCIS, CompTIA Server+, ASHRAE Class A4/W5 awareness, vendor certs ( , HPE, Dell) Plus
EXPERIENCE REQUIREMENTS Minimum:4–5 years in a datacenter infrastructure or datacenter operations (DCO) technician role
Must Have
  • 4+ years hands-on experience with Direct Liquid Cooling (DLC) systems in production datacenter environments, including CDU operation and coolant loop maintenance.
  • 4+ years experience with enterprise server hardware installation and break/fix servicing at component level (CPU, DIMM, GPU, storage, NIC).
  • Demonstrated experience with high-density rack platforms (≥20 kW/rack) including cabling, power distribution, and liquid-cooling integration.
  • Proven ability to follow and enforce LOTO, EHS, and physical security procedures in a critical facility environment.
  • Experience with out-of-band server management (BMC, IPMI, Redfish, iDRAC, iLO, or equivalent).
  • Familiarity with high-speed fiber optics (LC/MPO, QSFP-DD/OSFP) including cleaning, inspection, and patch panel management.
Strongly Preferred
  • 2+ years experience with GPU compute infrastructure ( Instinct, NVIDIA HGX, or equivalent OAM/SXM form factors).
  • Experience with EPYC server platforms (1P/2P) including BIOS, memory population, and diagnostic procedures.
  • Familiarity with OCP Open Rack / ORW standards and open datacenter architectures.
  • Exposure to AI or HPC datacenter deployments with cluster-scale interconnects (InfiniBand, Ethernet RoCE, or UALink).
  • Working knowledge of ROCm software stack for hardware validation tasks.
Nice to Have
  • BICSI RCDD, DCIS, or equivalent datacenter infrastructure certification.
  • CompTIA Server+, Linux+, or similar vendor/industry certifications.
  • Experience with Schneider Electric EcoStruxure, Motivair, or Vertiv CDU/DCIM platforms.
  • Familiarity with change management frameworks (ITIL) and CMDB tooling (ServiceNow or equivalent).
PHYSICAL REQUIREMENTS
  • Ability to lift and maneuver equipment up to 50 lbs (23 kg) unassisted; team-lift required for compute trays (~77 kg) and switch trays (~120 kg with tooling).
  • Ability to stand, bend, kneel, and work in confined rack spaces for extended periods.
  • Comfort working in raised-floor and hot/cold-aisle-contained datacenter environments.
  • Manual dexterity sufficient to handle fine connectors, optical fibers, and small form-factor electronics.
WORK ENVIRONMENT This position is on-site at a high-density AI datacenter facility operating Helios-class rack-scale infrastructure. Racks operate at up to 246 kW with fully liquid-cooled compute and switch trays. The environment involves high electrical density, pressurized coolant loops, heavy mechanical components, and stringent physical security controls. Candidates must be comfortable in this environment and committed to strict adherence to safety protocols.
Typical schedule is Monday–Friday with participation in an on-call rotation for critical-path incidents outside standard business hours. Shift work may be required during large-scale deployments or datacenter expansions. ABOUT THIS ENGAGEMENT This is a contract worker (CW) engagement. The position does not carry employee benefits or implied conversion guarantees unless separately agreed in writing. The CW will be engaged through an approved staffing agency and will work under the direction of the datacenter operations team. All work will be performed on-site at the designated datacenter facility.
Confidential — For Staffing Agency Use Only | Liquid-Cooled DC Infrastructure Technician (CW) | Rev 1.0 | Sept 2026

Technicians (Lab Technicians) - Datacenter Technician 1 · Mindlance

Auto apply with Likeremote