Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
Beam logo

GPU Cluster Infrastructure Engineer

Beam
  • 🇺🇸 United States
  • Hybrid
  • Manager or above
  • 9 hours ago
  • $10,500 – $18,000 / month
  • Beam
  • AI
  • Python
  • React.js
  • TypeScript
  • Snyk
  • Acceptance Testing
  • InfiniBand
  • Fabric
  • Node.js
  • triage
  • Ansible
  • Prometheus
  • Grafana
  • Equity
  • Learning budget
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Beam

AI-Native Cloud Platform

Tech

Our infra code is mostly Go, our backend APIs are Python, and our dashboard is React/Typescript. We work very closely with customers – they’re all in a communal Slack channel with us, so it’s important that you’re interested in interacting with them (luckily, our customers are all developers).

The role

**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.# **About the Role**We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.* Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.* Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.* Stand up and validate high-performance storage alongside vendor teams.* Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.* Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.* Write runbooks, as-builts, and remote-hands procedures.* Provide escalation support after go-live and help our team ramp up.**Skills & Experience*** You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.* Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.* GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.* Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.* Equally effective on the data center floor and remotely, including directing colo remote hands.* You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.* Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.# Benefits* Competitive salary and meaningful equity* Join a fast-growing pre-series A company at the ground floor* Health, dental, and vision benefits with 90% coverage for you and 50% for dependents* Opportunities to participate in events across the cloud native community* Fitness stipend, learning budget, and much, much more

GPU Cluster Infrastructure Engineer · Beam

Auto apply with Likeremote