UI
Rack Systems Architect, Technical Lead
Ursus, Inc.
- ๐บ๐ธ United States
- On-site
- Staff / Principal
- 6 hours ago
- $208,000 โ $263,000 / year
- AI
- Shopify Liquid
- AWS
- Oracle Cloud
- HPE
- InfiniBand
- Equity
- Pension
6 hours ago
LOCATION: New York, NY; Phoenix, AZ; or Austin, TX
DURATION: FTE Direct Hire
PAY RANGE: $208K โ $263K; offers equity
QUALIFICATIONS:
- 10+ years of experience in rack-scale compute systems, high-performance computing (HPC), AI infrastructure, hyperscale data centers, server hardware, or integrated compute platform development.
- Demonstrated ownership of rack-level system architecture, including electrical, mechanical, thermal, networking, and integration requirements.
- Experience architecting and deploying high-density rack-scale computing environments supporting large accelerator deployments, AI training infrastructure, GPU clusters, or HPC systems.
- Experience developing and managing system-level power budgets, thermal budgets, cooling requirements, rack weight constraints, interconnect architectures, and infrastructure interfaces.
- Experience owning Interface Control Documents (ICDs) and defining boundaries between rack systems, facility systems, networking infrastructure, power systems, and cooling systems.
- Strong understanding of liquid-cooled rack architecture including CDUs, direct-to-chip cooling, thermal distribution systems, facility cooling integration, and serviceability considerations.
- Experience leading multidisciplinary technical teams including Electrical Engineering, Mechanical Engineering, Thermal Engineering, Systems Engineering, and Integration Engineering functions.
- Demonstrated experience making architecture trade-off decisions involving cost, performance, reliability, manufacturability, scalability, cooling, power density, transportability, and serviceability.
- Experience driving design reviews, technical readiness reviews, integration reviews, and engineering decision-making throughout product development lifecycles.
- Experience working directly with accelerator vendors, server vendors, ODMs, hyperscalers, or AI infrastructure providers.
EDUCATION:
- Bachelor's degree in Electrical Engineering, Mechanical Engineering, Systems Engineering, Computer Engineering, or related technical discipline.
PREFERRED QUALIFICATIONS:
- Experience leading NVL-class, HGX-class, GB200-class, B100/B200-class, MI300-class, or equivalent rack programs.
- Deep familiarity with NVIDIA ecosystem architecture, accelerator integration, and GPU cluster deployments.
- Experience supporting hyperscale AI infrastructure environments at organizations such as NVIDIA, Meta, Microsoft, AWS, Google, Oracle Cloud, CoreWeave, Crusoe, Lambda, xAI, Dell, HPE, Supermicro, Quanta, Wistron, Celestica, or Foxconn.
- Experience with rack-level liquid cooling systems including direct-to-chip cooling, immersion cooling, rear-door heat exchangers, facility water systems, and CDU integration.
- Strong understanding of high-speed networking architectures, InfiniBand, Ethernet, optical interconnects, copper interconnects, scale-up fabrics, and scale-out architectures.
- Experience supporting transport qualification, seismic qualification, environmental qualification, or regulatory certification activities.
- Experience developing product roadmaps and aligning infrastructure designs to future accelerator generations.
- Experience supporting manufacturing, integration, rack deployment, system validation, or product lifecycle management activities.
- Advanced degree preferred in Electrical Engineering, Mechanical Engineering, Systems Engineering, Thermal Engineering, or Computer Engineering.
HOW WE OPERATE:
- Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.
- Insane urgency. We drive everything forward as fast as possible.
- Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
- Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
- Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.
ABOUT THE TEAM:
The team owns the rack as a product: the unit of compute that designs, integrates, and deploys into gigawatt-scale AI data centers.
Examples of key problems the team is working on:
- Architect 150kW-class liquid-cooled racks where power, thermal, weight, and interconnect budgets all bind at once and every tradeoff shows up somewhere else.
- Define and hold the interfaces between rack and facility so power, liquid, and network can each evolve without forcing a redesign of the other side.
- Keep the rack roadmap ahead of accelerator vendor roadmaps, so the next generation of dense compute lands in facilities designed before that hardware existed.
- Run a small engineering team with the review rigor of a large one, catching design faults before they reach integration.
POSITION SUMMARY:
- Own the rack as a product for 150kW-class liquid-cooled deployments: architecture, power budget, cooling, interconnect topology, structure, and serviceability, with tradeoffs made explicitly and defended on the record.
- Write and hold the interface control documents between rack and facility, precise enough that power, liquid, and network teams can each move independently without breaking the boundary.
- Lead the rack EE, ME, thermal, and integration engineers, owning the architecture they execute and the technical calls that unblock them.
- Run design reviews where a 7-person team catches what a 70-person team would, so faults die in review rather than in the field.
- Align the rack roadmap with accelerator vendors and the company's facility architecture, so each rack generation lands in a facility built to receive it.
WHAT WE'RE LOOKING FOR:
The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.
- 10+ years in system architecture of rack-scale compute at a hyperscaler, NVIDIA-ecosystem ODM, or AI systems company.
- You've owned a rack as a product end to end: power budget, cooling architecture, interconnect topology, weight, transport, and serviceability.
- You've shipped dense liquid-cooled rack-scale products, NVL-class or equivalent.
- You've led the EE, ME, and thermal engineers who executed your architecture, not just advised them.
- You've traded real things away on a shipping rack and taken the pushback that came with it.
- You write interface control documents that let both sides of a boundary move independently, and you've seen what happens when they don't.
- You hold a defensible view on copper versus optics for scale-up interconnect at 200G+ per lane, and you can argue where the crossover sits.
- Bonus: NVL-class rack programs, transport and seismic qualification, facility-side liquid cooling fluency, or direct engagement with accelerator vendor roadmaps.
BENEFITS SUMMARY:Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate or annual salary only, unless otherwise stated. In addition to base compensation, full-time roles are eligible for Medical, Dental, Vision, Commuter and 401K benefits with company matching.
IND123
BENEFITS SUMMARY:
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate or annual salary only, unless otherwise stated. In addition to base compensation, full-time roles are eligible for Medical, Dental, Vision, Commuter and 401K benefits with company matching.
FAIR CHANCE POLICY:
Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, qualified applicants will be considered for assignment with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation.
Rack Systems Architect, Technical Lead ยท Ursus, Inc.