Matcha
Beam logo

GPU Cluster Infrastructure Engineer

Beam

See similar roles
Apply

We’re bringing AI to welfare services, trusted by over 100 government partners. Join us 🚀

$125K Seed Round raised
$11k - $18kCONTRACTORRemote · United States349 employeesPosted Oct 5
nvidia hgx/dgx cluster operationsinfiniband fabric managementsubnet managementufmgpu node bring-up and burn-infirmware and bmc/redfishdcgmnccl testingpxe and imagingxid error triageparallel storagehardware and fabric troubleshootingtelemetry and observabilityrunbook and as-built documentationansible automationprometheus and grafana

Remote startup roles in your inbox

Matcha reads all job descriptions to surface the handful that actually matter in a zero noise email. Simple by design.

Describe your next role, cut the noise

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

About the Role

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

Skills & Experience

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.
Benefits
  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

More remote jobs at Beam

Similar remote jobs