
Systems Administrator Lead (Second Shift)
5C Data Centers USA, Inc.
Leading data centers for hyperscalers, cloud, and enterprises; fast, sustainable.
Remote startup roles in your inbox
Matcha reads all job descriptions to surface the handful that actually matter in a zero noise email. Simple by design.
Describe your next role, cut the noise
5C DATA CENTERS
Join the Future of Digital Infrastructure
Are you a passionate Systems Administrator Lead (Second Shift) looking to make a meaningful impact? We're building the next generation of digital infrastructure powering hyperscalers, AI innovation, and high-performance computing across North America.
JOB TITLE
Systems Administrator Lead (Second Shift)
DEPARTMENT
Infrastructure Operations
LOCATION
Remote (Ohio, USA)
WORK ARRANGEMENT
Remote
SALARY RANGE
$115,000 – $145,000
Role Summary
As a Systems Administrator Lead covering the Second Shift (4PM– 1AM EST), you will support, maintain, troubleshoot, and optimize large-scale HPC and AI infrastructure across our data center and cloud environments.
This role is primarily focused on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management. This position provides critical overnight coverage, ensuring infrastructure issues are resolved quickly outside standard business hours.
You will serve as a senior technical escalation point for complex infrastructure incidents affecting GPU clusters, compute nodes, networking, storage, and supporting management services, working closely with Network Engineering, Data Center Operations, Deployment Engineering, vendors, and customer technical teams to restore service and improve the reliability of production HPC environments.
How We Work at 5C
Our core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great.
What You Will Do
HPC Infrastructure Operations
- Administer, maintain, and troubleshoot Linux-based HPC and AI compute environments, including large-scale GPU clusters built on NVIDIA HGX, DGX, or equivalent accelerated computing platforms.
- Diagnose hardware, OS, driver, firmware, networking, and storage-related failures, and perform root-cause analysis to develop corrective actions for recurring issues.
- Serve as a senior escalation point for complex incidents affecting production compute infrastructure, troubleshooting degraded or unstable compute/GPU nodes to restore service.
- Support infrastructure across heterogeneous environments - bare-metal, virtualized, containerized, and cloud-hosted - and maintain operational documentation, runbooks, and shift-handoff notes.
GPU and Accelerated Computing Systems
- Install, configure, validate, and troubleshoot NVIDIA GPU drivers, CUDA components, firmware, and supporting system software using tools such as Nvidia Smi, DCGM, and NVIDIA Fabric Manager.
- Investigate GPU Xid errors, NVLink/NVSwitch faults, PCIe errors, ECC events, GPU resets, and thermal or power-related issues.
- Diagnose communication and performance issues involving NCCL, GPUDirect RDMA, PCIe topology, NUMA placement, and multi-GPU systems, validating GPU-to-GPU, GPU-to-network, and GPU-to-storage communication after repairs.
- Coordinate replacement and RMA activities for GPUs, GPU trays, baseboards, NVSwitch components, system boards, and NICs.
Linux Systems Administration
- Administer enterprise Linux distributions across the production fleet, troubleshooting boot failures, kernel issues, systemd services, filesystems, and OS performance.
- Analyze system and kernel logs (journalctl, dmesg, lspci, dmidecode, ipmitool, ethtool, ss, sar, perf) and manage kernel modules, device drivers, DKMS packages, and OS updates.
- Support Linux networking, storage mounts, authentication, and security controls, and investigate OS crashes, kernel panics, and out-of-memory conditions.
Server Hardware, Firmware, and Out-of-Band Management
- Support enterprise server platforms from vendors such as Dell, HPE, ASUS, Supermicro, and NVIDIA, troubleshooting processors, memory, PCIe devices, NICs, storage controllers, and other hardware components.
- Manage BIOS, BMC, CPLD, NIC, GPU, switch, and drive firmware, and perform remote console troubleshooting, power cycling, and log bundle generation.
- Partner with on-site Data Center technicians on component replacement, cabling validation, break-fix activity, and post-repair testing.
High-Performance Networking
- Troubleshoot Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments, including link state, VLAN/bonding issues, MTU mismatches, packet loss, and RDMA/GPUDirect RDMA connectivity.
- Collaborate with Deployment Engineering to troubleshoot fabric instability, link degradation, congestion, and node-level connectivity issues.
Storage and Filesystem Support
- Troubleshoot server connectivity to parallel, distributed, and enterprise storage platforms such as NFS, Lustre, BeeGFS, GPFS/Spectrum Scale, and Ceph, diagnosing mount failures, throughput degradation, and capacity issues.
- Validate storage network paths and coordinate with storage and networking teams during service-impacting events
Incident, Problem, and Change Management
- Lead or support incident response for critical infrastructure events, participating in technical bridges and providing clear status updates during outages.
- Create detailed timelines, root-cause analyses, and corrective-action plans, maintaining accurate records in Jira Service Management or similar ITSM platforms.
Leadership and Collaboration
- Mentor junior and mid-level Systems Administrators and HPC Engineers, providing technical guidance during complex troubleshooting and infrastructure changes.
- Promote consistent troubleshooting methods, documentation standards, and operational discipline across shift handoffs.
What You Bring
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
- 7+ years of experience in Linux systems administration, HPC/AI infrastructure or data center operations, including hands-on troubleshooting in production environments.
- Hands-on experience supporting enterprise server hardware, bare-metal compute systems, and large-scale NVIDIA GPU infrastructure or accelerated computing platforms.
- Strong understanding of server hardware, firmware, BIOS, BMC, PCIe, NUMA, memory, storage, and network interfaces, with experience diagnosing failures using OS logs, BMC logs, and vendor diagnostics.
- Experience with NVIDIA drivers, CUDA compatibility, and GPU-management tools such as DCGM, Fabric Manager, or nvidia-smi.
- Working knowledge of high-performance networking (InfiniBand, RDMA, RoCE, NVIDIA/Mellanox adapters) and out-of-band management tools (Redfish, IPMI, iDRAC, or iLO).
- Proficiency in one or more automation/scripting languages (Python, Bash, or Ansible) and familiarity with observability platforms such as Prometheus, Grafana, or Datadog.
- Proven ability to troubleshoot across hardware, firmware, OS, network, storage, and cluster-management layers, backed by strong documentation and communication skills.
- Willingness to work a fixed, recurring shift and participate in scheduled on-call escalation.
Why Join 5C Data Centers?
At 5C, we believe great people build great companies. You'll build a rewarding career while helping shape the future of digital infrastructure - one of the fastest-growing industries in the world.
Career Growth
Build a rewarding career in one of the world's fastest-growing industries.
Industry Leadership
Help power the infrastructure behind AI and high-performance computing.
Entrepreneurial Culture
Your ideas matter - we empower employees to help shape our future.
Comprehensive Rewards
Competitive pay plus meaningful, lasting impact on the work you do.
Life at 5C
We're more than a workplace - we're a team of builders, innovators, and problem-solvers united by a shared purpose: creating infrastructure that powers the technologies transforming our world. Your voice matters here, and we encourage fresh ideas at every level.
Ready to Apply?
If this opportunity sounds like the right fit for you, we'd love to hear your story. Apply today and discover what your future could look like at 5C Data Centers.
5C Data Centers is an equal opportunity employer.
We celebrate diversity and are committed to creating an inclusive environment where everyone can thrive. 5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.
#LI-LS1