Head of AI Infrastructure
L7
techshackNew York, NY2 days ago
$275,000–$350,000
Occupations
Computer and Information Systems ManagersComputer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsIndustries
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesComputer Systems Design ServicesComputer Facilities Management ServicesHead of AI Infrastructure New York City | Hybrid / Onsite We are hiring a Head of AI Infrastructure to lead the architecture, deployment and operation of a rapidly scaling GPU and HPC platform. This is a senior technical leadership role with end-to-end ownership across compute, networking, cluster software, workload orchestration and production operations. You will set the technical direction for the infrastructure function, remain close to the technology and help define how the next generation of large-scale GPU systems are designed, deployed and operated.
Responsibilities:
Own the technical strategy and architecture for large-scale GPU and HPC infrastructure supporting demanding AI training and inference workloads. Design and scale modern NVIDIA GPU environments, including HGX/DGX platforms and H100, H200 and Blackwell-generation systems. Lead infrastructure architecture across Infini Band, RoCE, RDMA, Ethernet, NCCL, NVLink/NVSwitch and high-performance storage. Establish standards across bare-metal provisioning, firmware, cluster deployment, burn-in, benchmarking, validation and production acceptance. Own cluster performance, reliability and availability across GPU compute environments. Drive improvements across GPU utilisation, node health, telemetry, observability, failure detection and automated remediation. Lead workload orchestration across technologies such as Slurm, Kubernetes and containerised environments. Develop an automation-first approach to provisioning, cluster operations, diagnostics and infrastructure management. Work directly with datacentre operators, OEMs, networking vendors, infrastructure suppliers and engineering teams. Build and lead a high-performing AI infrastructure team across GPU systems, HPC, networking and platform engineering. Define technical standards, engineering processes and operational best practices as the infrastructure estate scales.
Experience:
We are looking for someone with deep experience building and operating large-scale infrastructure across several of the following areas: NVIDIA HGX/DGX and H100/H200/Blackwell-class GPU platforms Infini Band, RoCE, RDMA and high-performance EthernetNCCL, NVLink, NVSwitch, CUDA and DCGMSlurm, Kubernetes, Docker and Linux systems engineering Bare-metal provisioning and infrastructure automationHPC benchmarking, GPU validation and production acceptance High-performance and distributed storage Datacentre and rack-scale infrastructure Python, Go, Bash or Ansible Distributed AI training and inference environments Background: You may currently be working as a Head of Infrastructure, Director of Infrastructure, GPU Infrastructure Lead, Head of HPC, Principal HPC Engineer, Staff or Principal Infrastructure Engineer, AI Infrastructure Architect or senior technical leader within a neocloud, hyperscaler, AI lab or HPC environment.
Compensation: $275,000–$350,000 base salary plus bonus and equity, depending on experience.
Apply now
Level
ManagerL7
Salary
$275,000–$350,000
Location
New York, NY
Occupation
Computer and Information Systems Managers
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
2 days ago
To get sharper similar jobs, create your profile using the link below.