AI & HPC Infrastructure Engineer
franklin fitchLas Vegas, NV
AI & HPC Infrastructure Engineer
L6
franklin fitchLas Vegas, NVtoday
Occupations
Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsComputer and Information Research ScientistsIndustries
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesOther Computer Related ServicesSenior AI & HPC Infrastructure Engineer Remote (US) | $150,000 - $170,000 Base + BonusOur client is a highly respected global consulting and research organization that supports leading commercial enterprises, public sector institutions, and professional services firms. As part of a significant investment in AI infrastructure, they are expanding their internal high performance computing (HPC) and GPU capabilities to support next-generation machine learning and large language model (LLM) initiatives. This is an exciting opportunity to join a small, highly skilled infrastructure team at a pivotal stage of growth. The successful candidate will play a key role in scaling a GPU environment from 8 to 32 NVIDIA H200 GPUs while helping shape the organization's long-term AI and HPC strategy. The role is heavily project-focused, with approximately 85% dedicated to engineering, architecture, and platform development activities, and a smaller proportion supporting operational needs.
The Opportunity:
You'll work at the intersection of traditional HPC, research computing, and modern AI infrastructure, supporting both analytical workloads and large-scale model training environments.
Key responsibilities include:
- Designing, deploying, and maintaining GPU-accelerated computing infrastructure
- Supporting large-scale AI/ML and LLM training environments
- Managing Linux-based HPC clusters and associated services
- Administering and optimizing parallel file systems, particularly IBM Spectrum Scale (GPFS)
- Managing NVIDIA GPU platforms, CUDA, cuDNN, NCCL, and related tooling
- Supporting resource scheduling through SLURM and related technologies
- Performance tuning for distributed and multi-GPU workloads
- Building automation, monitoring, reporting, and operational tooling
- Collaborating with researchers, data scientists, and technical stakeholders to translate business requirements into infrastructure solutions
- Evaluating emerging AI infrastructure technologies and recommending future platform enhancements
Required Experience:
- We're particularly interested in candidates who combine deep infrastructure expertise with an understanding of how end-users consume HPC resources.
- Essential SkillsStrong Linux systems administration background
- Experience supporting HPC or research computing environments
- Hands-on NVIDIA GPU infrastructure experience
- Strong experience with GPFS / IBM Spectrum Scale
- Familiarity with job schedulers such as SLURM, LSF, or similar
- Experience supporting distributed compute environments
- Ability to lead projects independently with minimal oversight
- Strong troubleshooting experience across compute, storage, networking, and hardware layers
- Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders
- Highly Desirable
- Experience supporting AI/ML infrastructure or LLM platforms
- Kubernetes and container orchestration experienceMLOps tooling exposure (MLflow, Kubeflow, etc.)Experience tuning large model training and inference workloads
- Bright Cluster Manager
- AnsibleDocker, Apptainer, or SingularitySAS platform experience alongside GPU/HPC expertise
Ideal Background: This role is particularly well suited to someone who:
- Has approximately 5-10 years of relevant infrastructure experience
- Is currently operating at a strong mid-level and ready for a senior step forward
- Can own and deliver significant technical projects independently
- Comes from an HPC, research computing, higher education, government, scientific computing, or technical enterprise environment
- Enjoys balancing infrastructure engineering with emerging AI technologies
- Team & Culture
You'll join a close-knit team of experienced infrastructure specialists covering HPC operations, automation, applications, and platform engineering. The group operates in a highly collaborative remote-first model, with team members distributed across the United States.
A notable area of planned growth is AI and LLM infrastructure expertise, making this a high-visibility hire with significant opportunity for progression and influence.
If you're an HPC infrastructure engineer looking to move into a highly visible AI-focused environment while remaining close to the hardware, architecture, and platform engineering side of the stack, we'd love to hear from you.
Apply now
Level
LeadL6
Location
Las Vegas, NV
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
today
To get sharper similar jobs, create your profile using the link below.