Tier 1 NOC Engineer (HPC & Bare-Metal Infrastructure)

L5

axe computeMiami Springs, FL3 days ago
We are seeking a technical, detail-oriented Tier 1 NOC Engineer to join our 24/7 Operations team. In this role, you will serve as the first line of defense for our global high-performance computing (HPC) bare-metal infrastructure, high-speed InfiniBand fabrics, and parallel storage systems. You will actively monitor cluster telemetry, perform rapid triage on live alerts, execute initial Linux OS and hardware troubleshooting, and coordinate on-site remote-hands data center technicians to maintain sub-15 minute SLA response targets. Core Responsibilities1. 24/7 Monitoring & Alert Triage Actively monitor global HPC clusters, bare-metal GPU/CPU nodes, network fabrics, and storage arrays using Grafana, Prometheus, and internal telemetry tools. Validate, categorize, and perform initial incident triage for automated alerts and customer ticket submissions. Ensure strict adherence to response time SLAs for critical hardware, OS, and network outages.2. Hardware & Infrastructure Troubleshooting Perform initial diagnostics on x86 bare-metal servers, inspect Out-of-Band (OOB) management interfaces (IPMI, Redfish, iDRAC/iLO), and review hardware logs. Identify component-level hardware failures, including GPU/PCIe errors, DIMM faults, drive dropouts, power supply failures, and thermal anomalies. Open, track, and direct vendor dispatch tickets (e.g., Supermicro, Dell, HPE, NVIDIA) for physical hardware replacements.3. Linux OS & Network Isolation Access non-responsive or degraded hosts via serial-over-LAN/SSH to inspect system logs (dmesg, syslog, journalctl).Check basic Layer 1–Layer 3 network connectivity, interface link states, IP assignment, and optical transceiver signal levels. Escalate complex fabric routing, InfiniBand subnet manager, or parallel storage issues to Tier 2/3 HPC and Network Engineers with clean, structured diagnostic notes.4. Remote-Hands Coordination & Dispatch Direct on-site data center technicians for physical cable pulls, server reboots, drive swaps, and component replacements. Track hardware RMA lifecycles from fault identification to physical repair verification and host re-commissioning.5. Incident Documentation & Runbooks Maintain detailed, real-time ticket notes documenting all troubleshooting steps, root causes, and resolutions. Follow established operational runbooks and assist in updating Standard Operating Procedures (SOPs) as new issues emerge.
Key Qualifications
Experience: 1–3 years in a Network Operations Center (NOC), data center technical support, or system administration role. Linux Proficiency: Solid hands-on experience navigating and troubleshooting Linux operating systems (Ubuntu, RHEL, Rocky Linux) via command line. Hardware Fundamentals: Clear understanding of enterprise x86 server architecture, IPMI/BMC management, RAID arrays, PCIe topologies, and component-level diagnostics. Networking Basics: Solid grasp of TCP/IP networking, DNS, VLANs, subnetting, and basic physical layer troubleshooting (copper, fiber optics, transceivers).Communication & Shift Work: Excellent written and verbal communication skills; comfortable working rotational, night, or weekend shifts in a 24/7 environment.
Nice-to-Have / Preferred Skills: Exposure to High-Performance Computing (HPC) environments or large-scale GPU clusters. Basic familiarity with InfiniBand fabrics, RDMA/RoCE, or parallel storage systems (Weka, Lustre, GPFS).Experience with monitoring and ticketing systems like Grafana, Prometheus, Datadog, Jira, or Service Now. Python or Bash scripting capabilities for task automation.
Apply now
Apply now

Level

SeniorL5

Location

Miami Springs, FL

Occupation

Network and Computer Systems Administrators

Industry

Computer Systems Design Services

Posted

3 days ago

To get sharper similar jobs, create your profile using the link below.

Create profile