Staff Software Engineer

L6

usa tech recruitSeattle, WA2 days ago
Staff Software Engineer – Fleet Management A rapidly growing US-based technology company specialising in GPU cloud infrastructure is looking for a Staff Software Engineer to join its infrastructure engineering team. The ideal candidate will have extensive experience building distributed systems, infrastructure automation and production platforms, with strong Python skills and experience managing large-scale compute, datacenter or hardware infrastructure. The key skills required for the Staff Software Engineer: Extensive experience building and operating distributed systems in production Strong Python software engineering skills Experience with workflow orchestration, state machines or event-driven architectures Experience building reliable automation for infrastructure provisioning and lifecycle management Strong understanding of idempotency, retries, checkpoints, recovery and failure handling Experience with bare-metal infrastructure, server provisioning or hardware lifecycle management Experience with BMC, IPMI, Redfish, PXE, MAAS, Ironic or NetBox is highly desirable Experience with GPU infrastructure, AI clusters, HPC or datacenter environments is advantageous Strong Kubernetes, Terraform, Ansible or cloud infrastructure experience Experience with networking and network automation, particularly at datacenter scale Experience building health monitoring, validation, observability and automated remediation systems Experience integrating infrastructure platforms, APIs, monitoring systems, credential stores and DCIM tools Strong understanding of production reliability, scalability and security Staff/Principal-level technical leadership with experience influencing architecture across teams Experience using modern AI-assisted development tools such as Claude, Cursor or similar is desirable The role will involve building and owning a Fleet Management platform responsible for provisioning, testing, monitoring and remediating GPU servers and network infrastructure at scale. You will design Python-based workflows covering the complete hardware lifecycle, from device enrolment and BMC configuration through provisioning, multi-day burn-in, health monitoring and automated remediation. The platform will need to support thousands of concurrent workflows, with strong guarantees around reliability, auditability, idempotency, resumability and failure recovery. You will work closely with infrastructure, hardware, networking and operations teams to build highly scalable systems powering next-generation AI infrastructure. Key words: Staff Software Engineer / Principal Software Engineer / Infrastructure Engineer / Platform Engineer / Distributed Systems / Python / Workflow Orchestration / Event-Driven / State Machines / Fleet Management / Infrastructure Automation / Bare Metal / GPU / AI Infrastructure / Datacenter / Kubernetes / Terraform / Ansible / MAAS / Ironic / IPMI / BMC / Redfish / PXE / NetBox / Hardware Lifecycle / Provisioning / Network Automation / Observability / SRE / Automated Remediation / HPC / NVIDIA / AI(By applying to this role you understand that we may collect your personal data and store and process it on our systems. For more information please see our Privacy Notice (https://eu-recruit.com/about-us/privacy-notice/)
Apply now
Apply now

Level

LeadL6

Location

Seattle, WA

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

2 days ago

To get sharper similar jobs, create your profile using the link below.

Create profile