Site Reliability Engineer — Scale & Resilience for AI Ops

L6

happyrobotMillbrae, CAyesterday
A high-growth AI startup in San Francisco is seeking a Site Reliability Engineer to lead the scaling of operational resilience. In this role, you will own system stability and debugging workflows while tackling complex failures and enhancing proactive operations. Ideal candidates will have over 3 years of experience in debugging production systems, strong problem-solving skills, and familiarity with tools like Datadog and Prometheus. Join a dynamic team dedicated to redefining enterprise operations with cutting-edge AI technology.
Apply now
Apply now

Level

LeadL6

Location

Millbrae, CA

Occupation

Computer Systems Engineers/Architects

Industry

Other Computer Related Services

Posted

yesterday

To get sharper similar jobs, create your profile using the link below.

Create profile
Site Reliability Engineer — Scale & Resilience for AI Ops at...