Site Reliability Engineer — Scale & Resilience for AI Ops
L6
happyrobotMillbrae, CAyesterday
Occupations
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems AdministratorsIndustries
Other Computer Related ServicesComputer Systems Design ServicesCustom Computer Programming ServicesA high-growth AI startup in San Francisco is seeking a Site Reliability Engineer to lead the scaling of operational resilience. In this role, you will own system stability and debugging workflows while tackling complex failures and enhancing proactive operations. Ideal candidates will have over 3 years of experience in debugging production systems, strong problem-solving skills, and familiarity with tools like Datadog and Prometheus. Join a dynamic team dedicated to redefining enterprise operations with cutting-edge AI technology.
Apply now
Level
LeadL6
Location
Millbrae, CA
Occupation
Computer Systems Engineers/Architects
Industry
Other Computer Related Services
Posted
yesterday
To get sharper similar jobs, create your profile using the link below.