Senior RL Researcher: Async Systems for Frontier Models

L5

thinking machines labMillbrae, CAyesterday
Thinking Machines is seeking a researcher to own the boundary between reinforcement learning algorithms and the underlying systems. The role focuses on high-training-compute, long-horizon RL, with a center of gravity around asynchronous RL and integration with inference constraints. You will co-design the RL recipe and the systems, advance async RL, and run frontier-scale RL end-to-end. Candidates should have strong Python and DL framework experience, and a PhD or equivalent research background
Apply now
Apply now

Level

SeniorL5

Location

Millbrae, CA

Occupation

Computer and Information Research Scientists

Industry

Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)

Posted

yesterday

To get sharper similar jobs, create your profile using the link below.

Create profile