LLM Evaluation Engineer - AI-Native Development - $175-200,000

L6

saragossaDenver, CO2 days ago
You are building the quality layer that takes generative AI from experiment to enterprise reality at a Private-Equity-backed professional services business. This is not a role where you simply test what someone else built. You are defining what “good” looks like for AI applications across the business. You set the standards. You own the evaluation process. You decide whether models, prompts, and retrieval changes are actually ready to ship. Day to day you are working with models - testing and tuning LLMs to make sure they are accurate, safe, and reliable. You are building evaluation pipelines, refining system prompts, creating golden datasets, monitoring model performance, and catching hallucinations, regressions, and quality issues before they reach users. You are not leaving your technical experience behind, you are applying it to one of the fastest growing areas in AI engineering. LLM evaluation, fine-tuning, red-teaming, prompt optimization, and AI quality, these are the skills that will define how enterprise AI gets built, and you are owning them here. You bring strong experience working with production LLMs and a deep understanding of AI evaluation and quality. You know how to test non-deterministic systems, measure whether they are actually improving, and turn subjective feedback into clear, defensible decisions about what is ready for production. This is a fully remote role Ready to stop hoping your models work and start proving they do?
Apply now
Apply now

Level

LeadL6

Location

Denver, CO

Occupation

Validation Engineers

Industry

Custom Computer Programming Services

Posted

2 days ago

To get sharper similar jobs, create your profile using the link below.

Create profile