AI Evaluation Engineer
Build the test systems that show whether production AI agents are getting better or quietly breaking.
What you'll do
- Turn real production cases into representative golden datasets.
- Design deterministic checks and LLM judges, then validate them against human labels.
- Wire evaluation suites into delivery pipelines so regressions block releases.
- Analyze failure clusters and translate them into measurable engineering work.
What you bring
- Strong Python and software testing fundamentals.
- Hands-on experience evaluating LLM applications, agents, or retrieval systems.
- Comfort with experiment design, metrics, and imperfect real-world data.
- Clear written communication in an async remote environment.
Experience with tracing platforms, prompt testing, or human annotation workflows.