“We spent 30 years perfecting quality engineering for deterministic software. Now AI has changed the rules. My premise is that LLM evals aren’t a completely new discipline, they’re the natural evolution of quality engineering, where confidence comes from statistical evidence, continuous evaluation, and engineering systems that learn rather than simply pass or fail.
The talk explores how familiar quality engineering concepts such as regression testing, coverage, release gates, observability, and production monitoring map to modern evaluation practices for AI-powered systems. Rather than focusing on a particular framework or tool, it provides a practical engineering mindset for building confidence in probabilistic software, highlighting both the lessons we can carry forward from traditional QE and the new challenges that emerge when software is no longer deterministic.”
“Christopher Hughes is Head of Screening Engineering at LSEG Risk Intelligence, where he leads the engineering behind World-Check with a focus on AI initiatives across Risk Intelligence. He is passionate about AI operating models and why adoption alone does not deliver transformation, arguing that AI pays off through workflow and authority redesign rather than just tool adoption.”
Join our community of testers and start your journey