All videos
0:00 / 0:00
research

Diego Fajardo - No single test is enough

Cohere16 June 2026Watch on YouTube

Description

How do we know a model is actually ready for high-stakes use? In healthcare and life sciences, that question gets complicated fast. A model can look strong on one task, weak on another, and still surprise you when the stakes become real. That makes the real problem bigger than evaluation alone. It is about understanding what a model can do, where it breaks, what different kinds of evidence really tell us, and how we can move from isolated results to a more complete picture of readiness. This talk explores that broader challenge, from benchmarks and expert review to interpretability and more realistic task-based testing, and asks what it would take to evaluate models in a way that actually matches how they will be used. Diego Fajardo leads evaluation work at Lumos, a startup focused on improving the real-world performance and safety of AI models and agents in healthcare and life sciences. His work spans benchmark design, expert evaluation, use of LLMs as judges, and interactive testing approaches to better understand model performance in complex settings. A major part of his current work focuses on AI patients: synthetic interactions designed to make benchmarking more realistic and scalable for patient-facing agents. This session is brought to you by the Cohere Labs Open Science Community - a space where ML researchers, engineers, linguists, social scientists, and lifelong learners connect and collaborate with each other. We'd like to extend a special thank you to Alif Munim and Abrar Frahman, Leads of our AI Safety and Alignment group for their dedication in organizing this event. If you’re interested in sharing your work, we welcome you to join us! Simply fill out the form at https://forms.gle/ALND9i6KouEEpCnz6 to express your interest in becoming a speaker. Join the Cohere Labs Open Science Community to see a full list of upcoming events (https://tinyurl.com/CohereLabsCommunityApp).

What you'll learn

  • No single test result or benchmark is sufficient to determine whether an AI model is ready for high-stakes healthcare applications.
  • Evaluating healthcare AI requires combining multiple perspectives: benchmarks, expert review, interpretability, and realistic testing.
  • Models can perform strongly on one task and poorly on another, which means you need to understand where the model breaks and under which conditions.
  • Interactive testing approaches and AI patients (synthetic interactions) make benchmarking more realistic and scalable for patient-facing agents.

Frequently asked questions

Why is a single benchmark insufficient for healthcare AI evaluation?
A benchmark tests only one aspect of model performance, while real-world healthcare use is far more complex. A model can perform well on an isolated metric but fail unexpectedly in practice, requiring multiple testing methods that together provide a more complete picture.
What types of evaluation methods are recommended for high-stakes AI models?
The speaker emphasizes combining benchmarking, expert review, interpretability, and realistic task-based testing. This helps understand not only what a model can do but also where it breaks and how it will actually be used.
What are AI patients and how do they contribute to better evaluation?
AI patients are synthetic interactions designed to make benchmarking more realistic and scalable for patient-facing agents. They bring evaluation closer to real use cases and help understand model performance in more complex settings.
How do you determine whether an AI model is truly ready for healthcare use?
This goes beyond evaluation alone. You must understand what the model can do, where it breaks, what different kinds of evidence really tell us, and move from isolated results to a complete picture of readiness by applying multiple testing methods.

Topics