All videos
0:00 / 0:00
research

ML Summer School 2026 - Evaluations with Laurie Voss

Cohere3 August 2026Watch on YouTube

Description

​How do you know if your LLM application is actually working, and how do you catch it when it isn't? This hands-on session walks through the practical mechanics of evaluation and observability for production AI systems. Using Arize, we'll cover how to instrument your application, trace what's happening across multi-step and agentic workflows, set up evals that catch failures, and use feedback from evals to automatically improve your app. You'll leave with a concrete picture of what production-grade evaluation looks like in practice. This session is hosted as part of Cohere Labs Open Science Community's ML Summer School. At Cohere Labs Open Science Community, we believe the future of machine learning isn’t gated by credentials, borders, or budgets, it’s built in open collaboration, powered by curiosity, and made stronger through community. This summer, we’re excited for the return of Cohere Labs Open Science Community Summer School, a learning initiative featuring some of the leading minds in machine learning from Meta, Google DeepMind, Cohere Labs and more. This initiative reflects the core mission of Cohere Labs: supporting fundamental research, expanding access, and enabling the next generation of ML thinkers — regardless of where they start. We are very grateful to our community leads for organizing this series of events!

What you'll learn

  • You learn how to instrument an LLM application in production using tools like Arize for complete visibility
  • You discover how to trace multi-step and agentic workflows to see exactly what's happening in your system
  • You understand how to set up evals that detect failures and automatically provide feedback for improvement

Frequently asked questions

What is the purpose of instrumentation and tracing in LLM applications?
Instrumentation and tracing give you complete visibility into what's happening in your production environment, so you can understand how your LLM application actually works and where it fails.
How do evals help you automatically improve your app?
You set up evals that catch failures, then use the feedback from these evaluations to automatically improve your application and learn from mistakes.
Why is observability important for agentic workflows?
Agentic workflows have multiple steps and decisions, so you need observability to see what happens at each stage and understand why the system takes specific actions.
What tools and frameworks are covered in this workshop?
The workshop uses Arize, an observability platform, to demonstrate practically how to implement evaluation and monitoring in production.

Topics