
AI Safety as a Reinforcement Learning Problem
Amii21 July 2026Watch on YouTube
Description
This talk will focus on the current research directions of the AI Trust and Safety team at the Alberta Machine Intelligence Institute (Amii). While the broader AI safety community heavily relies on static data and supervised learning paradigms, real-world risk is inherently dynamic. Despite this reality, Reinforcement Learning (RL) is almost entirely missing from the current AI safety conversation. Addressing this gap is critical to developing a rigorous, scientific understanding of deployed AI. We will present Amii's vision for safety research, leveraging the University of Alberta's world-class strengths in reinforcement learning (RL) with the Trust & Safety team's focus on real-world deployments and failures. To illustrate this research vision, we introduce two early-stage initiatives that explore the different ways RL applies to safety: first, as a theoretical lens through which we diagnose structural failure modes in agentic AI, and second, as a multi-step sequential optimization tool to automate adversarial safety testing (such as LLM jailbreaking). We conclude with a call to action, bringing together the people building agents with those studying them at a granular, mathematical level to ensure that safety is built in from the start, rather than added as an afterthought. BIO: Montaser is an Applied Research Scientist at the Alberta Machine Intelligence Institute (Amii), where he works on research bridging RL and safety. With over five years of experience spanning industry and academia, his work centers on developing safe, reliable AI systems that learn from dynamic experience rather than static datasets. Montaser completed his PhD in Computing Science at the University of Alberta under the supervision of Prof. Michael Bowling, where his dissertation focused on a principled treatment of RL under partially observable reward signals and how agents can learn to act cautiously. Prior to his current role, he spent three years as an AI Engineer at Sony AI in Tokyo, contributing to the ACE table-tennis robot project, a system achieving professional-level play that resulted in a Nature publication. His research on safe, cautious autonomous agents has been published in top-tier AI venues.
What you'll learn
- Reinforcement Learning is a critical yet largely missing perspective in current AI safety, while real-world risks in deployed systems are inherently dynamic
- Amii diagnoses structural failure modes in agentic AI by using RL as a theoretical lens
- Adversarial testing can be automated by leveraging RL as a multi-step sequential optimization tool, including detection of LLM jailbreaking
- AI safety must be built in from the start by bringing together agent builders with safety researchers who work at a granular, mathematical level