
AI is getting smarter—but it's still not aligned
Dr Waku16 June 2026Watch on YouTube
Part of series
Ep. 5 · Waarom AI liegt
Onderzoek naar de oorzaken van AI-hallucinaties en waarom gebruikers niet kunnen vertrouwen op de betrouwbaarheid van taalmodellen.
View the seriesDescription
There is a lot of online discussion about AI safety, but not much clarity on why exactly AI alignment is hard. We have to give AI systems goals and incentives for them to operate in the world, especially if they are agents. Since humans don't all agree on the right thing to do in all cases, how do we make sure AIs have the right goals? Alignment can fail in many ways including outer and inner misalignment. We examine an example of each, including reward hacking and deceptive alignment. We also describe the notion of mesa objectives. Once a system is misaligned, it enables misuse by humans and even loss of control to the AI itself. When this happens at scale, with reward hacking or deceptive misalignment inside the system, the AI could dramatically change its behavior to pursue a different goal. The so-called treacherous turn could happen after an arbitrary amount of time and could be very dangerous. Hence, more work on AI safety is needed for society to thrive. #ai #alignment #safety What is AI alignment? https://bluedot.org/blog/what-is-ai-alignment?from_site=aisf What is deceptive alignment? https://aisafety.info/questions/8EL6/What-is-deceptive-alignment Alignment faking in large language models [paper] https://arxiv.org/abs/2412.14093 Many-Shot Jailbreaking [paper] https://www.anthropic.com/research/many-shot-jailbreaking 0:00 Intro 0:28 Contents 0:36 Part 1: Alignment intuition 1:04 Coordination in AI agents 1:19 Example: Group projects in school 1:46 Example: Team at company building a product 2:14 Aspect 1: Putting goals into the machine 2:53 Human goals and values aren't consistent 3:23 Can't the legal system help? 3:41 Aspect 2: Setting out incentives 4:01 Unethical behavior from companies 4:25 Incentives create bad actors 4:52 Part 2: How alignment can fail 5:07 Outer and inner alignment 5:14 Example: outer misalignment in robots 5:43 Example: inner misalignment in mazes 6:03 Issue 1: Reward hacking 6:50 Example: Melting down silver coins 7:21 Isn't AI smart enough to know what you really mean? 7:58 Example: o3 engages in reward hacking 8:25 Issue 2: Deceptive alignment 8:49 Results of training will modify AI's goal 9:05 Model adjusts its behavior to look compliant 9:42 Range of mesa objectives 10:08 Paper: Alignment faking in large language models 10:19 Part 3: Why misalignment is dangerous 10:36 Category 1: Misuse by humans 11:08 Jailbreaks are inevitable 11:44 Paper: Many-shot Jailbreaking 11:56 Category 2: Loss of control 12:07 Agents are a powerful idea 12:23 Keeping yourself in the loop at first 12:40 Giving the system autonomy 13:26 Synthesis: What happens if these agents are misaligned 13:58 The treacherous turn 14:24 In the limit, we could be destroyed: they boil the seas 14:52 But can't human society tell us how to handle AI? 15:29 We may not survive a loss of control scenario 16:00 Conclusion 16:29 Consequences of alignment failure 17:00 Outro