All videos
0:00 / 0:00
research

Two AI Models Set to “stir government urgency”, But Will This Challenge Undo Them?

AI Explained15 June 2026Watch on YouTube

Part of series

Ep. 1 · Claude Fable Here

View the series

Description

First look at exclusive reports about OpenAI's new Spud model, and the model Anthropic think will stir governments to urgency, all in the context of the newly-launched ARC-AGI-3. What does the extreme difficulty of that benchmarks, and its quirky scoring metrics, mean for AI in 2026? https://assemblyai.com/aiexplained Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcouncil.ai AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:55 - OpenAI Side Quests 01:58 - Claude New Model Coming + Universal Equity? 03:13 - ARC-AGI 3 05:00 - Intentional or Unintentional Gaming? 07:11 - But is it AGI Harbinger? No Harness 09:41 - Not the First 12:32 - Automated Researcher 15:00 - Claw Caveat Spud: https://www.theinformation.com/articles/openai-ceo-shifts-responsibilities-preps-spud-ai-model?utm_campaign=Editorial&utm_content=Article&utm_medium=organic_social&utm_source=bluesky%2Cfacebook%2Clinkedin%2Cthreads%2Ctwitter&rc=sy0ihq FT: OpenAI Special Model: https://www.ft.com/content/de9bf0af-b241-424f-8229-5870b1c0d93d?syn-25a6b1a6=1 Jensen Huang: https://www.forbes.com/sites/antoniopequenoiv/2026/03/23/nvidias-jensen-huang-says-he-thinks-weve-achieved-agi/ Axios Article: https://archive.fo/20260326100140/https://www.axios.com/2026/03/26/anthropic-pentagon-ai-deal#selection-827.0-829.257 https://arcprize.org/arc-agi/3 ARC AGI 3 Paper: https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf NetHack Leaderboard: https://balrogai.com/ Paper: https://ai.meta.com/research/publications/the-nethack-learning-environment/ https://x.com/_rockt/status/2036864121585438995 Claw Shells: https://x.com/DrJimFan/status/2036494601750716711 OpenAI Automated Researcher: https://www.technologyreview.com/2026/03/20/1134438/openai-is-throwing-everything-into-building-a-fully-automated-researcher/ Patreon Post: https://www.patreon.com/c/aiexplained/posts Eng Jobs: https://x.com/lennysan/status/2036535460726767793 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/

What you'll learn

  • OpenAI is developing a new model called Spud, while Anthropic is building a model expected to prompt government urgency and action
  • The ARC-AGI-3 benchmark is being used to test these new AI models and reveals how advanced AI systems have become
  • The extreme difficulty of ARC-AGI-3 raises questions about whether models truly achieve AGI or simply optimize for benchmarks
  • Despite claims that AGI has been reached, the analysis shows current models still have limitations in solving complex problems

Frequently asked questions

What is ARC-AGI-3 and how does it differ from previous benchmarks?
ARC-AGI-3 is a new benchmark with extreme difficulty and unconventional scoring metrics designed to better test AI models on true generalization and problem-solving abilities.
How are OpenAI and Anthropic responding to this new benchmark?
OpenAI is working on Spud, while Anthropic is developing a new model expected to prompt government urgency, both being tested against ARC-AGI-3.
Do current AI models truly achieve AGI according to this video analysis?
According to the analysis, current models have not yet achieved true AGI despite some claims, showing limitations in complex problem-solving.
Are models truly improving or are they just optimizing for benchmarks?
The video questions this, suggesting some models may have simply optimized for benchmarks rather than achieving genuine generalization and intelligence progress.

Topics

In this video

Related reads