
Two AI Models Set to “stir government urgency”, But Will This Challenge Undo Them?
AI Explained15 June 2026Watch on YouTube
Part of series
Ep. 1 · Claude Fable Here
View the seriesDescription
First look at exclusive reports about OpenAI's new Spud model, and the model Anthropic think will stir governments to urgency, all in the context of the newly-launched ARC-AGI-3. What does the extreme difficulty of that benchmarks, and its quirky scoring metrics, mean for AI in 2026? https://assemblyai.com/aiexplained Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcouncil.ai AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:55 - OpenAI Side Quests 01:58 - Claude New Model Coming + Universal Equity? 03:13 - ARC-AGI 3 05:00 - Intentional or Unintentional Gaming? 07:11 - But is it AGI Harbinger? No Harness 09:41 - Not the First 12:32 - Automated Researcher 15:00 - Claw Caveat Spud: https://www.theinformation.com/articles/openai-ceo-shifts-responsibilities-preps-spud-ai-model?utm_campaign=Editorial&utm_content=Article&utm_medium=organic_social&utm_source=bluesky%2Cfacebook%2Clinkedin%2Cthreads%2Ctwitter&rc=sy0ihq FT: OpenAI Special Model: https://www.ft.com/content/de9bf0af-b241-424f-8229-5870b1c0d93d?syn-25a6b1a6=1 Jensen Huang: https://www.forbes.com/sites/antoniopequenoiv/2026/03/23/nvidias-jensen-huang-says-he-thinks-weve-achieved-agi/ Axios Article: https://archive.fo/20260326100140/https://www.axios.com/2026/03/26/anthropic-pentagon-ai-deal#selection-827.0-829.257 https://arcprize.org/arc-agi/3 ARC AGI 3 Paper: https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf NetHack Leaderboard: https://balrogai.com/ Paper: https://ai.meta.com/research/publications/the-nethack-learning-environment/ https://x.com/_rockt/status/2036864121585438995 Claw Shells: https://x.com/DrJimFan/status/2036494601750716711 OpenAI Automated Researcher: https://www.technologyreview.com/2026/03/20/1134438/openai-is-throwing-everything-into-building-a-fully-automated-researcher/ Patreon Post: https://www.patreon.com/c/aiexplained/posts Eng Jobs: https://x.com/lennysan/status/2036535460726767793 Non-hype Newsletter: https://signaltonoise.beehiiv.com/ Podcast: https://aiexplainedopodcast.buzzsprout.com/
What you'll learn
- OpenAI is developing a new model called Spud, while Anthropic is building a model expected to prompt government urgency and action
- The ARC-AGI-3 benchmark is being used to test these new AI models and reveals how advanced AI systems have become
- The extreme difficulty of ARC-AGI-3 raises questions about whether models truly achieve AGI or simply optimize for benchmarks
- Despite claims that AGI has been reached, the analysis shows current models still have limitations in solving complex problems
Frequently asked questions
What is ARC-AGI-3 and how does it differ from previous benchmarks?
How are OpenAI and Anthropic responding to this new benchmark?
Do current AI models truly achieve AGI according to this video analysis?
Are models truly improving or are they just optimizing for benchmarks?
Topics
In this video
Related reads
OpenAI mikt op lang draaiende agents met overname van Ona
OpenAI neemt Ona over, voorheen bekend als Gitpod, een startup die AI-agents laat draaien in cloud sandboxes. De overname versterkt Codex en stelt OpenAI in staat agents taken te laten uitvoeren die uren of dagen in beslag nemen.
OpenAI introduceert ChatGPT Work, een AI-agent die taken zelfstandig afrondt
OpenAI heeft ChatGPT Work gelanceerd, een AI-agent binnen ChatGPT die complexe projecten zelfstandig uitvoert en eindproducten oplevert zoals spreadsheets, presentaties en webapplicaties. De lancering gaat gepaard met het samenvoegen van de Codex-app in de ChatGPT-desktopapp en de introductie van de GPT-5.6-modelreeks.
Trump schrapt beperkingen op Anthropic-modellen Mythos en Fable
De Trump-regering heeft beperkingen op Anthropic's AI-modellen Mythos en Fable laten vallen. Het wisselvallige AI-beleid van de administratie zorgt voor onduidelijkheid bij bedrijven over welke regels gelden voor toekomstige modelreleases.
Microsoft vervangt OpenAI en Anthropic-modellen in Copilot
Microsoft voert eigen MAI-modellen in voor producten als Excel en Outlook om kosten van externe modellen te verlagen.
Anthropic werkt aan eigen AI-chip met Samsung
Anthropic ontwikkelt eigen AI-chips en voert gesprekken met Samsung voor productie, in concurrentie met OpenAI's recent gepresenteerde chip.
Amazon onderzoekt alternatieve LLM's vanwege Claude-prijsstijging
Amazon zou OpenAI en zijn eigen Nova-modellen evalueren voor interne use cases na prijsverhogingen van Anthropic's Claude.

