All videos
0:00 / 0:00
applications

Claude's Most Expensive Model Is Doing Your Grunt Work

Authority Hacker Podcast21 July 2026Watch on YouTube

Part of series

Ep. 10 · Claude token-gebruik optimaliseren

Technische handleidingen voor het drastisch verminderen van tokenverbruik bij Claude via de RTK-plugin en slim contextbeheer.

View the series

Description

Gael can write a full blog post using about 7% of his weekly Fable limit. Most people use 50% doing the same job. The difference isn't a better prompt. He stopped letting the smartest model do the work at all. Fable briefs the task, Sonnet handles research, Opus makes the editorial calls, Codex writes the code. Fable just manages. With Anthropic cutting weekly limits by roughly a third, this stops being optional. In this episode we break down: → Why Fable is a project manager and not a worker → How to build a routing table so delegation happens automatically → Why spawning 100 subagents costs you more, not less → Running eight threads at once without losing track of them → OpenAI's new model names and what the ChatGPT and Codex merge means Token efficiency is about to be a job. Start learning it now. 🔗 AI Accelerator: https://www.authorityhacker.com/ai-accelerator 💻 Learn Claude Code: https://www.authorityhacker.com/resources/learn-claude-code/ ⏱️ Timestamps: 00:00 - Why efficiency is the new game 02:50 - What orchestration actually means 05:49 - Your $1.1M model doing grunt work 07:54 - Building the delegate skill 10:30 - Why 100 subagents backfires 16:40 - Running eight threads at once 23:21 - OpenAI's new models and the app merge 33:26 - Kimi K3 and the proposed US ban 39:23 - Software companies picking sides

Topics

In this video

Related reads