
0:00 / 0:00
ai
Benchmarks are measuring the wrong thing
Authority Hacker Podcast1 September 2026Watch on YouTube
Description
Most AI benchmarks test coding or creative writing. Businesses need support replies, presentations, articles, and everyday copy. So Gael built a blind head-to-head test around real business writing. The surprising part: the expensive closed models often did not win. Watch the full episode: https://podcasts.apple.com/ie/podcast/youre-using-the-wrong-ai-model-for-writing/id1073349789?i=1000786574042
What you'll learn
- That most AI benchmarks focus on coding and creative writing, while businesses mainly need supportive copy.
- That Gael built a blind head-to-head test around realistic business writing.
- That expensive closed AI models do not always win in business writing.
- How to evaluate models based on real business needs instead of generic benchmarks.
- That a blind test can deliver surprising results about which models perform best for business writing.
Frequently asked questions
What topics do most AI benchmarks test?
Most AI benchmarks test coding and creative writing. These topics do not align well with businesses' everyday need for supportive copy.
What did Gael build to compare AI models?
Gael built a blind head-to-head test around realistic business writing, such as support replies, presentations, articles, and everyday copy.
What was the surprising outcome of the blind test?
The expensive closed models often did not win. This shows that a high price tag does not always mean a better result for business writing.