All videos
0:00 / 0:00
research

Cancel your subscriptions, Ox-Alpha is here! (GLM 5.3 Flash)

Matthew Berman29 August 2026Watch on YouTube

What you'll learn

  • GLM 5.3 Flash is a mixture of experts model with 320 billion parameters, 18 billion active, at a tenth of its predecessor's price.
  • It scores 63.4 on DeepSWE, passing Claude Opus 4.8, which sits at 58.0.
  • Cost per completed task is 9 cents, against 95 cents for GPT 5.6 Soul and 3.14 dollar for Claude Fable 5.
  • During the preview Z.ai served 100 billion tokens per day on Chinese AI chips alone, without Nvidia hardware.
  • GLM 5.3 Flash averages 47,000 tokens per task, more than double the 20,000 tokens of GPT 5.6 Luna Max.

Topics

Read next

Sources

What is known about this topic outside the broadcast, and where it says so.

Description from the channel

Check out Higgsfield API: https://higgsfield.ai/s/TvYKkQ Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://z.ai/blog/glm-5.3-flash https://x.com/OpenRouter/status/2090544970923184269 https://artificialanalysis.ai/ https://x.com/SemiAnalysis_/status/2092623833630998556