
Qwen3.8 27B Local Test with llama.cpp | Best Small Local Model? | Coding & Agentic Work | 🔴 Live
Venelin Valkov14 August 2026Watch on YouTube
Part of series
Ep. 9 · Cpp Llama Local
View the seriesWhat you'll learn
- Qwen3.8 27B runs in 4-bit quantization on a MacBook M5 Pro with 48 GB unified memory, using over 17 GB of memory.
- With the current llama.cpp build it reaches roughly 13 to 14 tokens per second on this machine, dropping below 9 while streaming.
- The multi-token prediction (MTP) feature gives no benefit so far, with the drafter enabled the speed drops two to three tokens per second.
- Complex visual coding tasks such as an SVG diagram take tens of minutes because of the low inference speed.
- Since only the text-to-text weights of Qwen3.8-Max are open, the 27B is the only open-weight Qwen3.8 variant that takes both images and text as input.
Frequently asked questions
On which hardware is Qwen3.8 27B tested in this broadcast?
What inference speed does the model reach on this setup?
Does multi-token prediction (MTP) work on this model?
How does this 27B version differ from the larger Qwen3.8-Max?
Topics
Sources
What is known about this topic outside the broadcast, and where it says so.
Description from the channel
Qwen3.8 27B is live and the weights are here! How much better is it than Qwen3.6 27B? Weights: https://huggingface.co/Qwen/Qwen3.8-27B AI Academy: https://MLExpert.io Work with me: https://mlexpert.io/consulting DaBench (open weight LLM benchmarks): https://dabench.ai LinkedIn: https://www.linkedin.com/in/venelin-valkov Follow me on X: https://twitter.com/venelin_valkov Discord: https://discord.gg/UaNPxVD6tv Subscribe: http://bit.ly/venelin-subscribe GitHub repository: https://github.com/curiousily/AI-Bootcamp 👍 Don't Forget to Like, Comment, and Subscribe for More Tutorials! Join this channel to get access to the perks and support my work: https://www.youtube.com/channel/UCoW_WzQNJVAjxo4osNAxd_g/join