All videos
0:00 / 0:00
research

Qwen3.8 27B Local Test with llama.cpp | Best Small Local Model? | Coding & Agentic Work | 🔴 Live

Venelin Valkov14 August 2026Watch on YouTube

Part of series

Ep. 9 · Cpp Llama Local

View the series

What you'll learn

  • Qwen3.8 27B runs in 4-bit quantization on a MacBook M5 Pro with 48 GB unified memory, using over 17 GB of memory.
  • With the current llama.cpp build it reaches roughly 13 to 14 tokens per second on this machine, dropping below 9 while streaming.
  • The multi-token prediction (MTP) feature gives no benefit so far, with the drafter enabled the speed drops two to three tokens per second.
  • Complex visual coding tasks such as an SVG diagram take tens of minutes because of the low inference speed.
  • Since only the text-to-text weights of Qwen3.8-Max are open, the 27B is the only open-weight Qwen3.8 variant that takes both images and text as input.

Frequently asked questions

On which hardware is Qwen3.8 27B tested in this broadcast?
Venelin Valkov runs the model locally on a MacBook M5 Pro with 48 GB of unified memory, in 4-bit dynamic quantization (Q4_K_XL) using the GGUF quants from Unsloth and the latest llama.cpp server.
What inference speed does the model reach on this setup?
The model reaches roughly 13 to 14 tokens per second when not streaming, dropping below 9 tokens per second while streaming. It uses around 17.4 GB of memory.
Does multi-token prediction (MTP) work on this model?
Not beneficially yet. With the MTP drafter enabled, Venelin gets two to three tokens per second fewer than without, and he expects improvement only with further versions of GGUF and the llama.cpp server.
How does this 27B version differ from the larger Qwen3.8-Max?
Qwen3.8-Max weights are text-to-text only and not open, which makes the 27B the only open-weight variant in the family that handles both images and text as input.

Topics

Sources

What is known about this topic outside the broadcast, and where it says so.

Description from the channel

Qwen3.8 27B is live and the weights are here! How much better is it than Qwen3.6 27B? Weights: https://huggingface.co/Qwen/Qwen3.8-27B AI Academy: https://MLExpert.io Work with me: https://mlexpert.io/consulting DaBench (open weight LLM benchmarks): https://dabench.ai LinkedIn: https://www.linkedin.com/in/venelin-valkov Follow me on X: https://twitter.com/venelin_valkov Discord: https://discord.gg/UaNPxVD6tv Subscribe: http://bit.ly/venelin-subscribe GitHub repository: https://github.com/curiousily/AI-Bootcamp 👍 Don't Forget to Like, Comment, and Subscribe for More Tutorials! Join this channel to get access to the perks and support my work: https://www.youtube.com/channel/UCoW_WzQNJVAjxo4osNAxd_g/join