
How DeepSeek built cutting-edge AI without top GPUs
Dr Waku16 June 2026Watch on YouTube
Description
The announcement of DeepSeek’s latest model R1 has taken the world by storm. There are two main reasons why: it was produced by a Chinese company, despite export controls on GPUs; and, it seems to have cost a lot less to train than other frontier reasoning models. The financial markets reacted strongly, with Nvidia stock dipping almost $600 billion in one day. We discuss where DeepSeek came from, which is in many ways a story of its founder Liang Wenfeng, who made a fortune in the finance industry and then invested it in GPUs. Despite export controls from the US, they obtained whichever GPUs they were able to, and worked with those. The model itself incorporates a number of impressive innovations. In particular, with engineers trained on quant trading from the financial firm High-Flyer, DeepSeek optimized the implementation of their training and inference. They used H800 gpus which were at one point available despite embargos, because they did not have powerful inter-GPU communication capabilities. DeepSeek reprogrammed the GPUs at the assembly level to support fast communication. A very interesting model and a very interesting company overall. #deepseek #r1 #ai DeepSeek’s first reasoning model R1-Lite-Preview turns heads, beating OpenAI o1 performance https://venturebeat.com/ai/deepseeks-first-reasoning-model-r1-lite-preview-turns-heads-beating-openai-o1-performance/ DeepSeek FAQ https://stratechery.com/2025/deepseek-faq/ DeepSeek Debates: Chinese Leadership On Cost, True Training Cost, Closed Model Margin Impacts https://semianalysis.com/2025/01/31/deepseek-debates/ Exploring DeepSeek’s R1 Training Process https://towardsdatascience.com/exploring-deepseeks-r1-training-process-5036c42deeb1/ DeepSeek's AI breakthrough bypasses industry-standard CUDA, uses Nvidia's assembly-like PTX programming instead https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseeks-ai-breakthrough-bypasses-industry-standard-cuda-uses-assembly-like-ptx-programming-instead “This level of optimization is nuts but would definitely allow them...” [reddit] https://www.reddit.com/r/LocalLLaMA/comments/1icaq2z/deepseeks_ai_breakthrough_bypasses_nvidias/ Why everyone is freaking out about DeepSeek https://www.theverge.com/ai-artificial-intelligence/598846/deepseek-big-tech-ai-industry-nvidia-impac DeepSeek-V3 Technical Report [paper] https://arxiv.org/html/2412.19437v1#S2 Nvidia shares sink as Chinese AI app spooks markets https://www.bbc.com/news/articles/c0qw7z2v1pgo Here’s how much AI firm DeepSeek and its founder are worth https://www.forbes.com.au/news/billionaires/how-much-ai-firm-deepseek-and-its-founder-are-worth/ 0:00 Intro 0:21 Contents 0:29 Part 1: DeepSeek meets geopolitics 0:49 DeepSeek is a new player in AI lab space 1:17 Two reasons DeepSeek captured everyone's attention 1:44 Reason 1: A Chinese firm catches up out of nowhere 2:13 US export restrictions to China 2:58 Timeline of DeepSeek release 3:30 Stock market impact of DeepSeek 3:54 Reason 2: Are companies spending 50x more than they need to? 4:31 My main lessons from DeepSeek's release 4:59 There is no moat anymore 5:25 Nvidia dropped but then recovered 5:34 Jevons Paradox says total GPU consumption may increase 6:12 Part 2: Where DeepSeek came from 6:15 Founder of DeepSeek is Liang Wenfeng 6:26 Starting out in finance 6:46 Serial entrepreneur leading to High-Flyer 7:11 Spending $155 million on 10,000 A100 GPUs 7:41 Why Liang might have invested in GPUs 8:09 DJI connection 8:26 Thoughts on Liang's personality 9:02 DeepSeek was founded from High-Flyer 9:29 Taking attention away from High-Flyer 9:55 Obtaining 10,000 H800 GPUs due to sanctions 10:35 The rest of the firm's GPUs (H20, black market) 11:16 How expensive were DeepSeek V3 and R1 really? $10 million 12:08 But many other experiments needed to be performed 12:41 Total hardware cost for company: $500 million 13:01 DeepSeek was not expecting popularity 13:39 Part 3: Breaking down DeepSeek-R1 13:53 Algorithmic improvement per year: 4x-10x 14:58 Did DeepSeek use distillation from OpenAI? 15:44 Distillation is only to be expected 16:01 DeepSeek released their models open weight 16:34 Algorithmic improvements in DeepSeek R1 16:37 Innovation 1: Multi-head attention 16:54 Innovation 2: Multi token prediction 17:13 Innovation 3: Mixture of Experts x256 18:06 Innovation 4: Using H800 GPUs despite disadvantages 18:55 Repurposing processing for communication 19:35 Bypassing CUDA and going to assembly (PTX) 20:25 List of suggestions for Nvidia 20:49 Innovation 5: Forward and backward pass simultaneously 21:47 Innovation 6: Moving experts with redundant copies 22:56 Training overview: pre-training 23:19 Training overview: reinforcement learning (R1-Zero) 23:59 Making full R-1 in many stages 24:48 AI-driven reinforcement learning 25:12 Conclusion 25:25 Founded by reclusive founder geek 26:05 DeepSeek model is way cheaper 26:25 Engineers solved amazing challenges 27:08 Outro