
Let's reproduce GPT-2 (124M)
Andrej Karpathy16 June 2026Watch on YouTube
Description
We reproduce the GPT-2 (124M) from scratch. This video covers the whole process: First we build the GPT-2 network, then we optimize its training to be really fast, then we set up the training run following the GPT-2 and GPT-3 paper and their hyperparameters, then we hit run, and come back the next morning to see our results, and enjoy some amusing model generations. Keep in mind that in some places this video builds on the knowledge from earlier videos in the Zero to Hero Playlist (see my channel). You could also see this video as building my nanoGPT repo, which by the end is about 90% similar. Links: - build-nanogpt GitHub repo, with all the changes in this video as individual commits: https://github.com/karpathy/build-nanogpt - nanoGPT repo: https://github.com/karpathy/nanoGPT - llm.c repo: https://github.com/karpathy/llm.c - my website: https://karpathy.ai - my twitter: https://twitter.com/karpathy - our Discord channel: https://discord.gg/3zy8kqD9Cp Supplementary links: - Attention is All You Need paper: https://arxiv.org/abs/1706.03762 - OpenAI GPT-3 paper: https://arxiv.org/abs/2005.14165 - OpenAI GPT-2 paper: https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf- The GPU I'm training the model on is from Lambda GPU Cloud, I think the best and easiest way to spin up an on-demand GPU instance in the cloud that you can ssh to: https://lambdalabs.com Chapters: 00:00:00 intro: Let’s reproduce GPT-2 (124M) 00:03:39 exploring the GPT-2 (124M) OpenAI checkpoint 00:13:47 SECTION 1: implementing the GPT-2 nn.Module 00:28:08 loading the huggingface/GPT-2 parameters 00:31:00 implementing the forward pass to get logits 00:33:31 sampling init, prefix tokens, tokenization 00:37:02 sampling loop 00:41:47 sample, auto-detect the device 00:45:50 let’s train: data batches (B,T) → logits (B,T,C) 00:52:53 cross entropy loss 00:56:42 optimization loop: overfit a single batch 01:02:00 data loader lite 01:06:14 parameter sharing wte and lm_head 01:13:47 model initialization: std 0.02, residual init 01:22:18 SECTION 2: Let’s make it fast. GPUs, mixed precision, 1000ms 01:28:14 Tensor Cores, timing the code, TF32 precision, 333ms 01:39:38 float16, gradient scalers, bfloat16, 300ms 01:48:15 torch.compile, Python overhead, kernel fusion, 130ms 02:00:18 flash attention, 96ms 02:06:54 nice/ugly numbers. vocab size 50257 → 50304, 93ms 02:14:55 SECTION 3: hyperpamaters, AdamW, gradient clipping 02:21:06 learning rate scheduler: warmup + cosine decay 02:26:21 batch size schedule, weight decay, FusedAdamW, 90ms 02:34:09 gradient accumulation 02:46:52 distributed data parallel (DDP) 03:10:21 datasets used in GPT-2, GPT-3, FineWeb (EDU) 03:23:10 validation data split, validation loss, sampling revive 03:28:23 evaluation: HellaSwag, starting the run 03:43:05 SECTION 4: results in the morning! GPT-2, GPT-3 repro 03:56:21 shoutout to llm.c, equivalent but faster code in raw C/CUDA 03:59:39 summary, phew, build-nanogpt github repo Corrections: I will post all errata and followups to the build-nanogpt GitHub repo (link above) SuperThanks: I experimentally enabled them on my channel yesterday. Totally optional and only use if rich. All revenue goes to to supporting my work in AI + Education.
What you'll learn
- You learn how to implement GPT-2 (124M) from scratch, covering neural network architecture through a complete training pipeline.
- GPU optimization techniques like mixed precision, torch.compile, and flash attention can dramatically speed up training from 1000ms to 93ms.
- The training process follows OpenAI standards for hyperparameters, learning rate scheduling, gradient accumulation, and distributed training.
- You see how to evaluate the trained model on benchmarks like HellaSwag and inspect generative model outputs.
- Practical implementation requires attention to details like parameter initialization, vocabulary alignment, and validation data splitting.
Frequently asked questions
Which GPU techniques are used to speed up the training process?
How does the implementation follow the guidelines from the original OpenAI GPT-2 and GPT-3 papers?
Which datasets are used for training and validation?
How are the training results evaluated?
Topics
In this video
Related reads
OpenAI aangeklaagd in zelfdodingszaak ChatGPT-gebruiker
OpenAI en ceo Sam Altman zijn in de VS aangeklaagd door een Canadese moeder die stelt dat haar dochter door ChatGPT is aangezet tot zelfdoding.
OpenAI mikt op lang draaiende agents met overname van Ona
OpenAI neemt Ona over, voorheen bekend als Gitpod, een startup die AI-agents laat draaien in cloud sandboxes. De overname versterkt Codex en stelt OpenAI in staat agents taken te laten uitvoeren die uren of dagen in beslag nemen.
OpenAI boekte bijna 39 miljard dollar verlies in 2025
Volgens uitgelekte documenten liep OpenAI's verlies vorig jaar fors op tot 38,53 miljard dollar bij een omzet van 13,07 miljard dollar.
OpenAI patcht beveiligingslekken in opensourcesoftware
OpenAI en Trail of Bits starten Patch the Planet om beveiligingsproblemen in opensourceprojecten op te sporen en te repareren.
OpenAI onthult eigen AI-chip Jalapeño
OpenAI kondigt Jalapeño aan, zijn eerste eigen AI-chip ontwikkeld met Broadcom, bedoeld voor LLM-inferentie met betere prestaties per watt dan bestaande chips.
OpenAI en Broadcom presenteren Jalapeño-chip
OpenAI en Broadcom onthullen Jalapeño, een AI-inferentiechip ontworpen voor LLM's, als eerste stap in een meerjarig hardwareplatform.
