All videos
0:00 / 0:00
research

On the Biology of a Large Language Model (Part 2)

Yannic Kilcher15 June 2026Watch on YouTube

Description

An in-depth look at Anthropic's Transformer Circuit Blog Post Part 1 here: https://youtu.be/mU3g2YPKlsA Discord here: https;//ykilcher.com/discord https://transformer-circuits.pub/2025/attribution-graphs/biology.html Abstract: We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology. Authors: Jack Lindsey†, Wes Gurnee*, Emmanuel Ameisen*, Brian Chen*, Adam Pearce*, Nicholas L. Turner*, Craig Citro*, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall◊, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, Joshua Batson*‡ Links: Homepage: https://ykilcher.com Merch: https://ykilcher.com/merch YouTube: https://www.youtube.com/c/yannickilcher Twitter: https://twitter.com/ykilcher Discord: https://ykilcher.com/discord LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this): SubscribeStar: https://www.subscribestar.com/yannickilcher Patreon: https://www.patreon.com/yannickilcher Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2 Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

What you'll learn

  • Circuit tracing reveals how Claude 3.5 Haiku processes and manipulates information internally across various contexts
  • The model's internal mechanisms can be characterized by tracing the flow of signals through neuron connections within the network
  • Language model interpretability is achievable by examining the computational structures and behavioral patterns inside the model

Frequently asked questions

What is circuit tracing and how is it used in this research?
Circuit tracing is a methodology used to map the internal workings of AI models. In this research, Anthropic scientists apply this technique to Claude 3.5 Haiku to understand how the model processes information and which internal mechanisms are responsible for specific behaviors.
Which model is the focus of this analysis?
Claude 3.5 Haiku, Anthropic's lightweight production model, is the focus of this analysis. The model has been examined using circuit tracing to gain insight into its internal processes.
Why is interpretability of language models important?
Interpretability allows researchers to understand how language models make decisions and process information. This is essential for building safer and more reliable AI systems.

Topics

In this video

Related reads