
On the Biology of a Large Language Model (Part 2)
Yannic Kilcher15 June 2026Watch on YouTube
Description
An in-depth look at Anthropic's Transformer Circuit Blog Post Part 1 here: https://youtu.be/mU3g2YPKlsA Discord here: https;//ykilcher.com/discord https://transformer-circuits.pub/2025/attribution-graphs/biology.html Abstract: We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology. Authors: Jack Lindsey†, Wes Gurnee*, Emmanuel Ameisen*, Brian Chen*, Adam Pearce*, Nicholas L. Turner*, Craig Citro*, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall◊, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, Joshua Batson*‡ Links: Homepage: https://ykilcher.com Merch: https://ykilcher.com/merch YouTube: https://www.youtube.com/c/yannickilcher Twitter: https://twitter.com/ykilcher Discord: https://ykilcher.com/discord LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this): SubscribeStar: https://www.subscribestar.com/yannickilcher Patreon: https://www.patreon.com/yannickilcher Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2 Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
What you'll learn
- Circuit tracing reveals how Claude 3.5 Haiku processes and manipulates information internally across various contexts
- The model's internal mechanisms can be characterized by tracing the flow of signals through neuron connections within the network
- Language model interpretability is achievable by examining the computational structures and behavioral patterns inside the model
Frequently asked questions
What is circuit tracing and how is it used in this research?
Which model is the focus of this analysis?
Why is interpretability of language models important?
Topics
In this video
Related reads
Trump wijzigt standpunt over Anthropic als veiligheidsrisico
President Trump stelt in een interview met Axios dat hij Anthropic niet langer als nationaal veiligheidsrisico beschouwt, nadat de VS het AI-bedrijf eerder dit jaar zo had bestempeld.
OpenAI mikt op lang draaiende agents met overname van Ona
OpenAI neemt Ona over, voorheen bekend als Gitpod, een startup die AI-agents laat draaien in cloud sandboxes. De overname versterkt Codex en stelt OpenAI in staat agents taken te laten uitvoeren die uren of dagen in beslag nemen.
Apple zou in overleg zijn met Anthropic en Google over Siri-alternatieven
Apple overlegt naar verluidt met Anthropic en Google over het beschikbaar stellen van hun AI-chatbots als alternatief voor Siri via een extensiesysteem.
Amazon zou overheid om blokkering Anthropic-model hebben gevraagd
Volgens berichtgeving zou Amazon de Amerikaanse overheid hebben benaderd om een Anthropic-model offline te halen, gevolgd door meldingen van kwetsbaarheden.
Claude Fable 5 en Mythos 5 geblokkeerd voor niet-Amerikanen na jailbreak-zorgen
Anthropic heeft op last van Washington de toegang tot Claude Fable 5 en Mythos 5 geblokkeerd voor gebruikers buiten de Verenigde Staten. De aanleiding is een jailbreak die volgens de Amerikaanse overheid de nationale veiligheid in gevaar brengt.
Anthropic beperkt toegang tot geavanceerde Mythos-modellen na overheidsingrijpen
Na inmenging van de Amerikaanse overheid stelt Anthropic de toegang tot zijn geavanceerde modellen grotendeels stil, hoewel een kleine groep organisaties toegang tot een experimentele versie behoudt.
