All videos
0:00 / 0:00
research

Understanding the Staged Dynamics of Transformers in Learning Latent Structure, Rohan Saha

Amii Intelligence16 June 2026Watch on YouTube

Description

The AI Seminar is a weekly meeting at the University of Alberta where researchers interested in artificial intelligence (AI) can share their research. Presenters include both local speakers from the University of Alberta and visitors from other institutions. Topics can be related in any way to artificial intelligence, from foundational theoretical work to innovative applications of AI techniques to new fields and problems. In this seminar from the Alberta Machine Intelligence Institute and the Department of Computing Science, Rohan Saha, PhD student, explains that a transformer model learns latent structure in stages. Training on the Alchemy benchmark reveals that the stages correspond to different latent structure components. Increasing task complexity does not delay model convergence when composing fundamental rules, but delays convergence when decomposing complex examples. Freezing interventions demonstrate the presence of layer-wise plasticity windows for stage acquisition. Bio: Rohan Saha is a PhD student at the University of Alberta studying the learning dynamics of transformer models.

What you'll learn

  • Transformer models learn latent structure in stages, with each stage representing a different component of the underlying structure
  • Task complexity has different effects: it does not delay convergence when composing fundamental rules, but delays convergence when decomposing complex examples
  • Layer-wise plasticity windows determine when models can acquire different stages of structure learning, as demonstrated by freezing interventions

Frequently asked questions

In what stages do transformer models learn latent structure?
Transformer models learn latent structure in multiple sequential stages, with each stage focusing on a specific component of the underlying structure. This was demonstrated through experiments on the Alchemy benchmark.
How does task complexity affect the learning process of transformers?
Task complexity has different effects depending on the task type. When composing fundamental rules, complexity does not delay convergence, but when decomposing complex examples, it does.
What are plasticity windows in transformer models?
Plasticity windows are layer-specific periods in which transformers can acquire new structure components. Freezing interventions demonstrate that these windows are critical for learning across different stages.

Topics