All videos
0:00 / 0:00
research

Oliver Sieberling - Dynamic Short Convolutions Improve Transformers

Cohere14 August 2026Watch on YouTube

What you'll learn

  • Dynamic short convolutions use input-dependent filters, unlike static convolutions, enabling adaptive local context aggregation
  • This improved primitive increases expressivity of Transformer-based language models and delivers stronger performance
  • Dynamic short convolutions can be efficiently implemented using custom Triton kernels for optimal hardware utilization
  • The technique combines efficient sequence modeling with scalable architecture design for language models

Frequently asked questions

What is the difference between dynamic and static short convolutions?
Dynamic short convolutions use input-dependent filters that adapt to the input, whereas static convolutions have fixed filters. This enables dynamic convolutions to perform context aggregation adaptively.
How are dynamic short convolutions implemented?
They are implemented efficiently using custom Triton kernels, which enable hardware-level optimization for improved performance.
What benefits does this technique offer for language models?
Dynamic short convolutions increase expressivity and deliver stronger performance in Transformer-based language models through adaptive local context aggregation.
Who researches this approach and with what focus?
Oliver Sieberling, a PhD student at MIT, investigates this under the guidance of Yoon Kim, focusing on efficient sequence modeling, scalable architectures, and hardware-algorithm co-design.

Topics

Read next

Description from the channel

​This talk introduces dynamic short convolutions as a scalable neural network primitive for improving Transformer-based language models. In contrast to static short convolutions, dynamic short convolutions use input-dependent filters, which allows them to adaptively aggregate local context. I will discuss how this simple modification improves expressivity and leads to stronger performance in Transformer-based language models, as well as how we implemented dynamic short convolutions efficiently using custom Triton kernels. ​Oliver Sieberling is a PhD student at MIT, advised by Yoon Kim. His research focuses on the pretraining of large neural networks, with particular interests in efficient sequence modeling, scalable architectures, and hardware-algorithm co-design. He received his BSc in Computer Science from ETH Zurich in 2025 and has worked on LLM efficiency and evolutionary algorithms. This session is brought to you by the Cohere Labs Open Science Community - a space where ML researchers, engineers, linguists, social scientists, and lifelong learners connect and collaborate with each other. We'd like to extend a special thank you to Harsha Nelaturu and Andrej Jovanović, Leads of our ML Systems and Theory group for their dedication in organizing this event. If you’re interested in sharing your work, we welcome you to join us! Simply fill out the form at https://forms.gle/ALND9i6KouEEpCnz6 to express your interest in becoming a speaker. Join the Cohere Labs Open Science Community to see a full list of upcoming events (https://tinyurl.com/CohereLabsCommunityApp).