
Oliver Sieberling - Dynamic Short Convolutions Improve Transformers
Cohere14 August 2026Watch on YouTube
What you'll learn
- Dynamic short convolutions use input-dependent filters, unlike static convolutions, enabling adaptive local context aggregation
- This improved primitive increases expressivity of Transformer-based language models and delivers stronger performance
- Dynamic short convolutions can be efficiently implemented using custom Triton kernels for optimal hardware utilization
- The technique combines efficient sequence modeling with scalable architecture design for language models
Frequently asked questions
What is the difference between dynamic and static short convolutions?
How are dynamic short convolutions implemented?
What benefits does this technique offer for language models?
Who researches this approach and with what focus?
Topics
Read next
Google Research publiceert TurboQuant, een compressiealgoritme voor grote taalmodellen
Google Research heeft TurboQuant gepubliceerd, een algoritme dat de KV-cache van grote taalmodellen comprimeert tot ongeveer 3 bits. Volgens de onderzoekers leidt dit tot zes keer minder geheugengebruik en acht keer hogere doorvoer op dezelfde GPU-hardware, zonder nauwkeurigheidsverlies.
Vier AI-onderzoeksresultaten die modellen sneller, slimmer en zuiniger maken
Vier recente onderzoeksontwikkelingen laten zien hoe AI-modellen kleiner, sneller en beter in redeneren worden: van spaarzame activering en zelfcorrectie tijdens training tot energiezuinige inferentie en een geheugenframework voor AI-agenten.
Description from the channel
This talk introduces dynamic short convolutions as a scalable neural network primitive for improving Transformer-based language models. In contrast to static short convolutions, dynamic short convolutions use input-dependent filters, which allows them to adaptively aggregate local context. I will discuss how this simple modification improves expressivity and leads to stronger performance in Transformer-based language models, as well as how we implemented dynamic short convolutions efficiently using custom Triton kernels. Oliver Sieberling is a PhD student at MIT, advised by Yoon Kim. His research focuses on the pretraining of large neural networks, with particular interests in efficient sequence modeling, scalable architectures, and hardware-algorithm co-design. He received his BSc in Computer Science from ETH Zurich in 2025 and has worked on LLM efficiency and evolutionary algorithms. This session is brought to you by the Cohere Labs Open Science Community - a space where ML researchers, engineers, linguists, social scientists, and lifelong learners connect and collaborate with each other. We'd like to extend a special thank you to Harsha Nelaturu and Andrej Jovanović, Leads of our ML Systems and Theory group for their dedication in organizing this event. If you’re interested in sharing your work, we welcome you to join us! Simply fill out the form at https://forms.gle/ALND9i6KouEEpCnz6 to express your interest in becoming a speaker. Join the Cohere Labs Open Science Community to see a full list of upcoming events (https://tinyurl.com/CohereLabsCommunityApp).