The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks. Sun, S., Canziani, A., LeCun, Y., & Zhu, J. March, 2026. arXiv:2603.05498 [cs]
Paper doi abstract bibtex We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters of the model. Attention sinks operate locally: they modulate attention outputs across heads and bias individual heads toward short-range dependencies. We identify the pre-norm configuration as the key choice that enables the co-occurrence, and show that ablating it causes the two phenomena to decouple.
@misc{sun2026Spike,
title = {The {Spike}, the {Sparse} and the {Sink}: {Anatomy} of {Massive} {Activations} and {Attention} {Sinks}},
shorttitle = {The {Spike}, the {Sparse} and the {Sink}},
url = {http://arxiv.org/abs/2603.05498},
doi = {10.48550/arXiv.2603.05498},
abstract = {We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters of the model. Attention sinks operate locally: they modulate attention outputs across heads and bias individual heads toward short-range dependencies. We identify the pre-norm configuration as the key choice that enables the co-occurrence, and show that ablating it causes the two phenomena to decouple.},
urldate = {2026-03-15},
publisher = {arXiv},
author = {Sun, Shangwen and Canziani, Alfredo and LeCun, Yann and Zhu, Jiachen},
month = mar,
year = {2026},
note = {arXiv:2603.05498 [cs]},
keywords = {Computer Science - Artificial Intelligence, Computer Science - Computation and Language, Improvements, Theoretical},
}
Downloads: 0
{"_id":"ZqonpPHqTFvj2L3ue","bibbaseid":"sun-canziani-lecun-zhu-thespikethesparseandthesinkanatomyofmassiveactivationsandattentionsinks-2026","author_short":["Sun, S.","Canziani, A.","LeCun, Y.","Zhu, J."],"bibdata":{"bibtype":"misc","type":"misc","title":"The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks","shorttitle":"The Spike, the Sparse and the Sink","url":"http://arxiv.org/abs/2603.05498","doi":"10.48550/arXiv.2603.05498","abstract":"We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters of the model. Attention sinks operate locally: they modulate attention outputs across heads and bias individual heads toward short-range dependencies. We identify the pre-norm configuration as the key choice that enables the co-occurrence, and show that ablating it causes the two phenomena to decouple.","urldate":"2026-03-15","publisher":"arXiv","author":[{"propositions":[],"lastnames":["Sun"],"firstnames":["Shangwen"],"suffixes":[]},{"propositions":[],"lastnames":["Canziani"],"firstnames":["Alfredo"],"suffixes":[]},{"propositions":[],"lastnames":["LeCun"],"firstnames":["Yann"],"suffixes":[]},{"propositions":[],"lastnames":["Zhu"],"firstnames":["Jiachen"],"suffixes":[]}],"month":"March","year":"2026","note":"arXiv:2603.05498 [cs]","keywords":"Computer Science - Artificial Intelligence, Computer Science - Computation and Language, Improvements, Theoretical","bibtex":"@misc{sun2026Spike,\n\ttitle = {The {Spike}, the {Sparse} and the {Sink}: {Anatomy} of {Massive} {Activations} and {Attention} {Sinks}},\n\tshorttitle = {The {Spike}, the {Sparse} and the {Sink}},\n\turl = {http://arxiv.org/abs/2603.05498},\n\tdoi = {10.48550/arXiv.2603.05498},\n\tabstract = {We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters of the model. Attention sinks operate locally: they modulate attention outputs across heads and bias individual heads toward short-range dependencies. We identify the pre-norm configuration as the key choice that enables the co-occurrence, and show that ablating it causes the two phenomena to decouple.},\n\turldate = {2026-03-15},\n\tpublisher = {arXiv},\n\tauthor = {Sun, Shangwen and Canziani, Alfredo and LeCun, Yann and Zhu, Jiachen},\n\tmonth = mar,\n\tyear = {2026},\n\tnote = {arXiv:2603.05498 [cs]},\n\tkeywords = {Computer Science - Artificial Intelligence, Computer Science - Computation and Language, Improvements, Theoretical},\n}\n\n","author_short":["Sun, S.","Canziani, A.","LeCun, Y.","Zhu, J."],"key":"sun2026Spike","id":"sun2026Spike","bibbaseid":"sun-canziani-lecun-zhu-thespikethesparseandthesinkanatomyofmassiveactivationsandattentionsinks-2026","role":"author","urls":{"Paper":"http://arxiv.org/abs/2603.05498"},"keyword":["Computer Science - Artificial Intelligence","Computer Science - Computation and Language","Improvements","Theoretical"],"metadata":{"authorlinks":{}}},"bibtype":"misc","biburl":"https://api.zotero.org/users/4032374/collections/4UCJZAVL/items?key=5f7T4OfDqkAYW22yql2BxO5c&format=bibtex&limit=100","dataSources":["h7GBqZX35Zw97tcZk"],"keywords":["computer science - artificial intelligence","computer science - computation and language","improvements","theoretical"],"search_terms":["spike","sparse","sink","anatomy","massive","activations","attention","sinks","sun","canziani","lecun","zhu"],"title":"The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks","year":2026}