arXiv · 2412.09925
Simulating Hard Attention Using Soft Attention
Abstract
We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Andy Yang, Lena Strobl, David Chiang, Dana Angluin. 2024-12-13. Simulating Hard Attention Using Soft Attention. https://arxiv.org/abs/2412.09925
Cite the original work for its findings. Save a collection to share your selection of sources.