arXiv · 2410.10810
Local and Global Decoding in Text Generation
Abstract
Text generation, a key component in applications such as dialogue systems, relies on decoding algorithms that sample strings from a language model distribution. Traditional methods, such as top-$k$ and top-$\pi$, apply local normalisation to the model's output distribution, which can distort it. In this paper, we investigate the effect of this distortion by introducing globally-normalised versions of these decoding methods. Additionally, we propose an independent Metropolis-Hastings algorithm to approximate sampling from globally-normalised distributions without explicitly computing them. Our empirical analysis compares the performance of local and global normalisation across two decoding algorithms (top-$k$ and top-$\pi$) with various hyperparameters, using Pythia language models. Results show that, in most configurations, global decoding performs worse than the local decoding version of the same algorithms -- despite preserving the distribution's integrity. Our results suggest that distortion is an important feature of local decoding algorithms.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daniel Gareev, Thomas Hofmann, Ezhilmathi Krishnasamy, Tiago Pimentel. 2024-10-14. Local and Global Decoding in Text Generation. https://arxiv.org/abs/2410.10810
Cite the original work for its findings. Save a collection to share your selection of sources.