arXiv · 2208.01561
Lost in Space Marking
Abstract
We look at a decision taken early in training a subword tokenizer, namely whether it should be the word-initial token that carries a special mark, or the word-final one. Based on surface-level considerations of efficiency and cohesion, as well as morphological coverage, we find that a Unigram LM tokenizer trained on pre-tokenized English text is better off marking the word-initial token, while one trained on raw text benefits from marking word ends. Our findings generalize across domains.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cassandra L. Jacobs, Yuval Pinter. 2022-08-02. Lost in Space Marking. https://arxiv.org/abs/2208.01561
Cite the original work for its findings. Save a collection to share your selection of sources.