arXiv · 2610.08794
Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages
Abstract
Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ben Gubler. 2026-03-25. Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages. https://arxiv.org/abs/2610.08794
Cite the original work for its findings. Save a collection to share your selection of sources.