arXiv · 2609.21676
The Spoken Wikipedia Presentation Corpus
Abstract
We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Thomas Ranzenberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer. 2026-09-18. The Spoken Wikipedia Presentation Corpus. https://arxiv.org/abs/2609.21676
Cite the original work for its findings. Save a collection to share your selection of sources.