arXiv · 2407.12838
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
Abstract
This paper presents two significant contributions: First, it introduces a novel dataset of 19th-century Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analysis in this region. Second, it develops a flexible framework that utilizes a Large Language Model for OCR error correction and linguistic surface form detection in digitized corpora. This semi-automated framework is adaptable to various contexts and datasets and is applied to the newly created dataset.
Explore related subjects
Keep this discovery
Laura Manrique-Gómez, Tony Montes, Arturo Rodríguez-Herrera, Rubén Manrique. 2024-07-04. Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction. https://doi.org/10.18653/v1/2024.nlp4dh-1.13
Cite the original work for its findings. Save a collection to share your selection of sources.