arXiv · 2602.16241
Are LLMs Ready to Replace Bangla Annotators?
Abstract
Large Language Models (LLMs) are increasingly used as automated annotators to scale dataset creation, yet their reliability as unbiased annotators--especially for low-resource and identity-sensitive settings--remains poorly understood. In this work, we study the behavior of LLMs as zero-shot annotators for Bangla hate speech, a task where even human agreement is challenging, and annotator bias can have serious downstream consequences. We conduct a systematic benchmark of 17 LLMs using a unified evaluation framework. Our analysis uncovers annotator bias and substantial instability in model judgments. Surprisingly, increased model scale does not guarantee improved annotation quality--smaller, more task-aligned models frequently exhibit more consistent behavior than their larger counterparts. These results highlight important limitations of current LLMs for sensitive annotation tasks in low-resource languages and underscore the need for careful evaluation before deployment.
Explore related subjects
Keep this discovery
Md. Najib Hasan, Touseef Hasan, Souvika Sarkar. 2026-02-18. Are LLMs Ready to Replace Bangla Annotators?. https://arxiv.org/abs/2602.16241
Cite the original work for its findings. Save a collection to share your selection of sources.