arXiv · 2101.11974
Disembodied Machine Learning: On the Illusion of Objectivity in NLP
Abstract
Machine Learning seeks to identify and encode bodies of knowledge within provided datasets. However, data encodes subjective content, which determines the possible outcomes of the models trained on it. Because such subjectivity enables marginalisation of parts of society, it is termed (social) `bias' and sought to be removed. In this paper, we contextualise this discourse of bias in the ML community against the subjective choices in the development process. Through a consideration of how choices in data and model development construct subjectivity, or biases that are represented in a model, we argue that addressing and mitigating biases is near-impossible. This is because both data and ML models are objects for which meaning is made in each step of the development pipeline, from data selection over annotation to model training and analysis. Accordingly, we find the prevalent discourse of bias limiting in its ability to address social marginalisation. We recommend to be conscientious of this, and to accept that de-biasing methods only correct for a fraction of biases.
Explore related subjects
Keep this discovery
Zeerak Waseem, Smarika Lulz, Joachim Bingel, Isabelle Augenstein. 2021-01-28. Disembodied Machine Learning: On the Illusion of Objectivity in NLP. https://arxiv.org/abs/2101.11974
Cite the original work for its findings. Save a collection to share your selection of sources.