arXiv ScienceSearch

arXiv subjects

Maria Milkova

Publications and source records attributed to Maria Milkova.

3 recordsLinked to original sources

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically validated, and transferable to an encoder model for scalable prediction. Using posts annotated according to Schwartz's theory of basic human values, we investigate how different LLMs and prompting strategies, which we call annotation regimes, operationalize the expression of values in text. Beyond standard classification metrics, we evaluate structural alignment, annotation ambiguity, error patterns, and stability across repeated runs. We find that different LLMs produce different value interpretations, and iterative prompt calibration through error analysis reduces misattributions and improves alignment with expert annotations. Error patterns are further used to derive targeted expert-verification rules for corpus annotation. We transfer soft LLM labels to an encoder model for prediction, retaining information about ambiguity in value expression. Finally, a sensitivity analysis on more than one million posts shows that regime-specific annotation differences propagate into predicted levels of value expression, whereas standardized temporal dynamics and the direction of major event responses are more robust.

cs.CL

Detecting Basic Values in A Noisy Russian Social Media Text Data: A Multi-Stage Classification Framework

This study presents a multi-stage classification framework for detecting human values in noisy Russian language social media, validated on a random sample of 7.5 million public text posts. Drawing on Schwartz's theory of basic human values, we design a multi-stage pipeline that includes spam and nonpersonal content filtering, targeted selection of value relevant and politically relevant posts, LLM based annotation, and multi-label classification. Particular attention is given to verifying the quality of LLM annotations and model predictions against human experts. We treat human expert annotations not as ground truth but as an interpretative benchmark with its own uncertainty. To account for annotation subjectivity, we aggregate multiple LLM generated judgments into soft labels that reflect varying levels of agreement. These labels are then used to train transformer based models capable of predicting the probability of each of the ten basic values. The best performing model, XLM RoBERTa large, achieves an F1 macro of 0.83 and an F1 of 0.71 on held out test data. By treating value detection as a multi perspective interpretive task, where expert labels, GPT annotations, and model predictions represent coherent but not identical readings of the same texts, we show that the model generally aligns with human judgments but systematically overestimates the Openness to Change value domain. Empirically, the study reveals distinct patterns of value expression and their co-occurrence in Russian social networks, contributing to a broader research agenda on cultural variation, communicative framing, and value based interpretation in digital environments. All models are released publicly.

cs.CL

Detecting value-expressive text posts in Russian social media

Basic values are concepts or beliefs which pertain to desirable end-states and transcend specific situations. Studying personal values in social media can illuminate how and why societal values evolve especially when the stimuli-based methods, such as surveys, are inefficient, for instance, in hard-to-reach populations. On the other hand, user-generated content is driven by the massive use of stereotyped, culturally defined speech constructions rather than authentic expressions of personal values. We aimed to find a model that can accurately detect value-expressive posts in Russian social media VKontakte. A training dataset of 5,035 posts was annotated by three experts, 304 crowd-workers and ChatGPT. Crowd-workers and experts showed only moderate agreement in categorizing posts. ChatGPT was more consistent but struggled with spam detection. We applied an ensemble of human- and AI-assisted annotation involving active learning approach, subsequently trained several classification models using embeddings from various pre-trained transformer-based language models. The best performance was achieved with embeddings from a fine-tuned rubert-tiny2 model, yielding high value detection quality (F1 = 0.77, F1-macro = 0.83). This model provides a crucial step to a study of values within and between Russian social media users.

cs.CL