arXiv Science⌕ Search

arXiv · 2610.06164

Predicting Engagement with Sponsored Content Across Account Types on Instagram

Abstract

In this work we explore how users interact with sponsored content on social media platforms by collecting and analyzing a large-scale dataset of sponsored Instagram posts. To maintain transparency, we favor robust statistical analysis and explainable models over deep learning techniques. Our pipeline categorizes Instagram accounts along multiple dimensions, including audience size and entity type. We complement this with semantic features extracted from post captions and hashtags, and use these elements to train regression models that forecast engagement. We validate our approach on a dataset comprising over 15M Instagram posts authored by over 700K accounts featuring sponsored content. Our analysis shows that per-post engagement is partly predictable and highly optimized when accounts are segmented by entity type or audience tier. Our models achieve competitive predictive power, while remaining fully transparent about which features drive engagement.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pedro Victor de Sousa Lima, Olga Goussevskaia. 2026-10-05. Predicting Engagement with Sponsored Content Across Account Types on Instagram. https://arxiv.org/abs/2610.06164

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Beyond the Clique: Comparing Clique and Dowker Complexes for Co-occurrence Data in Learning Analytics

Learning analytics increasingly represents the relations in co-occurrence data (codes in a window of discourse, participants in a thread, tags on a post) as simplicial complexes and analyses them with persistent homology. The standard approach uses the clique complex. Because its simplices are determined by the pairwise network alone, it cannot distinguish "three elements co-occurred together" from "each of the three pairs co-occurred separately". The Dowker complex, in contrast, takes as simplices the sets of entities actually observed to co-occur. We compare the two. First, we show that, on the same 1-skeleton, the Dowker complex is a subcomplex of the clique complex, the map on first homology induced by the inclusion is surjective, and its kernel is generated by phantom triangles that never co-occurred; that is, the clique construction can only erase holes. We also give a criterion for agreement that can be checked on triples of groups. We then test this on real data. On four Stack Exchange data sets, 46-91% of clique triangles are phantom and 522-1,970 holes are erased. On the example data of the learning-analytics R packages tna/Nestimate, the first Betti number $β_1$ of the clique complex is 0 at every threshold, whereas the Dowker complex detects holes. Against a degree-preserving null model, the Dowker complex departs strongly in five of the six data sets. This difference is invisible at the fixed thresholds used in practice. We conclude that the construction should follow the data type and that, for observed groups, the Dowker complex is the appropriate choice. Code is available at https://github.com/igu-lab/beyond-the-clique.

cs.SI↗

Two Americas of Well-Being: Divergent Rural-Urban Patterns of Life Satisfaction and Happiness from 2.6 B Social Media Posts

Using 2.6 billion geolocated social-media posts (2014-2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural-urban paradox: rural counties express higher life satisfaction while urban counties exhibit greater happiness. We reconcile this by treating the two as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Republican-leaning areas appear more satisfied in evaluative terms, but partisan gaps in happiness largely flatten outside major metros, indicating context-dependent political effects. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020-2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications and align with well-being theory. Interpreted as associations for the population of social-media posts, the results show that large-scale, language-based indicators can resolve conflicting findings about the rural-urban divide by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.

cs.SI↗

Linking Scalar-Intensity Language to Structural Polarization with Validated Signed-Network Measures

Polarization in online communities is often studied through either language or interaction structure, but the two views are rarely connected within a unified framework. Prior work has linked them by constructing interaction graphs from human judgements of agreement and disagreement, leaving a gap between language as observed text and structure as an engineered representation of that text. We address this gap with a language-grounded signed-network pipeline that derives signed relations directly from conversational exchanges and links window-level language patterns to structural polarization over time. Before examining this relationship, we compare spectral and frustration-based polarization measures on synthetic benchmarks and real interaction networks. We find that frustration-based measures normalized by the graph's cycle-space capacity provide a more suitable basis for comparing polarization across networks of different sizes and densities. We therefore carry forward two complementary frustration-based measures: a weighted form that incorporates stance-model confidence and a count form based on edge signs. The weighted measure provides the better-behaved structural estimate and aligns more closely with polarization measured from human-labelled interactions, while the count measure reveals a stronger relationship with language. Across monthly Reddit Brexit discussions, greater prevalence of scalar-intensity language is associated with greater structural polarization, with a similar rank-level pattern in the human-labelled network. We find little evidence that language in one month predicts polarization in the next, whereas contemporaneous scalar-intensity prevalence provides useful information about polarization within the same month.

cs.SI↗