arXiv ScienceSearch

arXiv · 2409.04649

Preserving Individuality while Following the Crowd: Understanding the Role of User Taste and Crowd Wisdom in Online Product Rating Prediction

Abstract

Numerous algorithms have been developed for online product rating prediction, but the specific influence of user and product information in determining the final prediction score remains largely unexplored. Existing research often relies on narrowly defined data settings, which overlooks real-world challenges such as the cold-start problem, cross-category information utilization, and scalability and deployment issues. To delve deeper into these aspects, and particularly to uncover the roles of individual user taste and collective wisdom, we propose a unique and practical approach that emphasizes historical ratings at both the user and product levels, encapsulated using a continuously updated dynamic tree representation. This representation effectively captures the temporal dynamics of users and products, leverages user information across product categories, and provides a natural solution to the cold-start problem. Furthermore, we have developed an efficient data processing strategy that makes this approach highly scalable and easily deployable. Comprehensive experiments in real industry settings demonstrate the effectiveness of our approach. Notably, our findings reveal that individual taste dominates over collective wisdom in online product rating prediction, a perspective that contrasts with the commonly observed wisdom of the crowd phenomenon in other domains. This dominance of individual user taste is consistent across various model types, including the boosting tree model, recurrent neural network (RNN), and transformer-based architectures. This observation holds true across the overall population, within individual product categories, and in cold-start scenarios. Our findings underscore the significance of individual user tastes in the context of online product rating prediction and the robustness of our approach across different model architectures.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Liang Wang, Shubham Jain, Yingtong Dou, Junpeng Wang, Chin-Chia Michael Yeh, Yujie Fan, Prince Aboagye, Yan Zheng, Xin Dai, Zhongfang Zhuang, Uday Singh Saini, Wei Zhang. 2024-09-06. Preserving Individuality while Following the Crowd: Understanding the Role of User Taste and Crowd Wisdom in Online Product Rating Prediction. https://arxiv.org/abs/2409.04649

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Higher-order Network phenomena of cascading failures in resilient cities

Modern urban resilience is threatened by cascading failures in multimodal transport networks, where localized shocks trigger widespread paralysis. Existing models, limited by their focus on pairwise interactions, often underestimate this systemic risk. To address this, we introduce a framework that confronts higher-order network theory with empirical evidence from a large-scale, real-world multimodal transport network. Our findings confirm a fundamental duality: network integration enhances static robustness metrics but simultaneously creates the structural pathways for catastrophic cascades. Crucially, we uncover the source of this paradox: a profound disconnect between static network structure and dynamic functional failure. We provide strong evidence that metrics derived from the network's static blueprint-encompassing both conventional low-order centrality and novel higher-order structural analyses-are fundamentally disconnected from and thus poor predictors of a system's dynamic functional resilience. This result highlights the inherent limitations of static analysis and underscores the need for a paradigm shift towards dynamic models to design and manage truly resilient urban systems.

cs.SI

Measuring Time-Horizon Engagement Effectiveness: Persistence, Recency, and Re-Emergence

Online attention is commonly summarized using cumulative volume, peak activity, or arithmetic averages, but such measures can obscure differences between activity that is sustained over time, concentrated near the present, or renewed after dormancy. This paper introduces the Time-Horizon Engagement Effectiveness (TH-EE) framework, which constructs interpretable temporal profiles of online attention by distinguishing three related but non-equivalent properties: persistence, recency, and re-emergence. We evaluate the framework through controlled engagement traces, a proof-of-concept application to 1,850 YouTube videos across 37 topics (18 cohorts selected a priori as exemplars of persistent, acute, cyclical, and recently originating attention, and 19 cohorts corresponding to authoritatively debunked claims), and an event-level validation on eleven years of daily Wikipedia pageview series for the same topics. The controlled analyses show that the framework distinguishes distributed activity from concentrated bursts, introduces temporal-order sensitivity through recency weighting, and identifies renewed activity after a defined dormant interval. On the event-level series, the framework's reactivations co-locate with Kleinberg burst onsets, and PELT change points far more often than chance. The YouTube application shows that debunked-claim cohorts do not occupy a unique region of temporal-profile space; they exhibit heterogeneous patterns that overlap substantially with benign topics. These results support treating persistence, recency, and re-emergence as separate dimensions of online attention. TH-EE is a descriptive and comparative measurement framework, not a classifier of misinformation, coordination, intent, or content veracity.

cs.SI

Enhancing Human Mobility Prediction with Spatially Aware LLM-based Multi-Agent Systems

Predicting a user's next POI is a task in human mobility modeling, yet LLM-based approaches focus on semantic reasoning from previous mobility records, while neglecting real-world spatial context. However, human mobility is inherently shaped by spatial cognition, including geographic distance and neighborhood context. This issue is further compounded by prior evidence that LLMs often struggle with spatial reasoning tasks, including distance estimation and geographically biased prediction. To address these limitations, we propose our framework, a multi-agent LLM framework that decomposes next-POI prediction into three stages: Firstly, a Pattern Extraction Agent that captures temporal and categorical mobility patterns from trajectory history; Secondly, a Spatial Reasoning Agent that structures candidate activity choices by combining behavioral preferences with real-world spatial constraints, including geographic distance, road network distance, and neighborhood affiliation; and Thirdly, a Decision Synthesis Agent that integrates behavioral patterns and spatial reasoning for final prediction. Experiments on the NYC benchmark dataset with two LLM backbones show improvements over baseline methods, with up to 493% Hit@1 improvement and 37% relative improvement in Hit@5. Ablations show that combining neighborhood affiliation with distance-based features generally outperforms distance-only settings, and that the Spatial Reasoning Agent plays a crucial role in final prediction by integrating behavioral preferences with real-world spatial constraints, especially for smaller models. Overall, the results highlight the importance of spatial reasoning in mobility prediction. Accurate next-POI prediction requires combining behavioral patterns with explicit real-world spatial constraints, and multi-agent decomposition provides an effective structure for organizing these forms of context.

cs.SI