arXiv · 2403.05006
Provable Pluralistic Alignment: Multi-Party RLHF under Offline Human Feedback
Abstract
Pluralistic alignment requires learning from feedback that reflects persistent and potentially conflicting stakeholder preferences while ultimately selecting a single collective policy. We study this problem in offline reinforcement learning from human feedback (RLHF), where the party associated with each comparison is observed. Under a shared low-rank linear reward model, we jointly estimate party-specific rewards and perform pessimistic policy optimization under Nash, Utilitarian, and Egalitarian social-welfare objectives. We establish nonasymptotic bounds for party-specific reward estimation and the resulting policy suboptimality under offline coverage conditions. We further consider general pairwise preferences that need not admit a scalar reward representation and may exhibit cycles. In this setting, we construct a pessimistic von Neumann winner policy and derive corresponding performance guarantees. Under these models, our results provide a unified finite-sample solution to a central challenge in pluralistic alignment: learning from limited, heterogeneous, and potentially cyclic feedback, and producing a single policy with explicit collective-welfare guarantees. Our framework thereby makes preference aggregation an explicit and statistically analyzable design choice rather than an implicit consequence of pooling human feedback.
Explore related subjects
Keep this discovery
Huiying Zhong, Tianwei Gao, Zhiwei Steven Wu, Linjun Zhang, Weijie J. Su, Zhun Deng. 2024-03-08. Provable Pluralistic Alignment: Multi-Party RLHF under Offline Human Feedback. https://arxiv.org/abs/2403.05006
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.