arXiv Science⌕ Search

arXiv · 2610.11375

Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

Abstract

Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rena Gao, Yue Dai, Hao Guan, Shengxiang Gao, Wangyang Wu, Yixin Shen, Jey Han Lau. 2026-10-08. Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions. https://arxiv.org/abs/2610.11375

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps

Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework for transforming a natural-language MLOps cloud engineering task into a verified repository and operational cloud deployment. The framework combines graph engineering, loop engineering, and agent harness engineering. A stateful Graph Orchestrator coordinates specialized agents for repository generation, review, execution, verification, release, and monitoring while governing workflow dependencies, evidence gates, retry bounds, recovery paths, and termination. Consequential lifecycle transitions proceed only when their required predicates are supported by verifiable execution or runtime evidence. Verification failures activate bounded reflection, repair, and re-verification, while runtime evidence of failure, drift, degradation, or policy violation can trigger bounded adaptation, recovery, or rollback. Agent harness engineering constrains repository generation, review, and repair, artifact execution, and cloud operations through controlled capabilities and isolated execution environments. We realize the framework on Google Cloud Platform and evaluate repository completeness, controlled execution, evidence-gated transitions, cloud promotion, and bounded recovery. Our experimental results show that the framework prevents unsupported lifecycle transitions and drives each run toward either a verified operational deployment or an auditable terminal failure.

cs.MA↗

Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination

Does authority in AI teams improve the outcome? Organizational theory asserts that authority facilitates decision making, improving quality. Meanwhile, some nascent AI research suggests that revision under authority makes LLM output worse. Multi-agent LLM frameworks default to giving a Manager agent the authority to send a worker's output back for revision. Prior comparisons test the effect of authority using verifiable tasks. We conduct an experiment on an open-ended task, business-intelligence reporting, using a sample of 43 paired laptop products and 86 runs. Each report is written once by a hierarchical team and once by a flat team. We find that flat teams produce higher-quality reports, scoring higher on Utility (d = 0.42, p = 0.009) and Writing Clarity (d = 0.34, p = 0.030). The reports are the same length, but hierarchical team reports use 53% more hedging words such as "may" and "could", and each revision is associated with a 0.14-point drop in Writing Clarity on a 1 to 5 scale. Before any revision, the hierarchical team's first draft is indistinguishable from the flat team's report. In other words, the quality gap can be traced to revision. Authority improves quality when the Manager can verify the work, else when it can only provide feedback it has a negative effect on quality.

cs.MA↗

Can Agents Adapt Their Institutions And Sacrifice Themselves During Collapse? Executable Self-Governance in a GovSim Commons

LLM agents are beginning to write the rules they live under: policies, contracts and code that other agents review and approve. What happens when those rules must decide who is sacrificed from a group? We built GovSim-SelfGovern, a commons in which agents legislate in executable Python, test each law in a sandbox, vote on it, and live with the result. When the fishery is abundant, self-rule helps, raising the chance that the community survives without anyone starving from 52.5% to 75.0%. When it cannot feed everyone though, adding governance is not enough to save the commons. We trace much of this to the vote: agents readily draft laws that remove members, but then frequently vote them down, thereby dooming the commons. In a controlled replay, the same exile law passes 8% of the time when the sandbox preview names its victim at once and 55% when the removal is deferred until a future crisis. Showing voters the full source code does not close the gap, though most describe the removal in their own words. Nor is sacrifice a fixed trait. A single sentence of framing can swing self-removal from 2.5% to 91%, and due process deters agents more than considering the potential death of a peer. In contrast, give the community a treasury and the question of who must go almost disappears. Our work shows that what our collectives legislate is decided more by the form in which a choice is put to them as compared to their ethical beliefs.

cs.MA↗