arXiv ScienceSearch

arXiv subjects

Yuanchun Shi

Publications and source records attributed to Yuanchun Shi.

At least 19 recordsLinked to original sources

StoryEcho: A Narrative Mirroring Loop Generative Storytelling System for Picky-Eating Intervention

Picky eating can limit children's dietary variety and create tension in family feeding routines. Existing food-related technologies often focus on mealtime intervention or standalone educational artifacts, offering limited support for connecting low-pressure narrative engagement with children's real-world food exploration over time. We present StoryEcho, a generative storytelling system centered on a narrative mirroring loop, in which personalized stories model sensory exploration through a persistent counterpart and children's subsequent food encounters are reflected back into narrative feedback and future story development. Informed by a formative study, we designed StoryEcho and evaluated it in a 14-day between-subjects field study with 26 families. Compared with food-personalized generative stories without the narrative mirroring loop, StoryEcho was associated with higher try level, approach, and intake, and lower resistance and caregiver pressure. These findings suggest that narrative mirroring can support children's low-pressure food exploration in family routines, while highlighting design tensions for future generative storytelling interventions.

cs.HC

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.

cs.CV

U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses

Uncertainty can appear throughout LLM-generated text (e.g., questionable claims, ambiguous terms). Prior work largely focuses on making such uncertainty visible through cues such as confidence scores, but seeing uncertainty is not the same as managing it. Through a formative study, we examine uncertainty management across interpretation, evaluation, and decision. From these insights, we derive design guidelines for uncertainty target representation, evaluative explanation, response guidance, and interactive presentation. We instantiate them in U-Lens, an uncertainty-management system that organizes uncertain information into contextual inspection targets, prioritizes them, and links each to evaluative context and response options. In an 18-participant within-subjects study comparing U-Lens with a confidence-cue baseline, U-Lens improved verification efficiency and effort allocation, reduced perceived workload, and strengthened support across all three stages. This work reframes uncertainty support for generative AI from text-centered cues to a user-centered process of interpreting, evaluating, and acting on uncertainty.

cs.HC

TraceMind: Predicting User Information Uptake from Low-Cost Interaction Traces during Human-LLM Content Co-Generation

In human-LLM content co-generation, AI-generated information can enter final artifacts without being adequately processed by users, creating risks when artifacts are shared or acted upon. We study whether recognition-level uptake of atomic information units can be assessed in open-ended co-generation and predicted from low-cost interaction traces. We collected data from 62 participants across three tasks. For each final draft, we extracted atomic information units and generated post-task recognition questions, yielding 1187 unit-level uptake labels. We present TraceMind, which tracks units across Chat and Draft histories, aligns interaction traces with changing on-screen layouts, and models spatial, temporal, and workflow-informed evidence. TraceMind outperformed all learned baselines across AUROC, AUPRC-non, balanced accuracy, and macro-F1. We found that uptake unfolds throughout interaction, with sustained active engagement providing informative evidence beyond isolated signals. Our work shifts human-LLM co-generation from content adoption toward what users actually take up, motivating uptake-aware systems grounded in low-cost interaction traces.

cs.HC

Reconstruction and Reflection of Positive Experiences through Resurfacing Laughter-indexed Everyday Moments

Positive everyday moments often escape deliberate recording, while continuous self-tracking can generate extensive records that are difficult to revisit. We explore laughter as a naturally occurring, sparse index for constructing contextualized personal records to support later reconstruction and reflection. A formative study with 12 participants characterized laughter as an affective but semantically incomplete index and informed \textit{LaughAnchor}, a mobile and wearable self-tracking system. During participant-initiated recording, the system assembles detected laughter and aligned context into candidate moments for later reconstruction and reflection, with layered context disclosure, user-controlled curation, and near-term and long-term resurfacing. In a three-week field deployment with 12 participants, passive indexing preserved moments they considered unlikely to record deliberately but valued retrospectively. During resurfacing, participants attributed affective re-experiencing to laughter and used additional context both to reconstruct episodes and to explore already-recalled experiences. Across moments and reviews, resurfacing supported rediscovery and broader awareness of relationships, routines, and emotional states. These findings inform self-tracking designs that use sparse affective indices to organize contextual records for reconstruction and reflection, while keeping interpretation and retention under user control.

cs.HC

Earinter: A Closed-Loop System for Eating Pace Regulation with Just-in-Time Intervention Using Commodity Earbuds

Rapid eating is common yet difficult to regulate in situ, partly because people seldom notice pace changes and sustained self-monitoring is effortful. We present Earinter, a commodity-earbud-based closed-loop system that integrates in-the-wild sensing, real-time reasoning, and theory-grounded just-in-time (JIT) intervention to regulate eating pace during daily meals. Earinter repurposes the earbud's bone-conduction voice sensor to capture chewing-related vibrations and estimate eating pace as chews per swallow (CPS) for on-device inference. With data collected equally across \modify{sessions in quiet and noisy environments}, Earinter achieves reliable chewing detection (F1 = 0.97) and accurate eating pace estimation (MAE: 0.18 $\pm$ 0.13 chews/min, 3.65 $\pm$ 3.86 chews/swallow), enabling robust tracking for closed-loop use. Guided by Dual Systems Theory and refined through two Wizard-of-Oz pilots, Earinter adopts a user-friendly design for JIT intervention content and delivery policy in daily meals. In a 13-day within-subject field study (N=14), the closed-loop system significantly increased CPS and slowed meal-level eating pace, with statistical signs of carryover on retention-probe days and generally favorable comfort and usability. These effects should be interpreted as changes in eating-pace regulation rather than evidence of weight loss. Our findings highlight how single-modality commodity earables can support practical, theory-driven closed-loop JIT interventions for regulating eating pace in the wild.

cs.HC

Towards Wearable Opportunistic Crowdsensing for Open-Vocabulary Activity Data Collection Through User-Scheduled Trigger-Action Routines

Collecting richly labeled wearable activity data in everyday settings remains difficult because retrospective annotation is costly and often imprecise. Prior data collection apps rely on a labor-intensive self-reporting strategy and primarily treat participants as crowd labelers. We present Pebbl, a feasibility-stage system that incentivizes in-situ labeling through opportunistic crowdsensing. Pebbl lets users author trigger-action recipes on a smartphone and receive just-in-time reminders for beneficial actions when a trigger is detected. In the prototype, triggers are a limited set with four common audio cues, while actions are described in open-vocabulary natural language. Each confirmed execution yields a short sensor window with explicit start/end boundaries and a user-authored action label. We evaluate Pebbl through an expert workshop with wearable Human Activity Recognition (HAR) researchers (N = 6), a within-subject in-lab study (N = 21), and a pilot deployment (N = 8). Experts viewed the approach as lower burden and more ecologically valid than common labeling workflows. In the lab, Pebbl produced reliable execution logs under controlled conditions (recall = 97.30%, precision = 97.15%) and was preferred over comparison workflows on perceived burden and confidence. The pilot deployment shows that the interaction and sensing pipeline can function in free-living use, while surfacing practical constraints such as false triggers and context dependence. Overall, Pebbl represents a step toward a low-burden, distributable collection approach of user-contributed wearable activity data.

cs.HC

VergeIO: Depth-Aware Eye Interaction on Glasses

There is growing industry interest in unobtrusive designs for electrooculography (EOG) sensing of eye gestures on glasses (e.g. JINS MEME and Apple eyewear). We present VergeIO, an EOG-based glasses system that enables depth-aware eye interaction by sensing vergence with a glasses-compatible electrode layout and smart glass prototype. It can distinguish between four depth-based eye gestures with 97% accuracy on unseen users without any calibration in a user study across 20 users and 1,520 gesture instances. To reduce false detections, we incorporate a motion artifact detection pipeline and a preamble-based activation scheme. The system uses dry sensors without any adhesives or gel and operates in real time with 3 mW power consumption by the analog sensing front-end.

cs.HC

PAGE: Towards Practical Human-level Gaze Target Estimation

Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation). As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2. The distilled student models retain most of the teacher's performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at our project page.

cs.CV

AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation

We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and two additional side-view cameras, paired with precise screen-space gaze targets collected under controlled fixation conditions. Each sample contains multi-view face observations together with structured facial region crops, enabling multimodal learning from both global and local visual cues. Unlike existing single-view gaze datasets, AA provides multi-view coverage from both screen-mounted and side-mounted perspectives, enabling more robust modeling under viewpoint variation and occlusion. The dataset includes subject-independent evaluation splits and a standardized data processing pipeline to support reproducible research in gaze estimation.

cs.CV

PAPEL: A Collaborative System for Parental Guidance during Preschool Play-Based English Learning

Play-based parent-child interaction offers preschoolers rich opportunities for everyday foreign language learning, yet many parents struggle to turn open-ended play into effective English-as-a-Foreign-Language (EFL) learning experiences at home. To explore how AI might support this process, we conducted formative studies through interviews and a Wizard-of-Oz study. We identified four key challenges: content selection, language expression, balancing instruction and play, and problem solving. To address these challenges, we present PAPEL, a parent-AI collaborative system that grounds suggestions in the ongoing play scene and organizes support into four core modules: content generation, language adaptation, balance assessment, and extended response. In a counterbalanced within-subjects study with 16 parent-child dyads, PAPEL was associated with more integrated parent utterances that combined playful and instructional content, as well as more parent-child conversational turns, than the lightweight chatbot baseline used in our study.

cs.HC

PACT: Proactive Asking for Continual Task Assistance in Human-Robot Collaboration

Robotic assistants in long-term human-robot collaboration need to assist users under partial observations while leveraging cross-day interaction history. However, human traits and routines are often unknown at the beginning of collaboration, making passive infer-then-act assistance ineffective and inefficient. To address this challenge, we study a cross-day proactive asking setting for continual task assistance and propose PACT (Proactive Asking for Continual Task Assistance), an ask-or-act framework that determines whether clarification should be sought before taking action. PACT leverages current observations together with accumulated interaction history to evaluate contextual sufficiency, enabling the robot to provide more reliable assistance and progressively adapt to the user over time. We implement its primary learned instantiation using reinforcement learning and evaluate alternative instantiations under the same framework. To assess such behavior, we further introduce a clarification utility metric that quantifies the trade-off between assistance accuracy and the frequency of clarification requests. Experiments in multi-day embodied collaboration scenarios demonstrate that, compared with passive inference baselines, PACT consistently improves both assistance accuracy and clarification utility, highlighting the importance of proactive asking in continual human-robot collaboration.

cs.RO

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we introduce EgoIntrospect, the first egocentric dataset captured in user-driven scenarios with self-annotations that explicitly reveal users' interactive intentions with AI assistants. EgoIntrospect was collected using a cross-device setup, providing synchronized video, audio, gaze, motion, and physiological signals. It consists of 180 hours of recordings from 60 subjects, with an average recording duration of 3 hours per subject. Leveraging EgoIntrospect, we formalize a suite of tasks centered on user internal states, including affective experience, interactive intent, and cognitive memory. We further process the annotations to construct benchmarks that evaluate the ability of modern multimodal large language models to reason about users' internal states from egocentric observations. Experiments on our benchmark suggest that existing multimodal large language models struggle to effectively leverage multimodal signals to infer users' subjective internal states. The dataset and annotations will be made publicly available to advance research in egocentric vision and wearable AI assistants. Project page: https://ego-introspect.github.io/

cs.CV

SpeakSoftly: Scaffolding Nonviolent Communication in Intimate Relationships through LLM-Powered Just-In-Time Interventions

Conflicts are common in text-based communication, particularly in intimate relationships, where misunderstandings can easily escalate into verbal aggression. To address this, we present SpeakSoftly, a system that applies Nonviolent Communication (NVC) principles to scaffold couples' conflict communication through LLM-powered just-in-time interventions. Informed by formative interviews with couples and NVC principles, we designed two core features: NVC-Prompt, which detects verbal aggression and suggests revisions to prevent escalation, and NVC-Guide, which analyzes dialogues to uncover users' feelings and needs, fostering self-awareness and perspective-taking. These features were implemented across three progressive intervention modes, each varying in intervention depth and tone: Basic Reminder, Neutral Guide, and Empathetic Guide. We conducted a mixed-methods user study with 18 couples across simulated and real-life conflict settings to evaluate the effectiveness of each mode. Results showed that Empathetic Guide significantly facilitated both behavioral and cognitive changes, while Neutral Guide was effective only for behavioral changes in simulated conflicts. In real-life conflicts, Neutral Guide showed distinct advantages due to lower cognitive load demands. We discuss the mechanisms behind these findings and propose design implications for in-situ interventions in high-stakes communication contexts.

cs.HC

Adapting AI to the Moment: Understanding the Dynamics of Parent-AI Collaboration Modes in Real-Time Conversations with Children

Parent-AI collaboration to support real-time conversations with children is challenging due to the sensitivity and open-ended nature of such interactions. Existing systems often simplify collaboration into static modes, providing limited support for adapting AI to continuously evolving conversational contexts. To address this gap, we systematically investigate the dynamics of parent-AI collaboration modes in real-time conversations with children. We conducted a co-design study with eight parents and developed COMPASS, a research probe that enables flexible combinations of parental support functions during conversations. Using COMPASS, we conducted a lab-based study with 21 parent-child pairs. We show that parent-AI collaboration unfolds through evolving modes that adapt systematically to contextual factors. We further identify three types of parental strategies--parent-oriented, child-oriented, and relationship-oriented--that shape how parents engage with AI. These findings advance the understanding of dynamic human-AI collaboration in relational, high-stakes settings and inform the design of flexible, context-adaptive parental support systems.

cs.HC

Exploring and Analyzing the Effect of Avatar's Visual Style on Anxiety of English as Second Language (ESL) Speakers

Virtual avatars offer new opportunities to reshape communication experiences beyond traditional live video. However, it remains unclear how avatar representations influence communication anxiety for English as a Second Language (ESL) speakers, and why such effects emerge. To take a first step to address this, we conducted a controlled laboratory study in which Mandarin-speaking ESL participants engaged in one-on-one conversations under three representation conditions: live video, stylized avatars, and realistic avatars. We assessed anxiety using both self-reported measures and physiological signals (EDA, ECG, PPG). Our results show that avatar style plays a critical role in shaping communication anxiety. While live video remained a strong baseline with low subjective anxiety, stylized avatars achieved comparable-and in some cases lower-physiological anxiety levels, whereas realistic avatars elicited higher anxiety. Beyond these effects, our findings reveal three underlying mechanisms that explain how avatar representations shape ESL communication anxiety: (1) facial expressiveness; (2) perceived feedback and fear of negative evaluation; and (3) contextual appropriateness. This work provides actionable design implications for developing avatar-mediated communication systems that support emotionally sustainable cross-linguistic interaction.

cs.HC

PACEE: Parent-Centered AI Scaffolding for Emotion Education in Early Childhood Conversations

Emotion education is critical for children aged 3 to 6. However, existing technologies largely focus on children's direct interaction with AI, overlooking the central role of parents in guiding early emotional development at home. To address this gap, we conducted co-design sessions with five kindergarten teachers and five parents to identify key parental challenges and opportunities for AI support in family emotion education. Based on these insights, we developed PACEE, an LLM-based assistant designed to support parents in guiding children's emotional development through conversations, rather than directly interacting with children. PACEE provides parent-centered AI scaffolding that supports parents in real-time conversation through personalized guidance, post-hoc reflection through trackable feedback, and understanding children's emotional states through modeling. We evaluated PACEE with 16 families. Results show that PACEE enhances parent-child engagement, fosters deeper emotional communication, and improves parents' expertise and overall experience in guiding their children. Our findings extend emotion coaching practices to the context of generative AI and offer design insights for building AI systems that support parent-centered family education.

cs.HC

Routine Computing: A Systematic Review of Sensing Daily Life Dimensions Towards Human-Centered Goals

Human routines structure daily life, yet remain challenging for computational systems to understand. This paper presents the first systematic review of routine computing, a previously implicit but increasingly recognized field that focuses on computationally sensing and modeling human behaviors. It synthesizes 203 studies published up to August 2025. The paper presents a new taxonomy of the literature, focusing on temporal structures, behavioral interactions, cognitive aspects, and how variability and deviations are addressed. The common goals of routine computing extend across four major application domains, including accessibility care, the promotion of healthy habits, adaptive and context-aware support, and large-scale population insights. Persistent challenges that limit the design of truly human-centered systems are identified, including the gap between low-level activity recognition and high-level intent, the tension between personalization and generalization, unresolved privacy concerns, and data-related limitations. By consolidating these findings, this paper provides a foundational framework for HCI researchers, outlining principles for designing ethical, adaptive, and human-centered routine-aware systems.

cs.HC