arXiv ScienceSearch

arXiv subjects

Pan Hui

Publications and source records attributed to Pan Hui.

At least 19 recordsLinked to original sources

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.

cs.CV

Proceedings of The First Reflection in Creative Experience (RiCE) Workshop

Reflection and metacognition are central to the creative user experience. However, most HCI research on reflection focuses on clear, task-oriented goals such as to reflect on personal data or pedagogical outcomes. This contrasts with the open-ended and challenging to articulate goals of creative user experiences. For the first time, this workshop brings together interdisciplinary researchers, designers, educators, and artists across HCI, Cognitive Science, Design, AI, Learning Sciences, and Digital Art to examine reflection in creative interaction. The workshop will discuss themes, drawn from earlier discussions with HCI researchers and artists, on: how best to capture reflection in creative contexts, how to leverage the arts to support reflection for ethical change, and how to design creative AI that enhances - not hinders - critical thinking. By bringing interdisciplinary perspectives on reflection into discussion, the workshop will develop a guiding taxonomy for reflection in creative interaction to inform future creative practice and tool development.

cs.HC

Toward Site-Aware MR Art Exhibitions: A SLAM-Based Deployment Pipeline for Spatial Coherence and Exhibition Experience

Mixed Reality (MR) is increasingly being used in exhibition settings to bring digital artworks into relation with the physical environment. However, existing MR exhibition systems are often confined to prototypes or case-specific deployments, offering limited guidance for large-scale practical implementation. To address this gap, this paper presents a practical pipeline for designing and deploying large-scale MR art exhibitions, treating spatial alignment not only as a technical mechanism but also as an experiential design decision. We first conducted a pilot study comparing marker-based and Simultaneous Localization and Mapping (SLAM)-based alignment methods in an MR exhibition setting. Based on the results, we developed a SLAM-based pipeline for MR exhibitions that integrates technical deployment with exhibition curation. We then evaluated the pipeline through both system overhead measures and users' experiential feedback. The results show that spatial alignment influences not only technical stability, but also overall exhibition coherence, visitors' sense of continuity and immersion, and artwork interpretation. These findings provide an empirically grounded reference for future large-scale MR art exhibition deployment.

cs.MM

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWell, a large-scale benchmark designed to systematically evaluate the spatio-temporal reasoning capabilities of MLLMs for urban wellbeing analytics through joint modeling of satellite and street view imagery. UrbanWell spans 38 cities across multiple years and includes diverse indicators covering (1) environmental conditions (CO$_2$, NO$_2$, PM${2.5}$, and Normalized Difference Vegetation Index), (2) spatial accessibility (minimum distance to supermarkets and restaurants), (3) urban form (road length, road density, and land use), (4) urban vitality (population, economic activity diversity, and land use diversity), and (5) subjective perception attributes (e.g., safety, beauty, liveliness, wealth, and quietness). All indicators are aligned at grid level to enable standardized evaluation. Beyond static prediction, UrbanWell defines temporal reasoning tasks, including future value forecasting from historical observations and temporal trend classification. We benchmark 15 state-of-the-art representative MLLMs in a zero-shot setting, providing a comprehensive comparative evaluation across spatial and temporal dimensions. Experimental results indicate that while MLLMs capture salient spatial and perceptual cues, their performance varies substantially across heterogeneous urban indicators spanning environment and subjective perception. UrbanWell serves as a unified benchmark for evaluating multimodal spatial and temporal reasoning in urban wellbeing analytics, offering a standardized testbed for systematic assessment and future research on multimodal urban intelligence. Our codes and datasets are accessible via https://github.com/axin1301/UrbanWell-Benchmark.

cs.AI

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.

cs.CV

NexusAI: Enabling Design Space Exploration of Ideas through Cognitive Abstraction and Functional Decomposition

Large Language Models (LLMs) offer vast potential for creative ideation; however, their standard interaction paradigm often produces unstructured textual outputs that lead users to prematurely converge on sub-optimal ideas-a phenomenon known as fixation. While recent creativity tools have begun to structure these outputs, they remain compositionally opaque: ideas are organized as monolithic units that cannot be decomposed, abstracted, or recombinable at a sub-idea level. To address this, we propose Cognitive Abstraction (CA), a computational pipeline that transforms raw LLM-generated inspiration into a navigable and transformable design space. We implement this pipeline in NexusAI, a prototype diagramming system that supports (I) decomposition of inspiration into typed functional fragments, (II) multi-level abstraction to externalize mental scaling, and (III) cross-dimensional recombination to spark novel design directions. A within-subject user study (N=14) demonstrates that NexusAI significantly improves design space exploration, reduces cognitive overhead, and facilitates perspective reframing compared to a baseline. Our work contributes: (1) a characterization of "compositional opacity" as a barrier in human-AI co-creation; (2) the CA pipeline for operationalizing creative cognitive primitives at scale; and (3) empirical evidence that structured, multi-level representations can effectively mitigate fixation and support divergent exploration.

cs.HC

CogInstrument: Modeling Cognitive Processes for Bidirectional Human-LLM Alignment in Planning Tasks

Although Large Language Models (LLMs) demonstrate proficiency in knowledge-intensive tasks, current interfaces frequently precipitate cognitive misalignment by failing to externalize users' underlying reasoning structures. Existing tools typically represent intent as "flat lists," thereby disregarding the causal dependencies and revisable assumptions inherent in human decision-making. We introduce CogInstrument, a system that represents user reasoning through cognitive motifs-compositional, revisable units comprising concepts linked by causal dependencies. CogInstrument extracts these motifs from natural language interactions and renders them as editable graphical structures to facilitate bidirectional alignment. This structural externalization enables both the user and the LLM to inspect, negotiate, and reconcile reasoning processes iteratively. A within-subjects study (N=12) demonstrates that CogInstrument explicitly surfaces implicit reasoning structures, facilitating more targeted revision and reusability over conventional LLM-based dialogue interfaces. By enabling users to verify the logical grounding of LLM outputs, CogInstrument significantly enhances user agency, trust, and structural control over the collaboration. This work formalizes cognitive motifs as a fundamental unit for human-LLM alignment, providing a novel framework for achieving structured, reasoning-based human-AI collaboration.

cs.HC

Beyond Compliance: How AI Could Help Creative Writers by Refusing Them

Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight the need for designs fostering reflection on balanced AI and non-AI resource use. Theoretically, intentional AI non-compliance, refusals (saying ``no'' to requests), could introduce such reflection through friction stronger than other bypass-able solutions. Practically, refusal content/language characteristics lead to nuanced reactions. However, little research empirically focuses on nuances beyond mandatory ethical/technical constraints, on turning refusals into strategic friction for `innocuous' requests. We address this through a qualitative study with 22 creative writers, exploring reactions to refusals to common requests across writing stages (planning, translating, reviewing). Findings suggest that reflective potential depends on heterogeneous preference alignment along situational (e.g., convergent/divergent thinking phases), cognitive (e.g., domain beliefs), and relational (e.g., AI roles) dimensions. We discuss implications for creativity support, broader issues (e.g., AI addiction), and frictional/seamful AI design (e.g., integrating different compliance levels).

cs.HC

The Decline of Online Knowledge Communities: Obstacles, Workarounds, and Sustainability

Online knowledge communities (OKC) such as Stack Exchange, Reddit, and Zhihu have long functioned as socio technical infrastructures for collective problem solving. The rapid adoption of Generative AI (GenAI) introduces both complementarity and substitution. Large language models (LLMs) offer faster, more accessible drafts, yet divert traffic and contributions away from OKC that also provided their training data. To understand how communities adapt under this systemic shock, we report a mixed-methods study combining an online survey (N=217) and interviews with 11 current users. Findings show that while users increasingly rely on AI for convenience, they still turn to OKC for complex, ambiguous, or trust sensitive questions. Participants express polarized attitudes toward AI, reflecting divergent hopes and uncertainties about its role. Yet across perspectives, sustaining sociability, empathy, and reciprocity emerges as essential for community resilience. We argue that GenAI's impact constitutes not a terminal decline but a design challenge: to reimagine socio-technical complementarities that balance automation's efficiency with human judgment, trust, and collective stewardship in the evolving knowledge commons. To decline or sustain, it is now or never to take action.

cs.HC

FlexiCamAR: Enhancing Everyday Camera Interactions on AR Glasses with a Flexible Additional Viewpoint

The recent emergence and popularity of consumer-grade augmented reality (AR) glasses from major technology companies highlight their potential to become the next daily computing platform. A dominant design trend in this context is the integration of a front-facing camera to deliver a first-person perspective. While this approach is intuitive, there is limited evidence that it is optimal (or sufficient) for supporting users in daily tasks. This paper explores a more effective camera interaction technique for AR glasses, which we term ``FlexiCamAR." This novel method aims to enhance both efficiency and the range of applications for AR glasses by offering flexible and comfortable secondary camera viewpoints. To investigate the applicability and usability of this approach, we developed a ring camera prototype that can be attached to users' fingers. We then conducted a user study with 12 participants, comparing FlexiCamAR against the baseline, a traditional front-facing AR camera setup, across two common tasks: taking photos and scanning QR codes. Our findings show that FlexiCamAR significantly reduces physical load. We also explore potential scenarios where the additional viewpoint afforded by FlexiCamAR proves valuable, such as capturing low-angle perspectives or navigating confined spaces. Participant feedback further suggests strong potential for additional applications, including selfie taking, video conferencing, and object scanning. Overall, FlexiCamAR presents a novel interaction approach that can serve as a powerful supplement or alternative to the first-person perspective, significantly improving the adaptability of AR glasses for everyday use.

cs.HC

Dream the Dream: Futuring Communication between LGBTQ+ and Cisgender Groups in Metaverse

Digital platforms frequently reproduce heteronormative norms and structural biases, limiting inclusive communication between LGBTQ+ and cisgender individuals. The Metaverse, with its affordances for identity fluidity, presence, and community governance, offers a promising site for reimagining such interactions. To investigate this potential, we conducted participatory design workshops involving LGBTQ+ and cisgender participants, situating them in speculative Metaverse contexts to surface barriers and co-create alternative futures. The workshops followed a three-phase process-identifying challenges, speculative problem-solving, and visualizing futures-yielding socio-spatial-technical solutions across four layers: activity, interaction, scene, and space. These findings highlight the importance of spatial cues and power dynamics in shaping digital encounters. We contribute by (1) articulating challenges of cross-group communication in virtual environments, (2) proposing inclusive design opportunities for the Metaverse, and (3) advancing principles for addressing power geometry in digital space. This work demonstrates futuring as a critical strategy for designing equitable, transformative communication infrastructures.

cs.HC

Hyper-learning and Unlearning: A Narrative Speculation on Urbanism in Media Ecologies

Hyper-learning and Unlearning is a speculative animation that reflect how learning is reconfigured within digital media ecologies. Using architectural education as a microcosm, the work reframes the city as a hyper-learning apparatus where urban space, algorithmic systems, and platform infrastructures condition cognition and agency. By staging both hyper-learning and the unlearning induced by machine-supported cognition, the work critiques institutional gatekeeping while revealing how platforms reshape expertise, memory, and spatial experience. This project invites viewers to reconsider how urban space becomes pedagogical infrastructure in a posthumanism era.

cs.CY

Multimodal Cyber-physical Interaction in XR: Hybrid Doctoral Thesis Defense

Academic events, such as a doctoral thesis defense, are typically limited to either physical co-location or flat video conferencing, resulting in rigid participation formats and fragmented presence. We present a multimodal framework that breaks this binary by supporting a spectrum of participation - from in-person attendance to immersive virtual reality (VR) or browser access - and report our findings from using it to organize the first ever hybrid doctoral thesis defense using extended reality (XR). The framework integrates full-body motion tracking to synchronize the user's avatar motions and gestures, enabling natural interaction with onsite participants as well as body language and gestures with remote attendees in the virtual world. It leverages WebXR to provide cross-platform and instant accessibility with easy setup. User feedback analysis reveals positive VR experiences and demonstrates the framework's effectiveness in supporting various hybrid event activities.

cs.MM

Where Digital Meets Place: Deriving Strategies for Curating Mixed Reality Exhibitions in Public Spaces

Mixed Reality (MR) technologies are increasingly being used to enrich exhibitions and public spaces by blending digital content with the physical environment in real time. However, little is known about curatorial strategies for embedding MR exhibitions into public spaces or promoting audience experiences. To explore this, we designed and curated a campus-based MR art exhibition, using contextualism as the fundamental concept. We conducted an interdisciplinary expert focus group alongside exhibition viewing to identify opportunities, challenges, and design strategies from multiple perspectives. In parallel, we conducted user studies with general audiences to examine how curatorial strategies foster ex-periential qualities. Our findings reveal insights from both experts and general users along with strategies in curating MR exhibitions and highlight the foundational role of contextualism in curating MR art exhibitions in urban public spaces.

cs.HC

The AI Amplifier Effect: Defining Human-AI Intimacy and Romantic Relationships with Conversational AI

What does it mean to fall in love with something we know is virtual? The proliferation of conversational AI enables users to create customizable companions, fostering new intimate relationships that, while virtual, are perceived as authentic. However, public understanding of these bonds is limited, and platform policies regarding these interactions remain inconsistent. There is a pressing need for further HCI research to investigate: (a) the design affordances in AI that construct bonds and a sense of intimacy, (b) how such long-term engagement impacts users' real lives, and (c) how to balance user autonomy with platform regulation in the design of these systems without compromising users' well-being and experiences. This paper takes a step toward addressing these goals by providing a concrete definition of human AI intimacy based on in depth interviews with 30 users engaged in romantic relationships with AI companions. We elucidate the complexities of these relationships, from their formation to sustainability, and identify key features of the bonds formed. Notably, we introduce the AI Amplifier Effect, where the AI serves as a medium that intensifies the user's existing emotional state, leading to divergent positive, neutral, and negative impacts. We argue that designing for emotion must extend beyond technical affordances to encompass the essence of human affection. This paper's contributions aim to initiate a conversation and guide future research on human AI relationships within the HCI community.

cs.HC

SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment

Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial components like arms or interacting objects. Generative models are promising alternatives to generate diverse hand images, but still suffer from misalignment issues. In this paper, we present SesaHand, which enhances controllable hand image generation from both semantic and structural alignment perspectives for 3D hand reconstruction. Specifically, for semantic alignment, we propose a pipeline with Chain-of-Thought inference to extract human behavior semantics from image captions generated by the Vision-Language Model. This semantics suppresses human-irrelevant environmental details and ensures sufficient human-centric contexts for hand image generation. For structural alignment, we introduce hierarchical structural fusion to integrate structural information with different granularity for feature refinement to better align the hand and the overall human body in generated images. We further propose a hand structure attention enhancement method to efficiently enhance the model's attention on hand regions. Experiments demonstrate that our method not only outperforms prior work in generation performance but also improves 3D hand reconstruction with the generated hand images.

cs.CV

Urban mobility network centrality predicts social resilience

Cities thrive on social interactions that foster well-being, innovation, and prosperity; yet, exogenous shocks such as pandemics, hurricanes, and wildfires can severely disrupt them. Different urban venues exhibit widely divergent response patterns, raising key questions about what factors contribute to these differences and how we can anticipate and respond. Understanding these questions is crucial for safeguarding social resilience, the capacity of urban venues to maintain both visitation and diversity. In this study, we analyze large-scale human mobility data from 15 US cities covering more than 103 million residents across three distinct urban shocks. Despite a general trend of declining visitation and weakened social mixing, 36.28%-53.01% of venues exhibit reduced segregation, and 21.04%-38.55% of venues exhibit increased visitation. By constructing a mobility network interlinking types of urban venues, we reveal that eigenvector network centrality tends to indicate the provision of essential services and robustly predicts social resilience across varied urban shocks. Specifically, centrality elevates the explanatory power by more than 80% in predicting both segregation and mobility change, compared with more intuitive features. Furthermore, compared to peripheral venues, core venues featuring shorter visit distances, broader neighborhood visitation, shorter visitor dwell times, and steadier popularity throughout the day. Such patterns imply a dual social mechanism: core venues sustain social ties through frequent informal interaction, while peripheral ones facilitate deeper engagement around specialized interests and their corresponding social circles. By bridging urban mobility research with economic theories that distinguish staple from discretionary products, we propose a well-and-pool analogy that suggests how people spend their varying urban mobility budgets.

cs.SI

Meflex: A Multi-agent Scaffolding System for Entrepreneurial Ideation Iteration via Nonlinear Business Plan Writing

Business plan (BP) writing plays a key role in entrepreneurship education by helping learners construct, evaluate, and iteratively refine their ideas. However, conventional BP writing remains a rigid, linear process that often fails to reflect the dynamic and recursive nature of entrepreneurial ideation. This mismatch is particularly challenging for novice entrepreneurial students, who struggle with the substantial cognitive demands of developing and refining ideas. While reflection and meta-reflection are critical strategies for fostering divergent and convergent thinking, existing writing tools rarely scaffold these higher-order processes. To address this gap, we present the Meflex System, a large language model (LLM)-based writing tool that integrates BP writing scaffolding with a nonlinear idea canvas to support iterative ideation through reflection and meta-reflection. We report findings from an exploratory user study with 30 participants that examined the system's usability and cognitive impact. Results show that Meflex effectively scaffolds BP writing, promotes divergent thinking through LLM-supported reflection, and enhances meta-reflective awareness while reducing cognitive load during complex idea development. These findings highlight the potential of non-linear LLM-based writing tools to foster deeper and coherent entrepreneurial thinking.

cs.HC