arXiv ScienceSearch

arXiv subjects

Xinyi Che

Publications and source records attributed to Xinyi Che.

6 recordsLinked to original sources

RobotEQ: Towards Social Proactive Intelligence in Embodied Agents

Embodied agents represent a prominent research focus across both academia and industry. The prevailing paradigm has gradually shifted from reactive assistance, which requires explicit user queries, to proactive assistance, capable of recognizing human needs and offering support without explicit instructions. Nevertheless, existing studies on proactive assistance remain confined to narrow scenarios and primarily emphasize task completeness, whereas real-world agents must operate in open-domain environments while adhering to social expectations. To bridge this gap, we extend the concept of proactive assistance to Social Proactive Intelligence (SPI), characterized by diverse scenarios, social understanding, and robot-centric behaviors. We further introduce RobotEQ, a dedicated benchmark for SPI. We first define two tasks: behavior judgment, emphasizing global contextual understanding, and spatial grounding, focusing on local perceptual details. Building on these tasks, we construct RobotEQ-Data, a dataset comprising 1,812 synthetic and 223 real-world scenarios, 7 social facets, 22K+ human annotations, 3K+ behavior judgment questions, and 3K+ spatial grounding questions. Furthermore, we establish RobotEQ-Bench to evaluate the performance of representative models. Experimental results demonstrate that current models fall short of achieving reliable SPI. Further analysis reveals that incorporating external social knowledge yields consistent improvements. This work aims to advance the development of socially desirable embodied agents in open-domain environments.

cs.RO

MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding

MER2026 marks the fourth edition of the MER series of challenges. The MER series provides valuable data resources to the research community and offers tasks centered on recent research trends, establishing itself as one of the largest challenges in the field. Throughout its history, the focus of MER has shifted from discriminative emotion recognition to generative emotion understanding. Specifically, MER2023 concentrated on discriminative emotion recognition, restricting the emotion recognition scope to fixed basic labels. In MER2024 and MER2025, we transitioned to generative emotion understanding and introduced two new tasks: fine-grained emotion recognition and descriptive emotion analysis, aiming to leverage the extensive vocabulary and multimodal understanding capabilities of Multimodal Large Language Models (MLLMs) to facilitate fine-grained and explainable emotion recognition. Building on this trajectory, MER2026 continues to follow these research trends and contains four tracks: MER-Cross shifts the focus from individual to dyadic interaction scenarios; MER-FG centers on fine-grained emotion recognition; MER-Prefer aims to predict human preferences regarding different emotion descriptions; MER-PS focuses on emotion recognition based on physiological signals. More details regarding the dataset and baselines are available at https://zeroqiaoba.github.io/MER-Challenge/.

cs.HC

Constraining the dynamical Chern-Simons gravity with future gravitational wave detectors

Dynamical Chern-Simons gravity, a parity-violating modification of general relativity, is regarded as a low-energy effective theory arising from string theory. Gravitational waves provide a powerful probe for testing its predictions. However, current gravitational wave observations are unable to place meaningful constraints on this theory through phase measurements, due to limitations from detector noise and the validity requirements of the waveform models. In this paper, we conduct a comprehensive assessment of the prospects for constraining the dynamical Chern-Simons gravity with future gravitational-wave detectors using stellar mass black holes binary. We quantify how the constraining capacities vary across different detectors and source parameters, and identify the regions of parameter space that satisfy the small-coupling condition. Furthermore, by incorporating an astrophysically motivated mass distribution model for stellar mass black hole binaries, we estimate the potential of upcoming observatories.

gr-qc

Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation

Multimodal Emotion Recognition in Conversation (MERC) aims to enhance emotion understanding by integrating complementary cues from text, audio, and visual modalities. Existing MERC approaches predominantly focus on cross-modal shared features, often overlooking modality-specific features that capture subtle yet critical emotional cues such as micro-expressions, prosodic variations, and sarcasm. Although related work in multimodal emotion recognition (MER) has explored disentangling shared and modality-specific features, these methods typically employ rigid orthogonal constraints to achieve full disentanglement, which neglects the inherent complementarity between feature types and may limit recognition performance. To address these challenges, we propose Angle-Optimized Feature Learning (AO-FL), a framework tailored for MERC that achieves partial disentanglement of shared and specific features within each modality through adaptive angular optimization. Specifically, AO-FL aligns shared features across modalities to ensure semantic consistency, and within each modality it adaptively models the angular relationship between its shared and modality-specific features to preserve both distinctiveness and complementarity. An orthogonal projection refinement further removes redundancy in specific features and enriches shared features with contextual information, yielding more discriminative multimodal representations. Extensive experiments confirm the effectiveness of AO-FL for MERC, demonstrating superior performance over state-of-the-art approaches. Moreover, AO-FL can be seamlessly integrated with various unimodal feature extractors and extended to other multimodal fusion tasks, such as MER, thereby highlighting its strong generalization beyond MERC.

cs.MM

Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation

Multimodal Emotion Recognition in Conversation (MERC) significantly enhances emotion recognition performance by integrating complementary emotional cues from text, audio, and visual modalities. While existing methods commonly utilize techniques such as contrastive learning and cross-attention mechanisms to align cross-modal emotional semantics, they typically overlook modality-specific emotional nuances like micro-expressions, tone variations, and sarcastic language. To overcome these limitations, we propose Orthogonal Disentanglement with Projected Feature Alignment (OD-PFA), a novel framework designed explicitly to capture both shared semantics and modality-specific emotional cues. Our approach first decouples unimodal features into shared and modality-specific components. An orthogonal disentanglement strategy (OD) enforces effective separation between these components, aided by a reconstruction loss to maintain critical emotional information from each modality. Additionally, a projected feature alignment strategy (PFA) maps shared features across modalities into a common latent space and applies a cross-modal consistency alignment loss to enhance semantic coherence. Extensive evaluations on widely-used benchmark datasets, IEMOCAP and MELD, demonstrate effectiveness of our proposed OD-PFA multimodal emotion recognition tasks, as compared with the state-of-the-art approaches.

cs.MM

Gravitational waves and cosmic boundary

Space-based gravitational wave detectors have the capability to detect signals from very high redshifts. It is interesting to know if such capability can be used to study the global structure of the cosmic space. In this paper, we focus on one particular question: if there exists a reflective cosmic boundary at the high redshift ($z>15$), is it possible to find it? We find that, with the current level of technology: 1) gravitational waves appear to be the only means with which that signatures from the cosmic boundary can possibly be detected; 2) a large variety of black holes, with masses roughly in the range $(10^3\sim 10^6) {\rm~M_\odot}$, can be used for the task; 3) in the presumably rare but physically possible case that two merger events from the growth history of a massive black hole are detected coincidentally, a detector network like TianQin+LISA is essential in help improving the chance to determine the orientation of the cosmic boundary; 4) the possibility to prove or disprove the presence of the cosmic boundary largely depends on how likely one can detect multiple pairs of coincident gravitational wave events.

gr-qc