arXiv ScienceSearch

arXiv subjects

Sejoon Lim

Publications and source records attributed to Sejoon Lim.

11 recordsLinked to original sources

Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality, such as facial expressions or speech, is often insufficient for capturing these complex emotional cues. To address this limitation, we propose a multimodal emotion recognition framework for the Expression (EXPR) task in the 10th Affective Behavior Analysis in-the-wild (ABAW) Challenge. Our framework builds on large-scale pre-trained models for visual and audio representation learning and integrates them in a unified multimodal architecture. To better capture temporal patterns in facial expression sequences, we incorporate temporal visual modeling over video windows. We further introduce a bi-directional cross-attention fusion module that enables visual and audio features to interact in a symmetric manner, facilitating cross-modal contextualization and complementary emotion understanding. In addition, we employ a text-guided contrastive objective to encourage semantically meaningful visual representations through alignment with emotion-related text prompts. Experimental results on the ABAW 10th EXPR benchmark demonstrate the effectiveness of the proposed framework, achieving a Macro F1 score of 0.32 compared to the baseline score of 0.25, and highlight the benefit of combining temporal visual modeling, audio representation learning, and cross-modal fusion for robust emotion recognition in unconstrained real-world environments.

cs.CV

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision-language models such as CLIP have demonstrated remarkable semantic alignment capabilities, their application to continuous regression tasks is often limited by the discrete nature of text prompts. In this paper, we propose a novel multimodal framework for VA estimation that introduces Distance-aware Soft Prompt Guidance to bridge the gap between semantic representations and continuous affective dimensions. Specifically, we partition the VA space into multiple discrete regions, each associated with distinct textual descriptions. Rather than relying on hard categorization, we employ a Gaussian kernel to compute soft labels based on the Euclidean distance between the ground-truth coordinates and the region centers, allowing the model to learn fine-grained emotional transitions. For multimodal integration, our architecture utilizes a CLIP image encoder and an Audio Spectrogram Transformer to extract robust visual and acoustic features. These features are temporally modeled using Gated Recurrent Units and integrated through a hierarchical fusion scheme that sequentially combines cross-modal attention for alignment and gated fusion for adaptive refinement. Experimental results on the Aff-Wild2 dataset show that the proposed semantic-guided approach outperforms the official baseline and demonstrates robust performance on in-the-wild data.

cs.CV

Emotion Recognition Using Transformers with Masked Learning

In recent years, deep learning has achieved innovative advancements in various fields, including the analysis of human emotions and behaviors. Initiatives such as the Affective Behavior Analysis in-the-wild (ABAW) competition have been particularly instrumental in driving research in this area by providing diverse and challenging datasets that enable precise evaluation of complex emotional states. This study leverages the Vision Transformer (ViT) and Transformer models to focus on the estimation of Valence-Arousal (VA), which signifies the positivity and intensity of emotions, recognition of various facial expressions, and detection of Action Units (AU) representing fundamental muscle movements. This approach transcends traditional Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) based methods, proposing a new Transformer-based framework that maximizes the understanding of temporal and spatial features. The core contributions of this research include the introduction of a learning technique through random frame masking and the application of Focal loss adapted for imbalanced data, enhancing the accuracy and applicability of emotion and behavior analysis in real-world settings. This approach is expected to contribute to the advancement of emotional computing and deep learning methodologies.

cs.CV

PU-MFA : Point Cloud Up-sampling via Multi-scale Features Attention

Recently, research using point clouds has been increasing with the development of 3D scanner technology. According to this trend, the demand for high-quality point clouds is increasing, but there is still a problem with the high cost of obtaining high-quality point clouds. Therefore, with the recent remarkable development of deep learning, point cloud up-sampling research, which uses deep learning to generate high-quality point clouds from low-quality point clouds, is one of the fields attracting considerable attention. This paper proposes a new point cloud up-sampling method called Point cloud Up-sampling via Multi-scale Features Attention (PU-MFA). Inspired by previous studies that reported good performance using the multi-scale features or attention mechanisms, PU-MFA merges the two through a U-Net structure. In addition, PU-MFA adaptively uses multi-scale features to refine the global features effectively. The performance of PU-MFA was compared with other state-of-the-art methods through various experiments using the PU-GAN dataset, which is a synthetic point cloud dataset, and the KITTI dataset, which is the real-scanned point cloud dataset. In various experimental results, PU-MFA showed superior performance in quantitative and qualitative evaluation compared to other state-of-the-art methods, proving the effectiveness of the proposed method. The attention map of PU-MFA was also visualized to show the effect of multi-scale features.

cs.CV

BYEL : Bootstrap Your Emotion Latent

With the improved performance of deep learning, the number of studies trying to apply deep learning to human emotion analysis is increasing rapidly. But even with this trend going on, it is still difficult to obtain high-quality images and annotations. For this reason, the Learning from Synthetic Data (LSD) Challenge, which learns from synthetic images and infers from real images, is one of the most interesting areas. In general, Domain Adaptation methods are widely used to address LSD challenges, but there is a limitation that target domains (real images) are still needed. Focusing on these limitations, we propose a framework Bootstrap Your Emotion Latent (BYEL), which uses only synthetic images in training. BYEL is implemented by adding Emotion Classifiers and Emotion Vector Subtraction to the BYOL framework that performs well in Self-Supervised Representation Learning. We train our framework using synthetic images generated from the Aff-wild2 dataset and evaluate it using real images from the Aff-wild2 dataset. The result shows that our framework (0.3084) performs 2.8% higher than the baseline (0.3) on the macro F1 score metric.

cs.LG

Multitask Emotion Recognition Model with Knowledge Distillation and Task Discriminator

Due to the collection of big data and the development of deep learning, research to predict human emotions in the wild is being actively conducted. We designed a multi-task model using ABAW dataset to predict valence-arousal, expression, and action unit through audio data and face images at in real world. We trained model from the incomplete label by applying the knowledge distillation technique. The teacher model was trained as a supervised learning method, and the student model was trained by using the output of the teacher model as a soft label. As a result we achieved 2.40 in Multi Task Learning task validation dataset.

cs.CV

Causal affect prediction model using a facial image sequence

Among human affective behavior research, facial expression recognition research is improving in performance along with the development of deep learning. However, for improved performance, not only past images but also future images should be used along with corresponding facial images, but there are obstacles to the application of this technique to real-time environments. In this paper, we propose the causal affect prediction network (CAPNet), which uses only past facial images to predict corresponding affective valence and arousal. We train CAPNet to learn causal inference between past images and corresponding affective valence and arousal through supervised learning by pairing the sequence of past images with the current label using the Aff-Wild2 dataset. We show through experiments that the well-trained CAPNet outperforms the baseline of the second challenge of the Affective Behavior Analysis in-the-wild (ABAW2) Competition by predicting affective valence and arousal only with past facial images one-third of a second earlier. Therefore, in real-time application, CAPNet can reliably predict affective valence and arousal only with past data.

cs.CV

Observation of broken inversion and chiral symmetries in the pseudogap phase in single and double layer bismuth-based cuprates

We deduce the symmetry of the pseudogap state in the single and double layer bismuth-based cuprate superconductors by measuring and analyzing their circular and linear photogalvanic responses, which are related linearly to the chirality and inversion breaking respectively of the order parameter. After separating out the trivial contribution arising from the surface where inversion symmetry is already broken, we show that both responses start below the pseudogap temperature $T^*$ and grow below it to a sizable magnitude, revealing the broken symmetries in the bulk of the crystal. Through a detailed analysis of the dependence of the signals on the angle of incidence, the polarization of the light, and the orientation of the crystal, we are able to discover that the point group symmetry below $T^*$ is limited to $mm2$ or $mm2\underline{1}$ groups. Taking into account formation of domains and previous measurements, our results narrow down the possible symmetries of the microscopic origin of the phase transition(s) at $T^*$.

cond-mat.str-el

Temperature-induced inversion of the spin-photogalvanic effect in WTe$_2$ and MoTe$_2$

We investigate the generation and temperature-induced evolution of optically-driven spin photocurrents in WTe$_2$ and MoTe$_2$. By correlating the scattering-plane dependence of the spin photocurrents with the symmetry analysis, we find that a sizeable spin photocurrent can be controllably driven along the chain direction by optically exciting the system in the high-symmetry $y$-$z$ plane. Temperature dependence measurements show that pronounced variations in the spin photocurrent emerge at temperatures that coincide with the onset of anomalies in their transport and optical properties. The decreasing trend in the temperature dependence starting below 150 K is attributed to the temperature-induced Lifshitz transition. The sign inversion of the spin photocurrent, observed around 50 K in WTe$_2$ and around 120 K in MoTe$_2$, may have its origin in an interaction that involves multiple kinds of carriers.

cond-mat.mtrl-sci

Matching rules from Al-Co potentials in an almost realistic model

We consider a model decagonal quasicrystal of composition Al$_{80.1}$Co$_{19.9}$ -- closely related to actual structures, and using realistic pair potentials -- on a quasilattice of candidate sites. Its ground state, according to simulations, is a Hexagon-Boat-Star tiling satisfying Penrose's matching rules. In this note, we rationalize these results in terms of the potentials; the Al-Co second-neighbor potential well is crucial.

cond-mat.mtrl-sci

Penrose Matching Rules from Realistic Potentials in a Model System

We exhibit a toy model of a binary decagonal Al-Co quasicrystal -- closely related to actual structures -- in which realistic pair potentials yield a ground state which appears to perfectly implement Penrose's matching rules, for Hexagon-Boat-Star (HBS) tiles of edge 2.45 A. The second minimum of the potentials is crucial for this result.

cond-mat.mtrl-sci