arXiv ScienceSearch

arXiv subjects

Sang Jun Lee

Publications and source records attributed to Sang Jun Lee.

13 recordsLinked to original sources

MFVINS: Multiple Fisheye Camera-Based Visual Inertial System

A simultaneous localization and mapping (SLAM) method using a monocular camera and a low-cost inertial measurement unit (IMU) sensor is an effective way to fulfill a low-cost sensor configuration. Using this sensor configuration, visual-inertial system (VINS) focuses on fusing data from a camera and an IMU sensor to estimate the six degrees-of-freedom (DOF) of the sensor pose. Typically, VINS uses only a single camera as visual input, which lead to problems such as error accumulation due to occlusion, various illumination, and textureless environments. In this paper, we propose a new multiple fisheye camera-based visual-inertial system called MFVINS. We present an IMU-aided FAST feature tracker for multiple cameras that enables efficient extraction and robust matching of local features. Then, the proposed method filters out outliers caused by fisheye distortion on the normalized image plane. Subsequently, a new reprojection error with physical validity constraints is proposed for bundle adjustment using learning-based depth estimation. The proposed method is applied to various scenarios, and its effectiveness is demonstrated by comparing previous VINS methods. In particular, MFVINS is implemented in real-time process to leverage the advantages of using multiple cameras -- robustness against occlusion and textureless regions -- while reducing the computational burden.

cs.RO

M2P-AD: Memory-to-Prototype Learning with Boundary-aware Score Refinement for 3D Anomaly Detection

3D anomaly detection has recently emerged as an important research topic in computer vision. Although existing methods have achieved high performance, excessive anomaly responses in normal regions and false positives near object boundaries remain unresolved challenges. To address these challenges, we propose a novel 3D anomaly detection model, Memory-to-Prototype Anomaly Detection (M2P-AD), which effectively models the distribution of normal features while suppressing excessive anomaly scores in normal regions and false positives near object boundaries. Specifically, we introduce a Memory-to-Prototype (M2P) module that learns representative prototypes from normal feature embeddings to preserve important structural information of objects. In addition, a Boundary extraction (BE) module is integrated to identify object boundaries, and a Boundary-aware score refinement (BSR) strategy is applied to recalibrate anomaly scores by incorporating boundary characteristics. The proposed method is evaluated on Real3D-AD, Anomaly-ShapeNet, and MulSen-AD, achieving state-of-the-art performance. Qualitative results demonstrate that excessive anomaly scores in normal regions are reduced and false positives near object boundaries are suppressed, resulting in more accurate and stable anomaly localization. The results indicate that the proposed approach enables more reliable 3D anomaly detection and provides a robust solution applicable to real-world industrial environments.

cs.CV

MS-rPPG: Multi-spectral State Space Model for Remote Photoplethysmography in Driver Monitoring Systems

Remote photoplethysmography (rPPG) is a camera-based technique for measuring physiological signals, particularly cardiac activity. From the remotely measured signals, heart rate can be estimated, which is crucial for health monitoring. In this study, we investigate a driver health monitoring system based on remote heart rate estimation. However, driving environments represent uncontrolled settings where videos are subject to varying illumination conditions and frequent head movements. We introduce MS-rPPG, a multi-spectral framework that combines RGB with near-infrared (NIR) face video to alleviate rPPG estimation under challenging driving conditions. To combine the complementary features from two spectral videos, we propose a cross-spectral linear modulation (CSLM) strategy based on frequency-domain analysis. Moreover, we introduce MS-Mamba, a novel state space model designed to effectively model long-range temporal dependencies while jointly capturing cross-channel interactions between multi-spectral features. We collected a real-world dataset called MS-Drive, which was recorded from 50 participants while driving the vehicle. The proposed method was evaluated on the MR-NIRP Car dataset and MS-Drive datasets. The experimental results indicate that MS-rPPG shows better robustness and heart rate estimation accuracy than previous methods, highlighting its promise for driver health monitoring. The codes are available at github.com/ziiho08/MS-rPPG.

cs.CV

Instance-Aware Knowledge Distillation for Semi-Supervised Learning of an On-Board Multi-Task Dense Prediction Model for Collision Avoidance System

Collision avoidance systems have evolved toward camera-based deep learning approaches for driving scene understanding. However, deployment in edge environments such as country clubs is constrained by limited computational resources and unreliable communication infrastructure. Moreover, constructing large-scale datasets for the target domain involves substantial annotation cost. To address these limitations, we propose an instance-aware knowledge distillation framework for semi-supervised learning. Specifically, we generate pseudo labels that mitigate teacher bias by leveraging domain priors from the teacher and instance-centric knowledge from foundation models. The trained lightweight student is deployed in the proposed collision avoidance system and performs multiple dense prediction tasks in real-time. The system detects frontal obstacles and encodes their spatial information into controller area network messages for automated guided vehicle operation. To achieve this, we construct a large-scale country club dataset and perform field validation of the proposed system. Experimental results demonstrate that the student outperforms the large teacher in instance segmentation while mitigating performance degradation in monocular depth estimation. Compared with the teacher, the student reduces FLOPs by 22.68$\times$ and parameters by 14.33$\times$, achieving 6.46 FPS on a low-cost edge device.

cs.CV

Beer-Lambert Guided Representation Learning for Unsupervised Anomaly Detection in Sub-THz Food Inspection Images

Food manufacturing requires reliable inspection systems to detect foreign material contamination and maintain product safety. Sub-THz transmission imaging provides material-dependent attenuation characteristics that are useful for detecting low-density contaminants in food products. However, existing unsupervised anomaly detection methods mainly rely on RGB-pretrained visual representations, which may not adequately capture the transmission behavior of Sub-THz images. This paper proposes a Beer-Lambert guided representation learning framework for unsupervised anomaly detection in Sub-THz food inspection images. The proposed method introduces an attenuation decomposition module as an auxiliary regularization module that constrains student representations through attenuation reconstruction during training. In addition to the conventional one-class setting, we introduce a Leave-One-Food-Out protocol to evaluate generalization capability under unseen food categories. Experimental results on the Inline-Food-Inspection-THz dataset show that the proposed method improves overall anomaly detection performance over the baseline method.

cs.CV

Weighted Knowledge Distillation for Semi-Supervised Segmentation of Maxillary Sinus in Panoramic X-ray Images

Accurate segmentation of maxillary sinus in panoramic X-ray images is essential for dental diagnosis and surgical planning; however, this task remains relatively underexplored in dental imaging research. Structural overlap, ambiguous anatomical boundaries inherent to two-dimensional panoramic projections, and the limited availability of large scale clinical datasets with reliable pixel-level annotations make the development and evaluation of segmentation models challenging. To address these challenges, we propose a semi-supervised segmentation framework that effectively leverages both labeled and unlabeled panoramic radiographs, where knowledge distillation is utilized to train a student model with reliable structural information distilled from a teacher model. Specifically, we introduce a weighted knowledge distillation loss to suppress unreliable distillation signals caused by structural discrepancies between teacher and student predictions. To further enhance the quality of pseudo labels generated by the teacher network, we introduce SinusCycle-GAN which is a refinement network based on unpaired image-to-image translation. This refinement process improves the precision of boundaries and reduces noise propagation when learning from unlabeled data during semi-supervised training. To evaluate the proposed method, we collected clinical panoramic X-ray images from 2,511 patients, and experimental results demonstrate that the proposed method outperforms state-of-the-art segmentation models, achieving the Dice score of 96.35\% while reducing boundary error. The results indicate that the proposed semi-supervised framework provides robust and anatomically consistent segmentation performance under limited labeled data conditions, highlighting its potential for broader dental image analysis applications.

cs.CV

Designing heterostructures to control oxygen stoichiometry in helimagnetic perovskite strontium ferrite

A large challenge in determining the physics of helimagnetic SrFeO3 is in stabilizing the stoichiometric chemical phase over long enough time scales to conduct extensive measurements. Degradation in SrFeO3 manifests mainly as a crossover from metallic to insulating behavior. Using a combination of electronic transport and density functional theory, we show that this degradation is dominated by oxygen loss, possibly on the order of one percent. We further demonstrate that high quality SrFeO3 thin films can be stabilized long-term by combining a nanoscale band insulator capping layer with an ex situ ozone anneal. We show that this produces a nearly-pristine cation sublattice and preserves metallicity for at least several weeks. These results establish a reliable pathway for producing chemically stable SrFeO3 thin films, enabling reproducible studies of its unusual helimagnetism.

cond-mat.mtrl-sci

Periodic-MAE: Periodic Video Masked Autoencoder for rPPG Estimation

In this paper, we propose Periodic-MAE, a self-supervised framework for learning generalizable spatio-temporal representations of periodic physiological signals from unlabeled facial videos. The proposed method leverages a masked autoencoder (MAE), which learns high-dimensional facial representations by reconstructing masked video tokens without relying on remote photoplethysmography (rPPG) specific supervision. To explicitly align representation learning with the characteristics of rPPG, we introduce a periodicity-aware frame masking strategy based on video resampling, enabling the encoder to learn representations that capture quasi-periodic temporal patterns relevant to pulse signal estimation. In addition, physiological bandlimit constraints are integrated into the MAE pre-training framework, exploiting the sparsity of pulse signals in the frequency domain to guide the learned representations toward physiologically meaningful patterns. After pre-training, the learned representations are transferred to downstream rPPG estimation, where the encoder serves as a generic feature extractor for recovering pulse-related signals from facial videos. We conduct extensive experiments on four benchmark datasets, including PURE, UBFC-rPPG, MMPD, and V4V. Moreover, we evaluate the proposed approach on a real-world rPPG dataset collected under unconstrained lighting conditions and subject motion. Experimental results demonstrate that Periodic-MAE consistently improves rPPG estimation performance, particularly in challenging cross-dataset and real-world evaluation settings. Our code is available at https://github.com/ziiho08/Periodic-MAE.

cs.CV

Phase-shifted remote photoplethysmography for estimating heart rate and blood pressure from facial video

Human health can be critically affected by cardiovascular diseases, such as hypertension, arrhythmias, and stroke. Heart rate and blood pressure are important biometric information for the monitoring of cardiovascular system and early diagnosis of cardiovascular diseases. Existing methods for estimating the heart rate are based on electrocardiography and photoplethyomography, which require contacting the sensor to the skin surface. Moreover, catheter and cuff-based methods for measuring blood pressure cause inconvenience and have limited applicability. Therefore, in this thesis, we propose a vision-based method for estimating the heart rate and blood pressure. This thesis proposes a 2-stage deep learning framework consisting of a dual remote photoplethysmography network (DRP-Net) and bounded blood pressure network (BBP-Net). In the first stage, DRP-Net infers remote photoplethysmography (rPPG) signals for the acral and facial regions, and these phase-shifted rPPG signals are utilized to estimate the heart rate. In the second stage, BBP-Net integrates temporal features and analyzes phase discrepancy between the acral and facial rPPG signals to estimate SBP and DBP values. To improve the accuracy of estimating the heart rate, we employed a data augmentation method based on a frame interpolation model. Moreover, we designed BBP-Net to infer blood pressure within a predefined range by incorporating a scaled sigmoid function. Our method resulted in estimating the heart rate with the mean absolute error (MAE) of 1.78 BPM, reducing the MAE by 34.31 % compared to the recent method, on the MMSE-HR dataset. The MAE for estimating the systolic blood pressure (SBP) and diastolic blood pressure (DBP) were 10.19 mmHg and 7.09 mmHg. On the V4V dataset, the MAE for the heart rate, SBP, and DBP were 3.83 BPM, 13.64 mmHg, and 9.4 mmHg, respectively.

cs.CV

Tuning Excited State Electron Transfer in Fe Tetracyano-Polypyridyl Complexes

We have investigated photoinduced intramolecular electron transfer dynamics following metal-to-ligand charge-transfer (MLCT) excitation of [Fe(CN)$_4$(2,2'-bipyridine)]$^{2-}$ (1), [Fe(CN)$_4$(2,3-bis(2-pyridyl)pyrazine)]$^{2-}$ (2) and [Fe(CN)$_4$(2,2'-bipyrimidine)]$^{2-}$ (3) complexes in various solvents with static and time-resolved UV-visible absorption spectroscopy and Fe 2p3d resonant inelastic X-ray scattering. We observe $^3$MLCT lifetimes from 180 fs to 67 ps over a wide range of MLCT energies in different solvents by utilizing the strong solvatochromism of the complexes. Intramolecular electron transfer lifetimes governing $^3$MLCT relaxation increase monotonically and (super)exponentially as the $^3$MLCT energy is decreased in 1 and 2 by changing the solvent. This behavior can be described with non-adiabatic classical Marcus electron transfer dynamics along the indirect $^3$MLCT->$^3$MC pathway, where the $^3$MC is the lowest energy metal-centered (MC) excited state. In contrast, the $^3$MLCT lifetime in 3 changes non-monotonically and exhibits a maximum. This qualitatively different behaviour results from direct electron transfer from the $^3$MLCT to the electronic ground state (GS). This pathway involves nuclear tunnelling for the high-frequency polypyridyl skeleton mode ($\hbar\omega$ = 1530 cm$^{-1}$), which is more displaced for 3 than for either 1 or 2, therefore making the direct pathway significantly more efficient in 3. To our knowledge, this is the first observation of an efficient $^3$MLCT->GS relaxation pathway in an Fe polypyridyl complex. Our study suggests that further extending the MLCT state lifetime requires (1) lowering the $^3$MLCT state energy with respect to the $^3$MC state and (2) suppressing the intramolecular distortion of the electron-accepting ligand in the $^3$MLCT excited state to suppress the rate of direct $^3$MLCT->GS electron transfer.

physics.chem-ph

Selective Distillation of Weakly Annotated GTD for Vision-based Slab Identification System

This paper proposes an algorithm for recognizing slab identification numbers in factory scenes. In the development of a deep-learning based system, manual labeling to make ground truth data (GTD) is an important but expensive task. Furthermore, the quality of GTD is closely related to the performance of a supervised learning algorithm. To reduce manual work in the labeling process, we generated weakly annotated GTD by marking only character centroids. Whereas bounding-boxes for characters require at least a drag-and-drop operation or two clicks to annotate a character location, the weakly annotated GTD requires a single click to record a character location. The main contribution of this paper is on selective distillation to improve the quality of the weakly annotated GTD. Because manual GTD are usually generated by many people, it may contain personal bias or human error. To address this problem, the information in manual GTD is integrated and refined by selective distillation. In the process of selective distillation, a fully convolutional network is trained using the weakly annotated GTD, and its prediction maps are selectively used to revise locations and boundaries of semantic regions of characters in the initial GTD. The modified GTD are used in the main training stage, and a post-processing is conducted to retrieve text information. Experiments were thoroughly conducted on actual industry data collected at a steelmaking factory to demonstrate the effectiveness of the proposed method.

cs.CV

L-Edge Spectroscopy of Dilute, Radiation-Sensitive Systems Using a Transition-Edge-Sensor Array

We present X-ray absorption spectroscopy and resonant inelastic X-ray scattering (RIXS) measurements on the iron L-edge of 0.5 mM aqueous ferricyanide. These measurements demonstrate the ability of high-throughput transition-edge-sensor (TES) spectrometers to access the rich soft X-ray (100-2000eV) spectroscopy regime for dilute and radiation-sensitive samples. Our low-concentration data are in agreement with high-concentration measurements recorded by conventional grating-based spectrometers. These results show that soft X-ray RIXS spectroscopy acquired by high-throughput TES spectrometers can be used to study the local electronic structure of dilute metal-centered complexes relevant to biology, chemistry and catalysis. In particular, TES spectrometers have a unique ability to characterize frozen solutions of radiation- and temperature-sensitive samples.

physics.ins-det

Metamaterial Perfect Absorber Analyzed by a Meta-cavity Model Consisting of Multilayer Metasurfaces

We demonstrate that the metamaterial perfect absorber behaves as a meta-cavity bounded between a resonant metasurface and a metallic thin-film reflector. The perfect absorption is achieved by the Fabry-Perot cavity resonance via multiple reflections between the "quasi-open" boundary of resonator and the "closed" boundary of reflector. The characteristic features including angle independence, ultra-thin thickness and strong field localization can be well explained by this model. With this model, metamaterial perfect absorber can be redefined as a meta-cavity exhibiting high Q-factor, strong field enhancement and extremely high photonic density of states, thereby promising novel applications for high performance sensor, infrared photodetector and cavity quantum electrodynamics devices.

physics.optics