arXiv ScienceSearch

arXiv subjects

Inki Kim

Publications and source records attributed to Inki Kim.

16 recordsLinked to original sources

Gaze Target Estimation Anywhere with Concepts

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.

cs.CV

Universal zero-crosstalk photonic integration via slab-engineered mode hybridization

Photonic integrated circuits have emerged as a scalable platform for optical computing, communication, and quantum technologies, where high-fidelity optical processing is essential. However, as photonic systems scale in complexity, inter-channel crosstalk accumulates across cascaded components, fundamentally degrading signal fidelity, limiting system-level performance, and constraining integration density. Existing crosstalk-suppression strategies rely on specialized nanostructures or platform-specific designs, hindering their adoption in standard foundry processes and across diverse material systems. Here we establish a universal and foundry-compatible route to eliminating crosstalk based on slab-engineered mode hybridization in standard rib waveguides. By tailoring the slab thickness, mode hybridization induces anisotropic modal perturbations that enable complete cancellation of coupling between adjacent waveguides. We experimentally demonstrate zero-crosstalk across diverse material platforms, including silicon-on-insulator, silicon nitride, thin-film lithium niobate, and germanium-on-insulator, spanning wavelengths from the visible to the mid-infrared. Our approach provides a manufacturable route toward scalable, high-fidelity, and high-density photonic integration, overcoming the long-standing trade-off between signal fidelity and integration density in large-scale photonic systems.

physics.optics

Design of Outage-Limit-Approaching Protograph LDPC Codes via Generalized Rootchecks

This paper presents a new protograph-based LDPC code design framework that simultaneously achieves full diversity over block-fading channels (BFCs) and near-capacity performance over additive white Gaussian noise channels. By leveraging a Boolean approximation-based analysis-Diversity Evolution-we derive structural constraints with generalized rootchecks that guarantee full diversity. Building on these constraints, we propose a diversity-aligned protograph template tailored for the two-block BFC (M=2) that ensures full diversity under iterative belief propagation decoding. Furthermore, a genetic algorithm guided by density evolution is employed to optimize the protograph edges within this family for improved coding gain. The resulting codes, termed DA-GRP-LDPC codes, simultaneously achieve full diversity and enhanced coding gain, reaching a 0.8 dB gap to the outage limit for the two-block BFC at a block length of 16,896. This demonstrates that the proposed framework effectively bridges the gap between diversity optimality in non-ergodic channels and high coding gain in ergodic channels.

cs.IT

5G LDPC Codes as Root LDPC Codes via Diversity Alignment

This paper studies the diversity of protographbased quasi-cyclic low-density parity-check (QC-LDPC) codes over nonergodic block-fading channels under iterative beliefpropagation decoding. We introduce diversity evolution (DivE), a Boolean-function-based analysis method that tracks how the fading dependence of belief-propagation messages evolves across decoding iterations. Under a Boolean approximation of block fading, DivE derives a Boolean fading function for each variable node (VN) output (i.e., the a-posteriori reliability after iterative decoding), from which the VN diversity order can be directly determined. Building on this insight, we develop a greedy blockmapping search that assigns protograph VNs to fading blocks so that all information VNs achieve full diversity, while including the minimum additional parity VNs when full diversity is infeasible at the nominal rate. Numerical results on the 5G New Radio LDPC codes show that the proposed search finds block mappings that guarantee full diversity for all information bits without modifying the base-graph structure, yielding a markedly steeper high-SNR slope and lower BLER than random mappings.

cs.IT

Estimation of Tissue Deformation and Interactive Force in Robotic Surgery through Vision-based Learning

Goal: A limitation in robotic surgery is the lack of force feedback, due to challenges in suitable sensing techniques. To enhance the perception of the surgeons and precise force rendering, estimation of these forces along with tissue deformation level is presented here. Methods: An experimental test bed is built for studying the interaction, and the forces are estimated from the raw data. Since tissue deformation and stiffness are non-linearly related, they are independently computed for enhanced reliability. A Convolutional Neural Network (CNN) based vision model is deployed, and both classification and regression models are developed. Results: The forces applied on the tissue are estimated, and the tissue is classified based on its deformation. The exact deformation of the tissue is also computed. Conclusions: The surgeons can render precise forces and detect tumors using the proposed method. The rarely discussed efficacy of computing the deformation level is also demonstrated.

eess.SY

Aberration Correcting Vision Transformers for High-Fidelity Metalens Imaging

Metalens is an emerging optical system with an irreplaceable merit in that it can be manufactured in ultra-thin and compact sizes, which shows great promise in various applications. Despite its advantage in miniaturization, its practicality is constrained by spatially varying aberrations and distortions, which significantly degrade the image quality. Several previous arts have attempted to address different types of aberrations, yet most of them are mainly designed for the traditional bulky lens and ineffective to remedy harsh aberrations of the metalens. While there have existed aberration correction methods specifically for metalens, they still fall short of restoration quality. In this work, we propose a novel aberration correction framework for metalens-captured images, harnessing Vision Transformers (ViT) that have the potential to restore metalens images with non-uniform aberrations. Specifically, we devise a Multiple Adaptive Filters Guidance (MAFG), where multiple Wiener filters enrich the degraded input images with various noise-detail balances and a cross-attention module reweights the features considering the different degrees of aberrations. In addition, we introduce a Spatial and Transposed self-Attention Fusion (STAF) module, which aggregates features from spatial self-attention and transposed self-attention modules to further ameliorate aberration correction. We conduct extensive experiments, including correcting aberrated images and videos, and clean 3D reconstruction. The proposed method outperforms the previous arts by a significant margin. We further fabricate a metalens and verify the practicality of our method by restoring the images captured with the manufactured metalens. Code and pre-trained models are available at https://benhenryl.github.io/Metalens-Transformer.

eess.IV

Room-temperature waveguide-integrated photodetector using bolometric effect for mid-infrared spectroscopy applications

Waveguide-integrated mid-infrared (MIR) photodetectors are pivotal components for the development of molecular spectroscopy applications, leveraging mature photonic integrated circuit (PIC) technologies. Despite various strategies, critical challenges still remain in achieving broadband photoresponse, cooling-free operation, and large-scale complementary-metal-oxide-semiconductor (CMOS)-compatible manufacturability. To leap beyond these limitations, the bolometric effect - a thermal detection mechanism - is introduced into the waveguide platform. More importantly, we pursue a free-carrier absorption (FCA) process in germanium (Ge) to create an efficient light-absorbing medium, providing a pragmatic solution for full coverage of the MIR spectrum without incorporating exotic materials into CMOS. Here, we present an uncooled waveguide-integrated photodetector based on a Ge-on-insulator (Ge-OI) PIC architecture, which exploits the bolometric effect combined with FCA. Notably, our device exhibits a broadband responsivity of 28.35 %/mW across 4030-4360 nm (and potentially beyond), challenging the state of the art, while achieving a noise-equivalent power of $4.03$x$10^{-7} W/Hz^{0.5}$ at 4180 nm. We further demonstrate label-free sensing of gaseous carbon dioxide (CO2) using our integrated photodetector and sensing waveguide on a single chip. This approach to room-temperature waveguide-integrated MIR photodetection, harnessing bolometry with FCA in Ge, not only facilitates the realization of fully integrated lab-on-a-chip systems with wavelength flexibility but also provides a blueprint for MIR PICs with CMOS-foundry-compatibility.

physics.optics

Radiative control of localized excitons at room temperature with an ultracompact tip-enhanced plasmonic nano-cavity

In atomically thin semiconductors, localized exciton (X$_L$) coupled to light shows single quantum emitting behaviors through radiative relaxation processes providing a new class of optical sources for potential applications in quantum communication. In most studies, however, X$_L$ photoluminescence (PL) from crystal defects has mainly been observed in cryogenic conditions because of their sub-wavelength emission region and low quantum yield at room temperature. Furthermore, engineering the radiative relaxation properties, e.g., emission region, intensity, and energy, remained challenging. Here, we present a plasmonic antenna with a triple-sharp-tips geometry to induce and control the X$_L$ emission of a WSe$_2$ monolayer (ML) at room temperature. By placing a ML crystal on the two sharp Au tips in a bowtie antenna fabricated through cascade domino lithography with a radius of curvature of <1 nm, we effectively induce tensile strain in the nanoscale region to create robust X$_L$ states. An Au tip with tip-enhanced photoluminescence (TEPL) spectroscopy is then added to the strained region to probe and control the X$_L$ emission. With TEPL enhancement of X$_L$ as high as ~10$^6$ in the triple-sharp-tips device, experimental results demonstrate the controllable X$_L$ emission in <30 nm area with a PL energy shift up to 40 meV, resolved by tip-enhanced PL and Raman imaging with <15 nm spatial resolution. Our approach provides a systematic way to control localized quantum light in 2D semiconductors offering new strategies for active quantum nano-optical devices.

physics.optics

MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement Learning

Inspired by human visual attention, we propose a novel inverse reinforcement learning formulation using Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) for predicting the visual attention of drivers in accident-prone situations. MEDIRL predicts fixation locations that lead to maximal rewards by learning a task-sensitive reward function from eye fixation patterns recorded from attentive drivers. Additionally, we introduce EyeCar, a new driver attention dataset in accident-prone situations. We conduct comprehensive experiments to evaluate our proposed model on three common benchmarks: (DR(eye)VE, BDD-A, DADA-2000), and our EyeCar dataset. Results indicate that MEDIRL outperforms existing models for predicting attention and achieves state-of-the-art performance. We present extensive ablation studies to provide more insights into different features of our proposed model.

cs.CV

Biomimetic Ultra-Broadband Perfect Absorbers Optimised with Reinforcement Learning

By learning the optimal policy with a double deep Q-learning network, we design ultra-broadband, biomimetic, perfect absorbers with various materials, based the structure of a moths eye. All absorbers achieve over 90% average absorption from 400 to 1,600 nm. By training a DDQN with motheye structures made up of chromium, we transfer the learned knowledge to other, similar materials to quickly and efficiently find the optimal parameters from the around 1 billion possible options. The knowledge learned from previous optimisations helps the network to find the best solution for a new material in fewer steps, dramatically increasing the efficiency of finding designs with ultra-broadband absorption.

physics.optics

A Case Study of Trust on Autonomous Driving

As autonomous vehicles have benefited the society, understanding the dynamic change of humans' trust during human-autonomous vehicle interaction can help to improve the safety and performance of autonomous driving. We designed and conducted a human subjects study involving 19 participants. Each participant was asked to enter their trust level in a Likert scale in real-time during experiments on a driving simulator. We also collected physiological data (e.g., heart rate, pupil size) of participants as complementary indicators of trust. We used analysis of variance (ANOVA) and Signal Temporal Logic (STL) to analyze the experimental data. Our results show the influence of different factors (e.g., automation alarms, weather conditions) on trust, and the individual variability in human reaction time and trust change.

cs.HC

Exploring Gaze Behavior to Assess Performance in Digital Game-Based Learning Systems

The recent growth of sophisticated digital gaming technologies has spawned an \$8.1B industry around using these games for pedagogical purposes. Though Digital Game-Based Learning Systems have been adopted by industries ranging from military to medical applications, these systems continue to rely on traditional measures of explicit interactions to gauge player performance which can be subject to guessing and other factors unrelated to actual performance. This study presents a novel implicit eye-tracking based metric for digital game-based learning environments. The proposed metric introduces a weighted eye-tracking measure of traditional in-game scoring to consider the mental schema of a player's decision making. In order to validate the efficacy of this metric, we conducted an experiment with 25 participants playing a game designed to evaluate Chinese cultural competency and communication. This experiment showed strong correlation between the novel eye-tracking performance metric and traditional measures of in-game performance.

cs.HC

The Effect of Whole-Body Haptic Feedback on Driver's Perception in Negotiating a Curve

It remains uncertain regarding the safety of driving in autonomous vehicles that, after a long, passive control and inattention to the driving situation, how the drivers will be effectively informed to take-over the control in emergency. In particular, the active role of vehicle force feedback on the driver's risk perception on curves has not been fully explored. To investigate it, the current paper examined the driver's cognitive and visual responses to the whole-body haptic feedback during curve negotiations. The effects of force feedback on drivers' responses on curves were investigated in a high-fidelity driving simulator while measuring EEG and visual gaze over ten participants. The preliminary analyses of the first two participants revealed that pupil diameter and fixation time on the curves were significantly longer when the driver received whole-body feedback, compared to none. The findings suggest that whole-body feedback can be used as an effective "advance notification" of hazards.

cs.HC

Nanophotonic modal dichroism: mode-multiplexed modulators

As the diffraction limit is approached, device miniaturization to integrate more functionality per area becomes more and more challenging. Here we propose a novel strategy to increase the functionality-per-area by exploiting the modal properties of a waveguide system. With such approach the design of a mode-multiplexed nanophotonic modulator relying on the mode-selective absorption of a patterned Indium-Tin-Oxide is proposed. Full-wave simulations of a device operating at the telecom wavelength of 1550nm show that two modes can be independently modulated, while maintaining performances in line with conventional single-mode ITO modulators reported in the recent literature. The proposed design principles can pave the way to a novel class of mode-multiplexed compact photonic devices able to effectively multiply the functionality-per-area in integrated photonic systems.

physics.app-ph

A highly efficient method for second and third harmonic generation from magnetic metamaterials

Second and third harmonic signals have been usually generated by using nonlinear crystals, but that method suffers from the low efficiency in small thicknesses. Metamaterials can be used to generate harmonic signals in small thicknesses. Here, we introduce a new method for amplifying second and third harmonic generation from magnetic metamaterials. We show that by using a grating structure under the metamaterial, the grating and the metamaterial form a resonator, and amplify the resonant behavior of the metamaterial. Therefore, we can generate second and third harmonic signals with high efficiency from this metamaterial-based nonlinear media.

physics.optics

The role of current loop in harmonic generation from magnetic metamaterials in two polarizations

In this paper, we investigate the role of the current loop in the generation of second and third harmonic signals from magnetic metamaterials. We will show that the fact that the current loop in the magnetic resonance acts as a source for nonlinear effects and it consists of two orthogonal parts, leads to the generation of two harmonic signals in two orthogonal polarizations.

physics.optics