arXiv ScienceSearch

arXiv subjects

Wonjun Kim

Publications and source records attributed to Wonjun Kim.

8 recordsLinked to original sources

EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder

We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-mounted devices or smart glasses. The challenge of this task lies in estimating the pose information of unobserved body parts based solely on a single joint (i.e., head) trajectory. Several studies have begun to adopt head-conditioned generative models, however, such previous methods are costly and time-consuming due to the diffusion-based iterative process. As an alternative, we propose a simple yet novel method that leverages the latent space of the guidance network, which is designed as a variational autoencoder taking full-body poses as inputs. By enforcing latent distributions of this guidance network and our head-to-motion network to be similar, latent features sampled from the 'guided' distribution, i.e., distribution learned in our head-to-motion network, can be reliably decoded for natural representations of full-body poses even only with the head pose. One important advantage of the proposed method is that one-step sampling scheme achieves remarkably fast inference (more than 50 times faster) compared to diffusion-based approaches. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of ego-body mesh reconstruction.

cs.CV

Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing methods rely on explicit geometric constraints or predefined structural representations to accurately position 3D Gaussians. Although these approaches have shown the remarkable progress in this field, they still suffer from inconsistent feature representations across multi-view inputs due to complex articulations of the human body and limited overlaps between different views. To address this problem, we propose a novel method to accurately localize 3D Gaussians and ultimately improve the quality of human rendering. The key idea is to unproject latent embeddings encoded from each viewpoint into a shared 3D space through predicted depth maps and recalibrate them belonging to the same body part based on cross-view attention. This helps the model resolve the spatial ambiguity occurring in highly textured regions as well as occluded body parts, thus leading to the accurate localization of 3D Gaussians. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of generalizable human Gaussian splatting from sparse-view inputs.

cs.CV

DropGaussian: Structural Regularization for Sparse-view Gaussian Splatting

Recently, 3D Gaussian splatting (3DGS) has gained considerable attentions in the field of novel view synthesis due to its fast performance while yielding the excellent image quality. However, 3DGS in sparse-view settings (e.g., three-view inputs) often faces with the problem of overfitting to training views, which significantly drops the visual quality of novel view images. Many existing approaches have tackled this issue by using strong priors, such as 2D generative contextual information and external depth signals. In contrast, this paper introduces a prior-free method, so-called DropGaussian, with simple changes in 3D Gaussian splatting. Specifically, we randomly remove Gaussians during the training process in a similar way of dropout, which allows non-excluded Gaussians to have larger gradients while improving their visibility. This makes the remaining Gaussians to contribute more to the optimization process for rendering with sparse input views. Such simple operation effectively alleviates the overfitting problem and enhances the quality of novel view synthesis. By simply applying DropGaussian to the original 3DGS framework, we can achieve the competitive performance with existing prior-based 3DGS methods in sparse-view settings of benchmark datasets without any additional complexity. The code and model are publicly available at: https://github.com/DCVL-3D/DropGaussian release.

cs.CV

Sampling is Matter: Point-guided 3D Human Mesh Reconstruction

This paper presents a simple yet powerful method for 3D human mesh reconstruction from a single RGB image. Most recently, the non-local interactions of the whole mesh vertices have been effectively estimated in the transformer while the relationship between body parts also has begun to be handled via the graph model. Even though those approaches have shown the remarkable progress in 3D human mesh reconstruction, it is still difficult to directly infer the relationship between features, which are encoded from the 2D input image, and 3D coordinates of each vertex. To resolve this problem, we propose to design a simple feature sampling scheme. The key idea is to sample features in the embedded space by following the guide of points, which are estimated as projection results of 3D mesh vertices (i.e., ground truth). This helps the model to concentrate more on vertex-relevant features in the 2D space, thus leading to the reconstruction of the natural human pose. Furthermore, we apply progressive attention masking to precisely estimate local interactions between vertices even under severe occlusions. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of 3D human mesh reconstruction. The code and model are publicly available at: https://github.com/DCVL-3D/PointHMR_release.

cs.CV

Towards Deep Learning-aided Wireless Channel Estimation and Channel State Information Feedback for 6G

Deep learning (DL), a branch of artificial intelligence (AI) techniques, has shown great promise in various disciplines such as image classification and segmentation, speech recognition, language translation, among others. This remarkable success of DL has stimulated increasing interest in applying this paradigm to wireless channel estimation in recent years. Since DL principles are inductive in nature and distinct from the conventional rule-based algorithms, when one tries to use DL technique to the channel estimation, one might easily get stuck and confused by so many knobs to control and small details to be aware of. The primary purpose of this paper is to discuss key issues and possible solutions in DL-based wireless channel estimation and channel state information (CSI) feedback including the DL model selection, training data acquisition, and neural network design for 6G. Specifically, we present several case studies together with the numerical experiments to demonstrate the effectiveness of the DL-based wireless channel estimation framework.

eess.SP

Sparse Vector Transmission: An Idea Whose Time Has Come

In recent years, we are witnessing bewildering variety of automated services and applications of vehicles, robots, sensors, and machines powered by the artificial intelligence technologies. Communication mechanism associated with these services is dearly distinct from human-centric communications. One important feature for the machine-centric communications is that the amount of information to be transmitted is tiny. In view of the short packet transmission, relying on today's transmission mechanism would not be efficient due to the waste of resources, large decoding latency, and expensive operational cost. In this article, we present an overview of the sparse vector transmission (SVT), a scheme to transmit a short-sized information after the sparse transformation. We discuss basics of SVT, two distinct SVT strategies, viz., frequency-domain sparse transmission and sparse vector coding with detailed operations, and also demonstrate the effectiveness in realistic wireless environments.

eess.SP

Deep Neural Network Based Active User Detection for Grant-free NOMA Systems

As a means to support the access of massive machine-type communication devices, grant-free access and non-orthogonal multiple access (NOMA) have received great deal of attention in recent years. In the grant-free transmission, each device transmits information without the granting process so that the basestation needs to identify the active devices among all potential devices. This process, called an active user detection (AUD), is a challenging problem in the NOMA-based systems since it is difficult to identify active devices from the superimposed received signal. An aim of this paper is to put forth a new type of AUD based on deep neural network (DNN). By applying the training data in the properly designed DNN, the proposed AUD scheme learns the nonlinear mapping between the received NOMA signal and indices of active devices. As a result, the trained DNN can handle the whole AUD process, achieving an accurate detection of the active users. Numerical results demonstrate that the proposed AUD scheme outperforms the conventional approaches in both AUD success probability and computational complexity.

eess.SP

Channel aware sparse transmission for ultra low-latency communications in TDD systems

Major goal of ultra reliable and low latency communication (URLLC) is to reduce the latency down to a millisecond (ms) level while ensuring reliability of the transmission. Since the current uplink transmission scheme requires a complicated handshaking procedure to initiate the transmission, to meet this stringent latency requirement is a challenge in wireless system design. In particular, in the time division duplexing (TDD) systems, supporting the URLLC is difficult since the mobile device has to wait until the transmit direction is switched to the uplink. In this paper, we propose a new approach to support a low latency access in TDD systems, called channel aware sparse transmission (CAST). Key idea of the proposed scheme is to encode a grant signal in a form of sparse vector. This together with the fact that the sensing mechanism preserves the energy of the sparse vector allows us to use the compressed sensing (CS) technique in CAST decoding. From the performance analysis and numerical evaluations, we demonstrate that the proposed CAST scheme achieves a significant reduction in access latency over the 4G LTE-TDD and 5G NR-TDD systems.

eess.SP