arXiv ScienceSearch

arXiv subjects

Taehyoung Kim

Publications and source records attributed to Taehyoung Kim.

9 recordsLinked to original sources

VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

Semantic 3D maps are increasingly constructed automatically for aerial robotics by integrating learned semantic predictions into 3D representations. While this avoids costly manual 3D annotation, errors in the perception and mapping pipeline can persist in the resulting map, reducing its reliability for downstream autonomous tasks. Existing 3D semantic map refinement methods either rely on the original observations, treat occupancy as part of the prediction problem, or apply non-learned local regularization to completed maps. Instead, we study post-hoc semantic correction, asking whether semantic accuracy can be recovered directly from the completed map while keeping its geometry and occupancy fixed. We introduce \method, a graph-based model that corrects voxel labels based on local geometry and neighboring semantic information. To obtain training pairs, we corrupt contiguous regions of annotated OccuFly maps according to class confusions observed in upstream maps. We evaluate \method on completed OccuFly maps generated from predictions of four independently trained 2D segmentation models. \method consistently improves mIoU by 4.23--5.00 percentage points, with gains broadly distributed across the evaluated semantic classes and particularly strong improvements for tree, roof, and wall. Results on an independently reconstructed out-of-distribution aerial scene further suggest that the learned correction can transfer beyond the environments seen during training.

cs.CV

FlexPath: Adapting Learned Connectivity Guidance to Path Preferences

Recent learning-based path planners use neural networks to process occupancy representations and approximate heuristics for classical search algorithms, yielding near-optimal paths with reduced search effort. However, these methods are tied to a fixed objective, usually the shortest-path objective, implicit in their supervision. This limits their flexibility to accommodate alternative criteria. We introduce $\textbf{FlexPath}$, a two-stage learned search-guidance framework that first learns a recall-oriented connectivity prior initialized from shortest-path planner demonstrations and then refines this prior using differentiable path-shape objectives, thereby separating demonstration-based learning of $\textbf{connectivity-biased guidance}$ from subsequent $\textbf{objective specific refinement}$. Beyond enabling adaptation to new routing preferences, the two-stage procedure improves standard shortest-path planning itself: on TMP, FlexPath improves optimal-path recovery from 75.0\% to 88.6\% over TransPath while reducing search expansions by 13.8\%. Ablations show that neither prior learning nor objective fine-tuning alone matches the full pipeline; their combination yields the strongest path cost and search efficiency. We further demonstrate the preference adaptation by adapting guidance to non-shortest-path objectives such as obstacle clearance, class-conditioned obstacle clearance and waypoint following. For clearance with $d_{\min}=2$, FlexPath achieves 96.2\% full clearance satisfaction on feasible instances while maintaining low search effort, and it reaches 98.4\% waypoint-following success.

cs.CV

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities. Our analysis shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning. This motivates closed-loop evaluation and highlights both the promise and risk of scaling-driven improvements for embodied AI.

cs.AI

Subsampling Confidence Bound for Persistent Diagram via Time-delay Embedding

Time-delay embedding is a fundamental technique in Topological Data Analysis (TDA) for reconstructing the phase space dynamics of time-series data. Persistent homology effectively identifies global topological features, such as loops associated with periodicity. Nevertheless, a statistically rigorous way to quantify uncertainty in the resulting topological features has remained underdeveloped -- a problem that we aim to challenge. First, we analyze the topological characterization of time-delay embeddings under both periodic and non-periodic conditions. Precisely, the embedded trajectory is homotopy equivalent to a circle ($S^1$) for periodic signals and is contractible for non-periodic ones. We also prove that the reach of the sliding window embedding is lower-bounded, ensuring stable persistence features. Next, we propose a subsampling-based method to construct confidence bounds for persistence diagrams derived from time-delay embeddings. Specifically, we derive confidence bounds with asymptotic guarantees, under the assumption that the support satisfies standard manifold regularity. Integrating the results, we propose a statistical testing framework to determine the periodicity of the underlying sampling function. This framework provides a principled statistical test for periodicity with asymptotically controlled type I and type II error rates. Simulation studies demonstrate that our method achieves detection performance comparable to the Generalized Lomb-Scargle Periodogram on periodic data while exhibiting superior robustness in distinguishing non-periodic signals with time-varying frequencies, such as chirp signals. Finally, it successfully captured the periodicity when applied to the BIDMC dataset.

math.ST

FDR Control via Neural Networks under Covariate-Dependent Symmetric Nulls

In modern multiple hypothesis testing, the availability of covariate information alongside the primary test statistics has motivated the development of more powerful and adaptive inference methods. However, most existing approaches rely on p-values that are precomputed under the assumption that their null distributions are independent of the covariates. In this paper, we propose a framework that derives covariate-adaptive p-values from the assumption of a symmetric null distribution of the primary variable given the covariates, without imposing any parametric assumptions. Building on these data-driven p-values, we employ a neural network model to learn a covariate-adaptive rejection threshold via the mirror estimation principle, optimizing the number of discoveries while maintaining valid false discovery rate control. Furthermore, our estimation of the conditional null distribution enables the computation of p-values directly from the raw data. The proposed method provides a principled way to derive covariate-adjusted p-values from raw data and allows seamless integration with previously established p-value based procedures. Simulation studies show that the proposed method outperforms existing approaches in terms of power. We further illustrate its applicability through two real data analyses: age-specific blood pressure data and U.S. air pollution data.

stat.ME

From Generation to Attribution: Music AI Agent Architectures for the Post-Streaming Era

Generative AI is reshaping music creation, but its rapid growth exposes structural gaps in attribution, rights management, and economic models. Unlike past media shifts, from live performance to recordings, downloads, and streaming, AI transforms the entire lifecycle of music, collapsing boundaries between creation, distribution, and monetization. However, existing streaming systems, with opaque and concentrated royalty flows, are ill-equipped to handle the scale and complexity of AI-driven production. We propose a content-based Music AI Agent architecture that embeds attribution directly into the creative workflow through block-level retrieval and agentic orchestration. Designed for iterative, session-based interaction, the system organizes music into granular components (Blocks) stored in BlockDB; each use triggers an Attribution Layer event for transparent provenance and real-time settlement. This framework reframes AI from a generative tool into infrastructure for a Fair AI Media Platform. By enabling fine-grained attribution, equitable compensation, and participatory engagement, it points toward a post-streaming paradigm where music functions not as a static catalog but as a collaborative and adaptive ecosystem.

cs.IR

High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight

Semantic segmentation from RGB cameras is essential to the perception of autonomous flying vehicles. The stability of predictions through the captured videos is paramount to their reliability and, by extension, to the trustworthiness of the agents. In this paper, we propose a lightweight video semantic segmentation approach-suited to onboard real-time inference-achieving high temporal consistency on aerial data through Semantic Similarity Propagation across frames. SSP temporally propagates the predictions of an efficient image segmentation model with global registration alignment to compensate for camera movements. It combines the current estimation and the prior prediction with linear interpolation using weights computed from the features similarities of the two frames. Because data availability is a challenge in this domain, we propose a consistency-aware Knowledge Distillation training procedure for sparsely labeled datasets with few annotations. Using a large image segmentation model as a teacher to train the efficient SSP, we leverage the strong correlations between labeled and unlabeled frames in the same training videos to obtain high-quality supervision on all frames. KD-SSP obtains a significant temporal consistency increase over the base image segmentation model of 12.5% and 6.7% TC on UAVid and RuralScapes respectively, with higher accuracy and comparable inference speed. On these aerial datasets, KD-SSP provides a superior segmentation quality and inference speed trade-off than other video methods proposed for general applications and shows considerably higher consistency. Project page: https://github.com/FraunhoferIVI/SSP.

cs.CV

Pseudo-Label Transfer from Frame-Level to Note-Level in a Teacher-Student Framework for Singing Transcription from Polyphonic Music

Lack of large-scale note-level labeled data is the major obstacle to singing transcription from polyphonic music. We address the issue by using pseudo labels from vocal pitch estimation models given unlabeled data. The proposed method first converts the frame-level pseudo labels to note-level through pitch and rhythm quantization steps. Then, it further improves the label quality through self-training in a teacher-student framework. To validate the method, we conduct various experiment settings by investigating two vocal pitch estimation models as pseudo-label generators, two setups of teacher-student frameworks, and the number of iterations in self-training. The results show that the proposed method can effectively leverage large-scale unlabeled audio data and self-training with the noisy student model helps to improve performance. Finally, we show that the model trained with only unlabeled data has comparable performance to previous works and the model trained with additional labeled data achieves higher accuracy than the model trained with only labeled data.

eess.AS

Understanding the Heart of the 5G Air Interface: An Overview of Physical Downlink Control Channel for 5G New Radio (NR)

New Radio (NR) is a new radio air interface developed by the 3rd Generation Partnership Project (3GPP) for the fifth generation (5G) mobile communications system. With great flexibility, scalability, and efficiency, 5G is expected to address a wide range of use-cases including enhanced mobile broadband (eMBB), ultra-reliable low-latency communications (URLLC), and massive machine type communications (mMTC). The physical downlink control channel (PDCCH) in NR carries Downlink Control Information (DCI). Understanding how PDCCH operates is key to developing a good understanding of how information is communicated over NR. This paper provides an overview of the 5G NR PDCCH by describing its physical layer structure, monitoring mechanisms, beamforming operation, and the carried information. We also share various design rationales that influence NR standardization.

cs.NI