arXiv ScienceSearch

arXiv subjects

An Le

Publications and source records attributed to An Le.

11 recordsLinked to original sources

Layerwise Tunable Lifting Scheme for the Convolutional Neural Network

This work introduces a family of tunable lifting schemes for biorthogonal wavelet filter banks. We propose three lifting strategies: low-pass tuning (LS-LayLatt-LP), high-pass tuning (LS-LayLatt-HP), and a sequential lifting scheme that jointly adapts low- and high-frequency branches (LS-LayLatt-Sequential). All proposed designs are formulated using a lattice-based lifting structure, which guarantees invertibility and stability for arbitrary parameter values within the lifting functions. We evaluated the proposed methods by integrating them into a ResNet-18 backbone for image classification on the Describable Textures Dataset (DTD), as well as for anomaly detection on hazelnut images from the MVTec-AD dataset and private KRC102S dataset. Experimental results demonstrate consistent performance improvements across all evaluated tasks.

cs.CV

Autonomous operation of the DIAG0 diagnostic line for 6D phase-space monitoring at LCLS-II

Characterizing the full 6-dimensional phase-space distribution of beams from the LCLS-II photoinjector is essential for understanding and optimizing downstream accelerator performance. Long-term monitoring of this distribution is equally important for detecting drifts in machine state and implementing timely corrective actions. Continuous phase space characterization during routine operation demands reliable tomographic diagnostic measurements and fast, efficient reconstruction methods. In this work, we demonstrate the first fully autonomous 6-dimensional beam-tomography system deployed on the DIAG0 parasitic beamline at LCLS-II. Using machine-learning-based control algorithms, the system autonomously configures DIAG0 and executes tomographic manipulations within operational constraints, adaptively re-optimizing beamline parameters and scan ranges in response to changes in the incoming beam. Tomographic measurements are streamed to the S3DF computing cluster where generative analysis methods reconstruct the phase-space distribution. We demonstrate that this framework produces detailed 6-dimensional beam reconstructions at a cadence of one reconstruction every 5 to 10 minutes, enabling real-time, multi-hour monitoring of injector beam evolution with unprecedented fidelity. These results represent a significant step toward fully autonomous operation of accelerator beamlines with real-time beam diagnostics for current and next-generation accelerator facilities.

physics.acc-ph

Learnable Multi-level Discrete Wavelet Transforms for 3D Gaussian Splatting Frequency Modulation

3D Gaussian Splatting (3DGS) has emerged as a powerful approach for novel view synthesis. However, the number of Gaussian primitives often grows substantially during training as finer scene details are reconstructed, leading to increased memory and storage costs. Recent coarse-to-fine strategies regulate Gaussian growth by modulating the frequency content of the ground-truth images. In particular, AutoOpti3DGS employs the learnable Discrete Wavelet Transform (DWT) to enable data-adaptive frequency modulation. Nevertheless, its modulation depth is limited by the 1-level DWT, and jointly optimizing wavelet regularization with 3D reconstruction introduces gradient competition that promotes excessive Gaussian densification. In this paper, we propose a multi-level DWT-based frequency modulation framework for 3DGS. By recursively decomposing the low-frequency subband, we construct a deeper curriculum that provides progressively coarser supervision during early training, consistently reducing Gaussian counts. Furthermore, we show that the modulation can be performed using only a single scaling parameter, rather than learning the full 2-tap high-pass filter. Experimental results on standard benchmarks demonstrate that our method further reduces Gaussian counts while maintaining competitive rendering quality.

eess.IV

Adiabatic reverse annealing is robust to low-temperature decoherence

Adiabatic reverse annealing (ARA) is an improvement to conventional quantum annealing (QA) that uses an initial guess at the desired ground state to circumvent problematic phase transitions. Despite encouraging results in the closed-system setting, Ref. [1] has suggested on the basis of numerical simulations that ARA may lose its advantage in the presence of decoherence. Here, we revisit this problem from a more analytical perspective. Using the $p$-spin model as a solvable example, together with the adiabatic master equation to describe the effects of the environment (valid at weak coupling), we show that ARA can in fact succeed in open systems but that the temperature of the environment plays a key role. We first demonstrate that, in the adiabatic limit, the system will follow the instantaneous equilibrium state as long as the protocol does not pass through any (finite-temperature) phase transitions. Given this, there are two distinct mechanisms by which ARA can break down at high temperature: either there are no paths that avoid transitions, or the equilibrium state itself is disordered. When the temperature is sufficiently low that neither of these occur, then ARA succeeds. Remarkably, there are even situations in which the environment benefits ARA: we find parameter values for which no transition-avoiding paths exist at zero temperature but such paths appear at non-zero temperature.

quant-ph

WaveletGaussian: Wavelet-domain Diffusion for Sparse-view 3D Gaussian Object Reconstruction

3D Gaussian Splatting (3DGS) has become a powerful representation for image-based object reconstruction, yet its performance drops sharply in sparse-view settings. Prior works address this limitation by employing diffusion models to repair corrupted renders, subsequently using them as pseudo ground truths for later optimization. While effective, such approaches incur heavy computation from the diffusion fine-tuning and repair steps. We present WaveletGaussian, a framework for more efficient sparse-view 3D Gaussian object reconstruction. Our key idea is to shift diffusion into the wavelet domain: diffusion is applied only to the low-resolution LL subband, while high-frequency subbands are refined with a lightweight network. We further propose an efficient online random masking strategy to curate training pairs for diffusion fine-tuning, replacing the commonly used, but inefficient, leave-one-out strategy. Experiments across two benchmark datasets, Mip-NeRF 360 and OmniObject3D, show WaveletGaussian achieves competitive rendering quality while substantially reducing training time.

cs.CV

DWTGS: Rethinking Frequency Regularization for Sparse-view 3D Gaussian Splatting

Sparse-view 3D Gaussian Splatting (3DGS) presents significant challenges in reconstructing high-quality novel views, as it often overfits to the widely-varying high-frequency (HF) details of the sparse training views. While frequency regularization can be a promising approach, its typical reliance on Fourier transforms causes difficult parameter tuning and biases towards detrimental HF learning. We propose DWTGS, a framework that rethinks frequency regularization by leveraging wavelet-space losses that provide additional spatial supervision. Specifically, we supervise only the low-frequency (LF) LL subbands at multiple DWT levels, while enforcing sparsity on the HF HH subband in a self-supervised manner. Experiments across benchmarks show that DWTGS consistently outperforms Fourier-based counterparts, as this LF-centric strategy improves generalization and reduces HF hallucinations.

cs.CV

Biorthogonal Tunable Wavelet Unit with Lifting Scheme in Convolutional Neural Network

This work introduces a novel biorthogonal tunable wavelet unit constructed using a lifting scheme that relaxes both the orthogonality and equal filter length constraints, providing greater flexibility in filter design. The proposed unit enhances convolution, pooling, and downsampling operations, leading to improved image classification and anomaly detection in convolutional neural networks (CNN). When integrated into an 18-layer residual neural network (ResNet-18), the approach improved classification accuracy on CIFAR-10 by 2.12% and on the Describable Textures Dataset (DTD) by 9.73%, demonstrating its effectiveness in capturing fine-grained details. Similar improvements were observed in ResNet-34. For anomaly detection in the hazelnut category of the MVTec Anomaly Detection dataset, the proposed method achieved competitive and wellbalanced performance in both segmentation and detection tasks, outperforming existing approaches in terms of accuracy and robustness.

cs.CV

Tunable Wavelet Unit based Convolutional Neural Network in Optical Coherence Tomography Analysis Enhancement for Classifying Type of Epiretinal Membrane Surgery

In this study, we developed deep learning-based method to classify the type of surgery performed for epiretinal membrane (ERM) removal, either internal limiting membrane (ILM) removal or ERM-alone removal. Our model, based on the ResNet18 convolutional neural network (CNN) architecture, utilizes postoperative optical coherence tomography (OCT) center scans as inputs. We evaluated the model using both original scans and scans preprocessed with energy crop and wavelet denoising, achieving 72% accuracy on preprocessed inputs, outperforming the 66% accuracy achieved on original scans. To further improve accuracy, we integrated tunable wavelet units with two key adaptations: Orthogonal Lattice-based Wavelet Units (OrthLatt-UwU) and Perfect Reconstruction Relaxation-based Wavelet Units (PR-Relax-UwU). These units allowed the model to automatically adjust filter coefficients during training and were incorporated into downsampling, stride-two convolution, and pooling layers, enhancing its ability to distinguish between ERM-ILM removal and ERM-alone removal, with OrthLattUwU boosting accuracy to 76% and PR-Relax-UwU increasing performance to 78%. Performance comparisons showed that our AI model outperformed a trained human grader, who achieved only 50% accuracy in classifying the removal surgery types from postoperative OCT scans. These findings highlight the potential of CNN based models to improve clinical decision-making by providing more accurate and reliable classifications. To the best of our knowledge, this is the first work to employ tunable wavelets for classifying different types of ERM removal surgery.

eess.IV

From Coarse to Fine: Learnable Discrete Wavelet Transforms for Efficient 3D Gaussian Splatting

3D Gaussian Splatting has emerged as a powerful approach in novel view synthesis, delivering rapid training and rendering but at the cost of an ever-growing set of Gaussian primitives that strains memory and bandwidth. We introduce AutoOpti3DGS, a training-time framework that automatically restrains Gaussian proliferation without sacrificing visual fidelity. The key idea is to feed the input images to a sequence of learnable Forward and Inverse Discrete Wavelet Transforms, where low-pass filters are kept fixed, high-pass filters are learnable and initialized to zero, and an auxiliary orthogonality loss gradually activates fine frequencies. This wavelet-driven, coarse-to-fine process delays the formation of redundant fine Gaussians, allowing 3DGS to capture global structure first and refine detail only when necessary. Through extensive experiments, AutoOpti3DGS requires just a single filter learning-rate hyper-parameter, integrates seamlessly with existing efficient 3DGS frameworks, and consistently produces sparser scene representations more compatible with memory or storage-constrained hardware.

cs.CV

Learning to Reason over Scene Graphs: A Case Study of Finetuning GPT-2 into a Robot Language Model for Grounded Task Planning

Long-horizon task planning is essential for the development of intelligent assistive and service robots. In this work, we investigate the applicability of a smaller class of large language models (LLMs), specifically GPT-2, in robotic task planning by learning to decompose tasks into subgoal specifications for a planner to execute sequentially. Our method grounds the input of the LLM on the domain that is represented as a scene graph, enabling it to translate human requests into executable robot plans, thereby learning to reason over long-horizon tasks, as encountered in the ALFRED benchmark. We compare our approach with classical planning and baseline methods to examine the applicability and generalizability of LLM-based planners. Our findings suggest that the knowledge stored in an LLM can be effectively grounded to perform long-horizon task planning, demonstrating the promising potential for the future application of neuro-symbolic planning methods in robotics.

cs.RO

Carleson measures and Toeplitz operators on small Bergman spaces on the ball

We study the Carleson measures and the Toeplitz operators on the class of so-called small weighted Bergman spaces, introduced recently by Seip. A characterization of Carleson measures is obtained which extends Seip's results from the unit disc of $\mathbb C$ to the unit ball of $\mathbb C^n$. We use this characterization to give necessary and sufficient conditions for the boundedness and compactness of Toeplitz operators. Finally, we study the Schatten $p$ classes membership of Toeplitz operators for $1<p<\infty$.

math.FA