arXiv ScienceSearch

arXiv subjects

Mei Yang

Publications and source records attributed to Mei Yang.

At least 19 recordsLinked to original sources

DualPathOcc: Dual-Resolution BEV Encoder for 3D Occupancy Prediction

Predicting 3D occupancy from multi-view images requires preserving geometric detail during 2D-to-3D lifting while reasoning over sparse, volumetric scene representations. We present DualPathOcc, a camera-based framework that combines a Spatial Enhancer for high-resolution feature aggregation before BEV compression, a SENet-augmented dual-path BEV encoder for local-global context modeling, and height-aware weighted cross-entropy for near-ground occupancy. The final model is optimized with occupancy supervision and no explicit depth loss. On single-frame Occ3D-nuScenes, DualPathOcc achieves 37.37 mIoU. We further analyze how surface-centered depth targets interact with volumetric occupancy learning.

cs.CV

Heavy Ball GMRESR method for Nonsymmetric Linear Systems

The heavy ball GMRES (HBGMRES) method is one of Krylov subspace methods for linear systems combined with the restarted GMRES and the heavy ball method which is applied in optimization. HBGMRES not only keeps benefit of the restarted GMRES in limiting memory usage and controlling orthogonalization cost, but also is able to cover up the slow convergence problem in the restarted GMRES. Another type of Krylov subspace methods is GMRESR which consists of GCR as the outer algorithm and certain steps of GMRES as an inner method. Compared with HBGMRES, GMRESR gives the approximately optimal solution over the new search vector gained from GMRES and all previously kept search vectors in GCR. Even though GMRESR performs better than HBGMRES, it has slow convergence which is similar to GMRES and the restarted GMRES. Inspired by HBGMRES, we present the heavy ball GMRESR method (HBGMRESR) by using HBGMRES as the inner loop instead of using GMRES to salvage the lost convergence speed while still keeping the benefit of GMRESR that in the sense of a global minimization over some specific part of the Krylov subspace is done. Numerical tests on real data are presented to demonstratee the superiority of the new methods over GMRESR and HBGMRES.

math.NA

Comb-enabled spectral-domain image transport through perturbation-prone multimode fibers

Multimode fibers (MMFs) offer a compact platform for imaging, sensing, and information transport, but their practical deployment is hindered by sensitivity to fiber perturbations, which alter modal coupling and invalidate conventional speckle-based calibrations. Here, we demonstrate perturbation-resilient image transport through MMFs by combining image-to-spectrum encoding with dual-comb spectroscopy. Two-dimensional images are converted into comb-line-resolved spectral signatures before fiber transmission, allowing spatial information to be carried in the spectral domain rather than in the output speckle field. After propagation, dual-comb heterodyne detection maps the encoded spectrum into the radio-frequency domain, enabling massively parallel spectral readout with a single photodetector. Neural-network-assisted compressive reconstruction further enables high-fidelity imaging from sparse, noisy, and spectrally aliased measurements. Our approach achieves Pearson correlation coefficients exceeding 0.9 under strong fiber perturbations and supports frame rates up to 2.5 MHz, allowing the observation of transient switching dynamics in a digital micromirror device. These results establish a powerful tool for robust, real-time image transport through flexible MMFs, with potential applications in remote sensing and fiber-based optical instrumentation.

physics.optics

Convergence Analysis of Two Alternating Iterative Schemes for Tucker Decomposition

The higher-order orthogonal iteration (HOOI) and the alternating subspace iteration (ASI) are two popular numerical methods for computing the Tucker decomposition of a multiple-mode tensor. Xu [Linear and Multilinear Algebra, 66(11):2247--2265, 2018] proposed a variation of HOOI, called the greedy HOOI, which has an extra alignment action between consecutive approximations. Kroonenberg and De Leeuw [Psychometrika, 45(1):69--97, 1980] analyzed the convergence of ASI but their analysis has gaps. These analysis were for a real tensor only. In this paper, we present detailed convergence analysis of the two methods that is applicable to a complex tensor with a real tensor being a special case, and it is shown both methods are globally convergent to stationary points under mild conditions while the objective function monotonically increases. Numerical examples are presented to demonstrate the convergence behavior of the methods.

math.NA

An NPDo Approach for Tensor Block-Diagonalization

This paper is concerned with Partial Tensor Block-Diagonalization of a multiway tensor by orthonormal matrices so that the extracted block-diagonal part optimally represents the tensor. The basic idea is to maximize the block-diagonal part via the tensor's mode-multiplications by orthonormal matrices. For that reason, it will be referred to Principal Tensor Block-Diagonalization (PTBD), which contains the Tucker decomposition (TD) of a tensor as a special case with just one block. Also as a special case is the approximate dominant tensor SVD in which each block-size is 1-by-1. An NPDo approach is proposed to optimize the block-diagonal part for computing \ptbd. It is shown the NPDo approach combined with Gauss-Seidel-type updating is globally convergent to a stationary point while the objective increases monotonically. Numerical experiments are presented to illustrate the efficiency of the NPDo approach.

math.NA

An NPDo Approach for Principal Joint SVD-type Block Diagonalization

This paper is concerned with partial Joint SVD-type Block Diagonalization of several matrices so that the extracted diagonal parts collectively optimally assume part of the total mass of all given matrices. For that reason, it will be referred also as Principal Joint SVD-type Block Diagonalization. When each block-size is 1-by-1, it is about finding a dominant partial joint SVD decomposition for the matrices of interests. An NPDo approach is proposed for maximizing the common dominant block-diagonal parts collectively. It is shown that the NPDo approach combined with Gauss-Seidel-type updating is globally convergent to a stationary point while the objective increases monotonically. Numerical experiments are presented to illustrate the efficiency of the NPDo approach.

math.NA

High-speed hyperspectral 3D ghost imaging LiDAR

Light detection and ranging (LiDAR) is widely used in autonomous systems and industrial metrology; however, the simultaneous acquisition of three-dimensional (3D) structure and broadband spectral information remains challenging, as conventional hyperspectral LiDAR relies on wavelength-scanning or spectrometer-based detection that limits speed. Here, we demonstrate a hyperspectral 3D ghost imaging LiDAR that eliminates these bottlenecks. By combining a stochastic broadband laser with single-pixel detection, and integrating spatiotemporal encoding with spectral ghost imaging in a time-of-flight framework, the system enables pulse-resolved recovery of spatial and spectral information. Consequently, we achieve a line-scanning rate of 60.5 MHz (point rate 1.8 GHz) and a ranging precision of 0.02 mm within a 10 {\mu}s integration time. Each voxel contains a 1.4 nm resolution spectrum over 1100-1250 nm, enabling simultaneous 3D imaging and chemical identification. This approach provides a route to high-speed hyperspectral LiDAR for environmental monitoring, precision agriculture, and industrial inspection.

physics.optics

Enhancing Image Aesthetics with Dual-Conditioned Diffusion Models Guided by Multimodal Perception

Image aesthetic enhancement aims to perceive aesthetic deficiencies in images and perform corresponding editing operations, which is highly challenging and requires the model to possess creativity and aesthetic perception capabilities. Although recent advancements in image editing models have significantly enhanced their controllability and flexibility, they struggle with enhancing image aesthetic. The primary challenges are twofold: first, following editing instructions with aesthetic perception is difficult, and second, there is a scarcity of "perfectly-paired" images that have consistent content but distinct aesthetic qualities. In this paper, we propose Dual-supervised Image Aesthetic Enhancement (DIAE), a diffusion-based generative model with multimodal aesthetic perception. First, DIAE incorporates Multimodal Aesthetic Perception (MAP) to convert the ambiguous aesthetic instruction into explicit guidance by (i) employing detailed, standardized aesthetic instructions across multiple aesthetic attributes, and (ii) utilizing multimodal control signals derived from text-image pairs that maintain consistency within the same aesthetic attribute. Second, to mitigate the lack of "perfectly-paired" images, we collect "imperfectly-paired" dataset called IIAEData, consisting of images with varying aesthetic qualities while sharing identical semantics. To better leverage the weak matching characteristics of IIAEData during training, a dual-branch supervision framework is also introduced for weakly supervised image aesthetic enhancement. Experimental results demonstrate that DIAE outperforms the baselines and obtains superior image aesthetic scores and image content consistency scores.

cs.CV

Dynamic stacking ensemble learning with investor knowledge representations for stock market index prediction based on multi-source financial data

The patterns of different financial data sources vary substantially, and accordingly, investors exhibit heterogeneous cognition behavior in information processing. To capture different patterns, we propose a novel approach called the two-stage dynamic stacking ensemble model based on investor knowledge representations, which aims to effectively extract and integrate the features from multi-source financial data. In the first stage, we identify different financial data property from global stock market indices, industrial indices, and financial news based on the perspective of investors. And then, we design appropriate neural network architectures tailored to these properties to generate effective feature representations. Based on learned feature representations, we design multiple meta-classifiers and dynamically select the optimal one for each time window, enabling the model to effectively capture and learn the distinct patterns that emerge across different temporal periods. To evaluate the performance of the proposed model, we apply it to predicting the daily movement of Shanghai Securities Composite index, SZSE Component index and Growth Enterprise index in Chinese stock market. The experimental results demonstrate the effectiveness of our model in improving the prediction performance. In terms of accuracy metric, our approach outperforms the best competing models by 1.42%, 7.94%, and 7.73% on the SSEC, SZEC, and GEI indices, respectively. In addition, we design a trading strategy based on the proposed model. The economic results show that compared to the competing trading strategies, our strategy delivers a superior performance in terms of the accumulated return and Sharpe ratio.

cs.CE

Reconstructive comb spectroscopy: A single-pixel detection paradigm beyond dual-comb limitations

Frequency comb spectroscopy has revolutionized broadband molecular fingerprinting with mode-defined resolution. While dual-comb spectroscopy stands as a dominant paradigm for high-resolution measurements, it relies on mutually coherent dual combs, and its applicability to non-cooperative sensing is limited by the requirement for phase-sensitive detection and controlled optical returns. Here, we introduce reconstructive comb spectroscopy, a fundamentally different paradigm that eliminates these constraints. By integrating a mode-programmable optical comb with a computational sensing scheme based on single-pixel detection, our method achieves picometer-level spectral resolution over a 10-nm (1.27-THz) instantaneous bandwidth, with single-photon sensitivity down to 10^-4 photons per pulse, and compressed spectral acquisition at 2.5% sampling while maintaining reconstruction errors below 10%. We demonstrate robust performance through scattering media and from non-cooperative targets. These capabilities establish reconstructive comb spectroscopy as a new platform for gas sensing, with broad applicability in remote atmospheric monitoring, industrial leak detection, and standoff chemical-threat identification.

physics.optics

CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process

Engineering design operates through hierarchical abstraction from system specifications to component implementations, requiring visual understanding coupled with mathematical reasoning at each level. While Multi-modal Large Language Models (MLLMs) excel at natural image tasks, their ability to extract mathematical models from technical diagrams remains unexplored. We present \textbf{CircuitSense}, a comprehensive benchmark evaluating circuit understanding across this hierarchy through 8,006+ problems spanning component-level schematics to system-level block diagrams. Our benchmark uniquely examines the complete engineering workflow: Perception, Analysis, and Design, with a particular emphasis on the critical but underexplored capability of deriving symbolic equations from visual inputs. We introduce a hierarchical synthetic generation pipeline consisting of a grid-based schematic generator and a block diagram generator with auto-derived symbolic equation labels. Comprehensive evaluation of six state-of-the-art MLLMs, including both closed-source and open-source models, reveals fundamental limitations in visual-to-mathematical reasoning. Closed-source models achieve over 85\% accuracy on perception tasks involving component recognition and topology identification, yet their performance on symbolic derivation and analytical reasoning falls below 19\%, exposing a critical gap between visual parsing and symbolic reasoning. Models with stronger symbolic reasoning capabilities consistently achieve higher design task accuracy, confirming the fundamental role of mathematical understanding in circuit synthesis and establishing symbolic reasoning as the key metric for engineering competence.

cs.CV

Rapid and precise distance measurement using balanced cross-correlation of a single frequency-modulated electro-optic comb

Ultra-rapid, high-precision distance metrology is critical for both advanced scientific research and practical applications. However, current light detection and ranging technologies struggle to simultaneously achieve high measurement speed, accuracy, and a large non-ambiguity range. Here, we present a time-of-flight optical ranging technique based on a repetition-frequency-modulated femtosecond electro-optic comb and balanced nonlinear cross-correlation detection. In this approach, a target distance is determined as an integer multiple of the comb repetition period. By rapidly sweeping the comb repetition frequency, we achieve absolute distance measurements within 500 ns and real-time displacement tracking at single-pulse resolution (corresponding to a refresh rate of 172 MHz). Furthermore, our system attains an ultimate ranging precision of 5 nm (with 0.3 s integration time). Our method uniquely integrates nanometer-scale precision, megahertz-level refresh rates, and a theoretically unlimited ambiguity range within a single platform, while also supporting multi-target detection. These advances pave the way for high-speed, high-precision ranging systems in emerging applications such as structural health monitoring, industrial manufacturing, and satellite formation flying.

physics.optics

Revealing Microscopic Objects in Fluorescence Live Imaging by Video-to-video Translation Based on A Spatial-temporal Generative Adversarial Network

In spite of being a valuable tool to simultaneously visualize multiple types of subcellular structures using spectrally distinct fluorescent labels, a standard fluoresce microscope is only able to identify a few microscopic objects; such a limit is largely imposed by the number of fluorescent labels available to the sample. In order to simultaneously visualize more objects, in this paper, we propose to use video-to-video translation that mimics the development process of microscopic objects. In essence, we use a microscopy video-to-video translation framework namely Spatial-temporal Generative Adversarial Network (STGAN) to reveal the spatial and temporal relationships between the microscopic objects, after which a microscopy video of one object can be translated to another object in a different domain. The experimental results confirm that the proposed STGAN is effective in microscopy video-to-video translation that mitigates the spectral conflicts caused by the limited fluorescent labels, allowing multiple microscopic objects be simultaneously visualized.

eess.IV

Detecting Wildfire Flame and Smoke through Edge Computing using Transfer Learning Enhanced Deep Learning Models

Autonomous unmanned aerial vehicles (UAVs) integrated with edge computing capabilities empower real-time data processing directly on the device, dramatically reducing latency in critical scenarios such as wildfire detection. This study underscores Transfer Learning's (TL) significance in boosting the performance of object detectors for identifying wildfire smoke and flames, especially when trained on limited datasets, and investigates the impact TL has on edge computing metrics. With the latter focusing how TL-enhanced You Only Look Once (YOLO) models perform in terms of inference time, power usage, and energy consumption when using edge computing devices. This study utilizes the Aerial Fire and Smoke Essential (AFSE) dataset as the target, with the Flame and Smoke Detection Dataset (FASDD) and the Microsoft Common Objects in Context (COCO) dataset serving as source datasets. We explore a two-stage cascaded TL method, utilizing D-Fire or FASDD as initial stage target datasets and AFSE as the subsequent stage. Through fine-tuning, TL significantly enhances detection precision, achieving up to 79.2% mean Average Precision (mAP@0.5), reduces training time, and increases model generalizability across the AFSE dataset. However, cascaded TL yielded no notable improvements and TL alone did not benefit the edge computing metrics evaluated. Lastly, this work found that YOLOv5n remains a powerful model when lacking hardware acceleration, finding that YOLOv5n can process images nearly twice as fast as its newer counterpart, YOLO11n. Overall, the results affirm TL's role in augmenting the accuracy of object detectors while also illustrating that additional enhancements are needed to improve edge computing performance.

cs.CV

Flexible Modified LSMR Method for Least Squares Problems

LSMR is a widely recognized method for solving least squares problems via the double QR decomposition. Various preconditioning techniques have been explored to improve its efficiency. One issue that arises when implementing these preconditioning techniques is the need to solve two linear systems per iterative step. In this paper, to tackle this issue, among others, a modified LSMR method (MLSMR), in which only one linear system per iterative step needs to be solved instead of two, is introduced, and then it is integrated with the idea of flexible GMRES to yield a flexible MLSMR method (FMLSMR). Numerical examples are presented to demonstrate the efficiency of the proposed FMLSMR method.

math.NA

Automated Quantification of White Blood Cells in Light Microscopic Images of Injured Skeletal Muscle

White blood cells (WBCs) are the most diverse cell types observed in the healing process of injured skeletal muscles. In the course of healing, WBCs exhibit dynamic cellular response and undergo multiple protein expression changes. The progress of healing can be analyzed by quantifying the number of WBCs or the amount of specific proteins in light microscopic images obtained at different time points after injury. In this paper, we propose an automated quantifying and analysis framework to analyze WBCs using light microscopic images of uninjured and injured muscles. The proposed framework is based on the Localized Iterative Otsu's threshold method with muscle edge detection and region of interest extraction. Compared with the threshold methods used in ImageJ, the LI Otsu's threshold method has high resistance to background area and achieves better accuracy. The CD68-positive cell results are presented for demonstrating the effectiveness of the proposed work.

eess.IV

Learning-to-solve unit commitment based on few-shot physics-guided spatial-temporal graph convolution network

This letter proposes a few-shot physics-guided spatial temporal graph convolutional network (FPG-STGCN) to fast solve unit commitment (UC). Firstly, STGCN is tailored to parameterize UC. Then, few-shot physics-guided learning scheme is proposed. It exploits few typical UC solutions yielded via commercial optimizer to escape from local minimum, and leverages the augmented Lagrangian method for constraint satisfaction. To further enable both feasibility and continuous relaxation for integers in learning process, straight-through estimator for Tanh-Sign composition is proposed to fully differentiate the mixed integer solution space. Case study on the IEEE benchmark justifies that, our method bests mainstream learning ways on UC feasibility, and surpasses traditional solver on efficiency.

eess.SY

Detection of Problem Gambling with Less Features Using Machine Learning Methods

Analytic features in gambling study are performed based on the amount of data monitoring on user daily actions. While performing the detection of problem gambling, existing datasets provide relatively rich analytic features for building machine learning based model. However, considering the complexity and cost of collecting the analytic features in real applications, conducting precise detection with less features will tremendously reduce the cost of data collection. In this study, we propose a deep neural networks PGN4 that performs well when using limited analytic features. Through the experiment on two datasets, we discover that PGN4 only experiences a mere performance drop when cutting 102 features to 5 features. Besides, we find the commonality within the top 5 features from two datasets.

cs.LG