arXiv ScienceSearch

arXiv subjects

Juan Zhang

Publications and source records attributed to Juan Zhang.

At least 19 recordsLinked to original sources

A counterexample to Huang's weak majorization conjecture

We give a counterexample to a weak majorization problem, proposed by Z. Huang (Linear Algebra Appl., 434 (2) (2011) 457--462), for singular values of Hadamard products of nonnegative matrices. The problem asks whether, for all nonnegative matrices $A$ and $B$, \begin{equation*} \bigl\{s_j^2(A\Had B)\bigr\} \prec_w \bigl\{s_j(A\Had A)s_j(B\Had B)\bigr\} \end{equation*} holds, where $\Had$ denotes the Hadamard product, $\prec_w$ means weak majorization, and $s_j(\cdot)$ is the $j$th largest singular value of a matrix. To answer this problem, we construct two $3\times3$ entrywise positive symmetric matrices. We give rigorous upper and lower bounds for the relevant singular-value sums by Sturm's theorem, and show that the inequality fails for the partial sum with $k=2$.

math.RA

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Deploying large language models (LLMs) is challenging due to their massive parameters and high computational costs. Ultra low-bit quantization can significantly reduce storage and accelerate inference, but extreme compression (i.e., mean bit-width <= 2) often leads to severe performance degradation. To address this, we propose Squeeze10-LLM, effectively "squeezing" 16-bit LLMs' weights by 10 times. Specifically, Squeeze10-LLM is a staged mixed-precision post-training quantization (PTQ) framework and achieves an average of 1.6 bits per weight by quantizing 80% of the weights to 1 bit and 20% to 4 bits. We introduce Squeeze10LLM with two key innovations: Post-Binarization Activation Robustness (PBAR) and Full Information Activation Supervision (FIAS). PBAR is a refined weight significance metric that accounts for the impact of quantization on activations, improving accuracy in low-bit settings. FIAS is a strategy that preserves full activation information during quantization to mitigate cumulative error propagation across layers. Experiments on LLaMA and LLaMA2 show that Squeeze10-LLM achieves state-of-the-art performance for sub-2bit weight-only quantization, improving average accuracy from 43% to 56% on six zero-shot classification tasks--a significant boost over existing PTQ methods.

cs.LG

A \(3\times 3\) counterexample to Lin and Wimmer's rank-minimization conjecture associated with Roth's similarity theorem

We give a \(3\times 3\) counterexample, valid over every field, to a rank-minimization conjecture of Lin and Wimmer (Bull. Aust. Math. Soc., 84 (3) (2011), 441--443) related to Roth's similarity theorem for the Sylvester matrix equation. We also prove that, over the complex field, no counterexample can occur when one of the two matrix sizes is less than \(3\). Hence the example is dimensionally minimal over the complex field.

math.RA

A High-Accuracy Numerical Homogenization Framework for Quasiperiodic Hamilton--Jacobi Equations

In this work, we develop an accurate numerical homogenization framework for computing effective Hamiltonians of quasiperiodic Hamilton--Jacobi equations (QHJEs) with convex Hamiltonians of the form $H(x,p) = |p|^k/{k}-f(x), ~k>1$, where $f$ is quasiperiodic. Computing effective Hamiltonians in the quasiperiodic setting requires solving QHJEs posed on the whole space. Their solutions generally possess neither translational symmetry nor decay and may exhibit low regularity. These features pose substantial challenges for numerical computation. To address these difficulties, we introduce a quasiperiodic boundary condition, which allows the original whole-space problem to be treated on a bounded domain while preserving quasiperiodicity at the boundary. We then propose an SL--FPR scheme that combines a semi-Lagrangian approximation with the finite points recovery method and establish stability and error estimates for the resulting scheme. We also extend the quasiperiodic homogenization result from the quadratic case to general $k>1$ and apply the proposed method to accurately approximate the corresponding effective Hamiltonians. Numerical experiments illustrate the convergence and applicability of the method and validate the extended homogenization results.

math.NA

Ab initio-based Deep-Learning Prediction of Carrier Mobility in Strongly Anharmonic Materials

Predicting charge transport in strongly anharmonic materials, particularly ultralow thermal conductors, remains a major challenge for first-principles methods. In such systems, perturbative treatments of electron-phonon interactions and the harmonic phonon picture often break down, necessitating non-perturbative approaches. The ab initio Kubo-Greenwood(aiKG) formalism provides a rigorous framework for evaluating temperature-dependent carrier transport beyond the harmonic approximation. Nevertheless, its practical application is computationally demanding because it requires large supercells, extensive statistical sampling, and extrapolation to the zero-frequency limit. In this work, we introduce an artificial-intelligence(AI)-assisted aiKG framework that incorporates the deep-learning Hamiltonian model. By predicting the Kohn-Sham Hamiltonian with sub-meV accuracy for supercells of up to 250 atoms, the model bypasses the costly iterative self-consistent field calculations while retaining first-principles reliability within the scope of effects captured by the training data. Using a strongly anharmonic thermal insulator, potassium iodide(KI) as a benchmark system, we demonstrate that the proposed approach enables efficient simulations of electronic structure and transport properties from a large supercell. The framework reproduces temperature-dependent carrier mobilities, spectral functions, and effective masses in close agreement with the underlying density functional theory while reducing computational cost to 10%. These results suggest that the AI-assisted aiKG framework can make non-perturbative transport calculations tractable for strongly anharmonic materials, opening a scalable route towards realistic simulations and accelerated discovery of new functional materials.

cond-mat.mtrl-sci

A Structure-Exploiting Implicit-Explicit Trust Region Method for Computing Second-Order Stationary Points of the Landau-Brazovskii Model

This work focuses on the reliable computation of second-order stationary points in the high-dimensional nonconvex energy landscape of the Landau-Brazovskii (LB) model, a fundamental model for studying phases and phase transitions. For this purpose, we develop an efficient implicit-explicit trust region (IMEX-TR) method. Trust region (TR) methods can avoid saddle-point stagnation and guarantee convergence to second-order stationary points under appropriate conditions. However, their direct application to the LB model has been impractical because the Hessian is dense if treated directly. The proposed IMEX-TR method overcomes this difficulty by exploiting the Hessian's special structure: the linear interaction part is diagonal in reciprocal space, whereas the nonlinear bulk-energy part is diagonal in physical space. Based on this structure, we design an efficient solver for the TR subproblem that with globally convergent guarantee and enjoys FFT-based acceleration, with $\mathcal O(N \log N)$ complexity per iteration. Existing first-order gradient-based methods for the LB model only guarantee convergence to first-order stationary points and may stagnate at saddle points. In contrast, the proposed IMEX-TR method inherits the theoretical guarantee of converging to second-order stationary points while remaining computationally practical. Numerical experiments verify the theoretical properties of the algorithm and demonstrate its robustness in locating stable phases from different initial conditions. Numerical results also show that IMEX-TR can escape unstable stationary states reached by first-order schemes and converge to physically meaningful second-order stationary points. These results suggest that targeting second-order stationary points provides an effective computational paradigm for exploring complex free-energy landscapes and identifying stable or metastable states.

math.NA

NanoBTE: Fast Iterative Solution of the Phonon Boltzmann Transport Equation for Nanoscale Heat Transport

Nanoscale heat dissipation has become a critical challenge in advanced semiconductor devices, where phonon transport can strongly deviate from the classical Fourier description due to the boundary scattering and ballistic effects. In this work, we propose NanoBTE, a deterministic finite-volume solver for the non-gray phonon Boltzmann transport equation under the relaxation-time approximation. The solver supports complex two- and three-dimensional geometries, band-resolved phonon properties, discrete-ordinates angular quadrature, volumetric heat generation, and multiple phonon boundary conditions, including thermalizing, diffuse, and specular reflections. %To improve the efficiency of multiscale simulations, both sequential and synthetic iterative schemes are implemented, where the latter couples the microscopic phonon transport equation with a macroscopic diffusion-type temperature equation to accelerate convergence in near-diffusive regimes. Both sequential and synthetic iterative options are implemented for the steady-state solution. Furthermore, NanoBTE adopts a band-direction task decomposition strategy, enabling efficient MPI-based CPU parallelization and GPU acceleration of the dominant sparse transport operations.

cond-mat.mtrl-sci

Newton-Based Mixed Precision Iterative Refinement for Large-Scale Sparse Continuous-Time Algebraic Riccati Equations

We propose a Newton-based mixed precision iterative refinement framework for solving large-scale sparse continuous-time algebraic Riccati equations (CAREs). The framework computes the initial approximation and the inner Lyapunov correction equations in lower precision, while evaluating residuals and updating the solution in higher precision. To handle indefinite residuals and Newton correction terms in low-rank form, we introduce factor decomposition procedures with truncation strategies that preserve positive semidefiniteness and control rank growth. A first-order rounding error analysis derives a residual recurrence for the refinement process and relates stable mixed precision refinement to a Lyapunov operator conditioning threshold governed by the unit roundoff of the lower precision inner solves. We then present a concrete ADI-based realization, using NLR-ADI for the initial CARE approximation and LR-ADI for the inner Lyapunov correction equations. Compared with dense Lyapunov correction implementations, this realization reduces the main computations to shifted linear solves and low-rank factor operations, and we provide a solver-dependent complexity analysis. Numerical experiments on dense CARE over a range of condition numbers illustrate the conditioning effect described by the error analysis, and experiments on large-scale sparse CAREs show that the mixed precision framework is faster than the full double precision implementation while maintaining the same level of accuracy.

math.NA

InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both RGB and infrared (IR) inputs, and treat them equally during both training and inference, which compromises their robustness when the RGB modality becomes unreliable or unavailable. In this case, we propose \textbf{InfraNet}, an IR-centric quality-aware framework that regulates RGB guidance during training and supports flexible RGB--IR or IR-only deployment. InfraNet employs an asymmetric architecture where the primary IR pathway extracts multi-scale infrared features for predictions, while the auxiliary RGB pathway provides reliability-controlled supervisory signals. The core of InfraNet is \textbf{QualGate}, a quality-aware fusion module that learns a task-oriented control signal to suppress unreliable RGB guidance and compensate IR features during cross-modal training. Built upon InfraNet, we design two architectural variants: a lightweight IR-only architecture InfraNet-IR and an RGB--IR architecture InfraNet-RGB-IR. Our method is evaluated through extensive experiments on four benchmark datasets (LLVIP, FLIR-Aligned, M$^3$FD, and DroneVehicle), showing strong or competitive accuracy in challenging low-light and adverse weather conditions. Notably, InfraNet maintains high efficiency in IR-only inference, making it both accurate and computationally efficient.

cs.CV

Teaching Vision-Language-Action Models What to See and Where to Look

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather than action-grounded planning. As a result, the learned representations capture semantic knowledge but lack spatial dependencies crucial for reliable trajectory prediction. We propose DriveTeach-VLA, a framework that explicitly teaches VLAs what to see and where to look. Driving-aware Vision Distillation (DVD) injects driving-specific perceptual priors into the vision encoder, while 2D Trajectory-Guided Prompts (2D-TGP) provide spatial conditioning aligned with feasible driving trajectories. Together, they form a vision-guided learning pipeline: what to see (DVD pretraining) - where to look (TGP-guided SFT) - how to act (TGP-guided GRPO). DriveTeach-VLA achieves the state-of-the-art performance on NAVSIM and nuScenes. Our code is available at: https://github.com/ShivaTeam/DriveTeach-VLA.

cs.CV

Kernel-learning parameter prediction and evaluation in algebraic multigrid method for several PDEs

This paper explores the application of kernel learning methods for parameter prediction and evaluation in the Algebraic Multigrid Method (AMG), focusing on several Partial Differential Equation (PDE) problems. AMG is an efficient iterative solver for large-scale sparse linear systems, particularly those derived from elliptic and parabolic PDE discretizations. However, its performance heavily relies on numerous parameters, which are often set empirically and are highly sensitive to AMG's effectiveness. Traditional parameter optimization methods are either computationally expensive or lack theoretical support. To address this, we propose a Gaussian Process Regression (GPR)-based strategy to optimize AMG parameters and introduce evaluation metrics to assess their effectiveness. Trained on small-scale datasets, GPR predicts nearly optimal parameters, bypassing the time-consuming parameter sweeping process. We also use kernel learning techniques to build a kernel function library and determine the optimal kernel function through linear combination, enhancing prediction accuracy. In numerical experiments, we tested typical PDEs such as the constant-coefficient Poisson equation, variable-coefficient Poisson equation, diffusion equation, and Helmholtz equation. Results show that GPR-predicted parameters match grid search results in iteration counts while significantly reducing computational time. A comprehensive analysis using metrics like mean squared error, prediction interval coverage, and Bayesian information criterion confirms GPR's efficiency and reliability. These findings validate GPR's effectiveness in AMG parameter optimization and provide theoretical support for AMG's practical application.

math.NA

EditSR: Enhancing Neural Symbolic Regression via Edit-based Rectification

Neural symbolic regression models improve inference efficiency by shifting structural search to pretraining, but their one-pass autoregressive decoding is prone to error accumulation, which may lead to generating structurally incorrect expressions, especially in complex expression generation scenarios. Existing rectification strategies can alleviate this issue, but they often depend on restarting global search, thereby weakening the efficiency advantage of neural models, and remain susceptible to error accumulation. In this paper, we propose EditSR, a two-layer framework that combines a neural symbolic regression model in the first layer with an edit-based Rectifier in the second layer to achieve efficient prediction and post-hoc rectification. Instead of restarting the global search, we maintain rectification efficiency by pretraining the Rectifier. Specifically, we formulate the rectification process as a step-by-step state-transition chain starting from an incorrect expression, and develop a state-transition algorithm to construct supervised rectification chains for training the Rectifier. To ensure syntactic validity throughout rectification, each edit action is restricted to a syntactically valid space so that every edited expression remains parseable. In addition, because each edit decision is conditioned on the current state rather than the history, the Rectifier allows errors made in earlier steps to be rectified by subsequent edits, thereby reducing the risk of error accumulation. Extensive experiments and ablation studies show that EditSR substantially improves symbolic structure recovery with limited extra cost, with more pronounced gains on complex expressions, where one-pass autoregressive decoding is more susceptible to error accumulation.

cs.AI

Circuit-Inspired High-Order Neural Networks with Unified Neural Dynamics Modeling for PDE Solving and Visual Perception

Deep networks often rely on architectural heuristics to shape representation evolution, limiting their ability to model data governed by intrinsic dynamics. We present the Circuit-inspired High-Order Neural Network (CHONN), a modular framework that treats representation evolution as a latent potential process and increases its effective order through Kirchhoff-inspired cascade composition. A single Kirchhoff Neural Cell implements a stable first-order update, while serially composed cells form higher-order dynamical operators within one block. This construction is interpretable, numerically stable and compatible with common neural backbones. Theoretical analysis shows that cascaded cells induce end-to-end high-order operators, and controlled experiments demonstrate that intra-block high-order construction differs from generic depth stacking, especially on derivative-sensitive measures. Across steady-state operator learning, long-horizon physical forecasting and ImageNet-1K recognition, CHONN improves structural fidelity, rollout stability and visual representation learning. These results identify high-order circuit composition as a general principle for neural dynamics modeling.

cs.LG

Multiwavelength Analysis of PSR J0437-4715 with Pulse Profile Modeling

We present a multi-wavelength analysis of the nearby millisecond pulsar PSR J0437--4715, combining Hubble Space Telescope (HST) far-ultraviolet, ROSAT soft X-ray, and XMM-Newton X-ray data, to model its broadband emission and energy-resolved pulse profiles, and infer key stellar parameters via Bayesian inference. The broadband emission includes cold thermal, hot thermal, and non-thermal components: cold bulk surface emission is modeled with a non-magnetized partially-ionized hydrogen atmosphere; hot-spot emission adopts the pulse profile modeling technique with a non-magnetized fully-ionized hydrogen atmosphere model; and non-thermal emission is included as a phase-invariant power-law component. By adopting an informative prior on the hot-spot geometry informed by radio polarization position angle measurements, the joint multi-instrument analysis yields a statistically viable and radio-consistent solution with a gravitational mass of 1.38$\pm$0.03~M$_\odot$ and an equatorial circumferential radius of 13.25$_{-0.35}^{+0.34}$~km (68\% confidence intervals). The hot-spot geometry consists of two spherical caps with uniform temperature distributions: the primary hot spot is situated at a colatitude of $\approx$130$^{\circ}$, and the secondary hot spot lies at a colatitude of $\approx$9$^{\circ}$, close to the north pole. It yields tighter radius constraints than HST+ROSAT fits and shifts the radius posterior distribution to larger values relative to NICER-only fits. This work demonstrates the importance of multi-wavelength data in refining neutron star mass-radius measurements and resolving geometric degeneracies.

astro-ph.HE

Coefficient-level output-feedback stabilization of linear port-Hamiltonian descriptor systems

This paper studies coefficient-level, structure-preserving output-feedback stabilization of linear port-Hamiltonian (pH) descriptor systems. Existing stabilization conditions generally require explicit pH representations, which may be costly to compute. We consider descriptor systems for which only the coefficient matrices are available and for which a pH representation is known to exist but is not explicitly given. For proportional output feedback, we derive coefficient-level conditions that are equivalent to the known solvability criteria in the explicit pH setting. These conditions ensure that the closed-loop system is regular, impulse-free, asymptotically stable, and remains port-Hamiltonian. We further extend the framework to proportional-derivative output feedback and enable the assignment of a prescribed dynamical order. Under the proposed conditions, the proportional gain may be chosen as any symmetric positive definite matrix, and the derivative gain is constructed from coefficient-based decompositions, without computing a pH representation.

math.OC

HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing

Edge computing faces unprecedented resource orchestration challenges from multi-dimensional heterogeneity across device architectures, diverse task requirements in CPU-intensive, GPU-intensive, I/O-intensive, and dynamic network conditions. The edge environments demand real-time task processing within strict energy budgets, yet conventional approaches struggle with mixed continuous-discrete optimization while meeting deadline and energy constraints. This paper presents HiRL, a hierarchical reinforcement learning framework that decomposes complex resource orchestration into coordinated power control and task allocation decisions. Our approach separates continuous power management using the Twin Delayed Deep Deterministic Policy Gradient (TD3) and discrete task placement using Double Deep Q-Network (DDQN), unified through a coordination engine with five-dimensional queue state representation. We propose a heterogeneous assessment of resource compatibility with deadline-oriented prioritization and failure-penalized adaptive sampling to enhance decision quality under resource constraints. To improve practical applicability, the framework models comprehensive system dynamics including device mobility, queue congestion patterns, infrastructure heterogeneity, and priority-sensitive scheduling demands. Experimental results show that HiRL achieves effective latency-energy trade-offs with 28% latency reduction compared to Single-DDQN and maintains nearly 100% task completion rates under all load conditions. Compared to baseline algorithms, HiRL reduces energy consumption by up to 51% under low load while achieving 24% better latency performance than static optimization approaches under high load, establishing effective resource orchestration in heterogeneous edge environments.

cs.DC

The trigger and localization system of SVOM-GRM

The Space multi-band Variable Object Monitor (SVOM) is an astronomical satellite jointly developed by China and France, primarily focused on the detection of gamma-ray bursts (GRBs) and transient sources. The SVOM satellite was launched on 22nd June, 2024 with four payloads installed onboard. As one of payload, GRM comprises 3 gamma-ray detectors (each detector has an effective area of approximately 200~cm$^{2}$) with distinct pointing directions, enabling the temporal and spectral measurements as well as localization of GRBs in the energy range of 15-5000 keV. This article firstly introduces the on-board localization algorithm design for GRM and presents preliminary test results. Then, leveraging abundant ground-based computational resources, a joint fitting method for spectral and localization analysis using Monte Carlo Markov Chain (MCMC) is implemented. In contrast to the on-board localization algorithm, the on-ground MCMC method comprehensively considers the influence of spectral characteristics, thereby mitigating systematic biases. Finally, a systematic analysis based on this method is provided, highlighting the localization and spectral measurement capabilities of GRM. The preliminary localization analysis result for the on-board detected GRB 240629A by both GRM and Fermi/GBM shows that the localization result (error$\sim$4.14$^{\circ}$) of GRM is consistent with the Fermi/GBM result.

astro-ph.IM

Study on the detector energy response of SVOM/GRM

The SVOM mission is specifically designed to for the detection and localization of Gamma-Ray Bursts (GRBs) and subsequent follow-up observations. Among the four telescopes installed on the SVOM satellite, the Gamma-Ray Monitor (GRM) plays a crucial role in capturing the prompt emission of GRBs due to its wide field of view (FOV) and broad energy range. Accurate determination of the detector's energy response is vital for analyzing GRM data, particularly considering the significant impact of the atmospheric albedo effect on this response. This research focuses on deriving the detector's energy response and establishing a calibration database for the GRM, with particular emphasis on investigating the atmospheric albedo effect. The study shows that the contribution of albedo photons to the detector's effective area depends strongly on the orientation of the GRD line of sight (LoS) relative to Earth and on the incident direction of the GRB. When the GRD LoS is anti-Earth oriented, the albedo effect is minimal, with the highest proportion of albedo effective area accounting for approximately 10% of the total effective area. This occurs when the incident angle of the GRB is nearly perpendicular to the LoS. Conversely, if the GRD LoS is not pointing away from Earth and the GRB arrives from angles greater than about 90$^{\circ}$, the albedo component can become predominant, contributing up to around 100% of the total effective area. This is especially pronounced in the 8-20 keV range, where the direct effective area drops to zero due to the large GRB injection angle. Our results show that, it is necessary for GRM to consider the atmospheric albedo effects in detector response, otherwise the spectral and localization analyses will result in biased measurements.

astro-ph.HE