arXiv ScienceSearch

arXiv subjects

Ziqing Zhang

Publications and source records attributed to Ziqing Zhang.

14 recordsLinked to original sources

Kullback-Leibler Mirror-Prox for Measure-Valued Variational Inequalities and Mean-Field Equilibria

We study the computation of static mean-field equilibria on a compact state space by formulating the equilibrium condition as a variational inequality over probability measures. We propose an entropic variant of Korpelevich's extragradient algorithm---the Kullback--Leibler Mirror-Prox method---in which Euclidean projections are replaced by relative-entropy proximal steps. Each half-step is therefore an explicit exponential reweighting of the current measure, implemented on a finite state-space discretization. Under Lasry--Lions monotonicity and continuity assumptions, we prove convergence of mesh-refined ergodic averages and obtain finite-iteration Minty-residual and approximate-equilibrium bounds that jointly quantify iteration and discretization errors. Under strong monotonicity, we derive metric convergence rates for the last, best, and averaged iterates. We also develop a KL-type Tikhonov regularization that selects the equilibrium minimizing relative entropy with respect to a reference measure. The framework applies to potential and nonpotential cost operators and does not require differentiability or convexity of the cost in the individual state.

math.OC

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi-step inference. Such capability is critical for real-world applications where models must reliably interpret cluttered environments. Yet existing training signals provide no explicit grounding between reasoning steps and the underlying visual entities and relations, leaving lightweight models free to generate fluent but visually unanchored reasoning chains. To address this gap, we first introduce DRBench, a benchmark of 14,573 questions across 2,943 images, organized into five task categories spanning three progressive reasoning layers. Building on DRBench, we propose DRScaffold, a supervised fine-tuning framework that decomposes the supervision target into four causally ordered stages, enforcing grounded reasoning without architectural modification. Experiments on three lightweight VLMs demonstrate substantial gains on DRBench while preserving or improving performance on general-purpose benchmarks. Notably, Qwen2.5-VL-3B trained with DRScaffold surpasses the frozen Qwen2.5-VL-32B on DRBench, demonstrating that structured supervision can substitute for a significant portion of model scale in dense-scene reasoning. Our code and models are available at https://github.com/irene-shi/DRScaffold .

cs.CV

The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high-resolution (HR) infrared images from low-resolution (LR) inputs generated through bicubic downsampling with a x4 scaling factor. The objective is to develop effective models or solutions that achieve state-of-the-art performance for infrared image SR in remote sensing scenarios. To reflect the characteristics of infrared data and practical application needs, the challenge adopts a single-track setting. A total of 115 participants registered for the competition, with 13 teams submitting valid entries. This report summarizes the challenge design, dataset, evaluation protocol, main results, and the representative methods of each team. The challenge serves as a benchmark to advance research in infrared image super-resolution and promote the development of effective solutions for real-world remote sensing applications.

cs.CV

The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview

This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze recent advances in the field. To reflect the evolving objectives of image super-resolution, the challenge includes two tracks: (1) a restoration track, which emphasizes pixel-wise fidelity and ranks submissions based on PSNR; and (2) a perceptual track, which focuses on visual realism and evaluates results using a perceptual score. A total of 194 participants registered for the challenge, with 31 teams submitting valid entries. This report summarizes the challenge design, datasets, evaluation protocol, main results, and methods of participating teams. The challenge provides a unified benchmark and offers insights into current progress and future directions in image super-resolution.

cs.CV

Continuous-time Online Learning via Mean-Field Neural Networks: Regret Analysis in Diffusion Environments

We study continuous-time online learning where data are generated by a diffusion process with unknown coefficients. The learner employs a two-layer neural network, continuously updating its parameters in a non-anticipative manner. The mean-field limit of the learning dynamics corresponds to a stochastic Wasserstein gradient flow adapted to the data filtration. We establish regret bounds for both the mean-field limit and finite-particle system. Our analysis leverages the logarithmic Sobolev inequality, Polyak-Lojasiewicz condition, Malliavin calculus, and uniform-in-time propagation of chaos. Under displacement convexity, we obtain a constant static regret bound. In the general non-convex setting, we derive explicit linear regret bounds characterizing the effects of data variation, entropic exploration, and quadratic regularization. Finally, our simulations demonstrate the outperformance of the online approach and the impact of network width and regularization parameters.

cs.LG

Combining scEEG and PPG for reliable sleep staging using lightweight wearables

Reliable sleep staging remains challenging for lightweight wearable devices such as single-channel electroencephalography (scEEG) or photoplethysmography (PPG). scEEG offers direct measurement of cortical activity and serves as the foundation for sleep staging, yet exhibits limited performance on light sleep stages. PPG provides a low-cost complement that captures autonomic signatures effective for detecting light sleep. However, prior PPG-based methods rely on full night recordings (8 - 10 hours) as input context, which is less practical to provide timely feedback for sleep intervention. In this work, we investigate scEEG-PPG fusion for 4-class sleep staging under short-window (30 s - 30 min) constraints. First, we evaluate the temporal context required for each modality, to better understand the relationship of sleep staging performance with respect to monitoring window. Second, we investigate three fusion strategies: score-level fusion, cross-attention fusion enabling feature-level interactions, and Mamba-enhanced fusion incorporating temporal context modeling. Third, we train and evaluate on the Multi-Ethnic Study of Atherosclerosis (MESA) dataset and perform cross-dataset validation on the Cleveland Family Study (CFS) and the Apnea, Bariatric surgery, and CPAP (ABC) datasets. The Mamba-enhanced fusion achieves the best performance on MESA (Cohen's Kappa $\kappa$ = 0.798, Acc = 86.9%), with particularly notable improvement in light sleep classification (F1-score: 85.63% vs. 77.76%, recall: 82.85% vs. 69.95% for scEEG alone), and generalizes well to CFS and ABC datasets with different populations. These findings suggest that scEEG-PPG fusion is a promising approach for lightweight wearable based sleep monitoring, offering a pathway toward more accessible sleep health assessment. Source code of this project can be found at: https://github.com/DavyWJW/scEEG-PPGFusion

eess.SP

Entropy-Based Dimension-Free Convergence and Loss-Adaptive Schedules for Diffusion Models

Diffusion generative models synthesize samples by discretizing reverse-time dynamics driven by a learned score (or denoiser). Existing convergence analyses of diffusion models typically scale at least linearly with the ambient dimension, and sharper rates often depend on intrinsic-dimension assumptions or other geometric restrictions on the target distribution. We develop an alternative, information-theoretic approach to dimension-free convergence that avoids any geometric assumptions. Under mild assumptions on the target distribution, we bound KL divergence between the target and generated distributions by $O(H^2/K)$ (up to endpoint factors), where $H$ is the Shannon entropy and $K$ is the number of sampling steps. Moreover, using a reformulation of the KL divergence, we propose a Loss-Adaptive Schedule (LAS) for efficient discretization of reverse SDE which is lightweight and relies only on the training loss, requiring no post-training heavy computation. Empirically, LAS improves sampling quality over common heuristic schedules.

cs.LG

InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution

Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) inefficiency due to the heavy cost of multi-step denoising for full-length sequences; and (2) poor consistency is hindered by temporal decomposition that causes artifacts and discontinuities. To break these limits, we propose InfVSR, which reformulates VSR as an autoregressive-one-step-diffusion paradigm, and enables streaming inference with video diffusion priors. First, we adapt the pretrained DiT into a causal structure, maintaining both local and global coherence via rolling KV-cache and joint visual guidance. Second, we distill the diffusion process into a single step efficiently, with patch-wise pixel supervision and cross-chunk distribution matching. To fill the gap in long-form video evaluation, we build a new benchmark tailored for extended sequences and further introduce semantic-level metrics to comprehensively assess temporal consistency. Our method pushes the frontier of long-form VSR, achieves state-of-the-art quality with enhanced semantic consistency, and delivers up to 58x speed-up over existing methods such as MGLD-VSR. Our code and models are available at https://github.com/Kai-Liu001/InfVSR.

cs.CV

Bridging Quantum Mechanics to Liquid Properties via a Universal Organic Force Field

Molecular dynamics simulations are essential tools for unraveling atomic-level insights into the structure and behavior of condensed-phase systems. However, the universal and accurate prediction of macroscopic properties based on quantum mechanical calculations remains a significant challenge, often hindered by the trade-off between computational cost and simulation accuracy. Here we present ByteFF-Pol, a polarizable force field parameterized by a graph neural network and trained exclusively on high-level quantum mechanical data. By leveraging physically-motivated force field forms and training strategies, ByteFF-Pol predicts thermodynamic and transport properties for a wide range of small-molecule liquids and electrolytes with high accuracy, surpassing current classical and machine learning force fields. This ability to make predictions without system-specific training bridges the gap between microscopic calculations and macroscopic liquid properties, enabling the exploration of previously intractable chemical spaces. This advancement enables the precise design of new electrolytes and custom-tailored solvents, establishing a robust foundation for data-driven materials discovery.

physics.comp-ph

A Unified Predictive and Generative Solution for Liquid Electrolyte Formulation

Liquid electrolytes are critical components of next-generation energy storage systems, enabling fast ion transport, minimizing interfacial resistance, and ensuring electrochemical stability for long-term battery performance. However, measuring electrolyte properties and designing formulations remain experimentally and computationally expensive. In this work, we present a unified framework for designing liquid electrolyte formulation, integrating a forward predictive model with an inverse generative approach. Leveraging both computational and experimental data collected from literature and extensive molecular simulations, we train a predictive model capable of accurately estimating electrolyte properties from ionic conductivity to solvation structure. Our physics-informed architecture preserves permutation invariance and incorporates empirical dependencies on temperature and salt concentration, making it broadly applicable to property prediction tasks across molecular mixtures. Furthermore, we introduce -- to the best of our knowledge -- the first generative machine learning framework for molecular mixture design, demonstrated on electrolyte systems. This framework supports multi-condition-constrained generation, addressing the inherently multi-objective nature of materials design. As a proof of concept, we experimentally identified three liquid electrolytes with both high ionic conductivity and anion-concentrated solvation structure. This unified framework advances data-driven electrolyte design and can be readily extended to other complex chemical systems beyond electrolytes.

cond-mat.mtrl-sci

Dog-IQA: Standard-guided Zero-shot MLLM for Mix-grained Image Quality Assessment

Image quality assessment (IQA) serves as the golden standard for all models' performance in nearly all computer vision fields. However, it still suffers from poor out-of-distribution generalization ability and expensive training costs. To address these problems, we propose Dog-IQA, a standard-guided zero-shot mix-grained IQA method, which is training-free and utilizes the exceptional prior knowledge of multimodal large language models (MLLMs). To obtain accurate IQA scores, namely scores consistent with humans, we design an MLLM-based inference pipeline that imitates human experts. In detail, Dog-IQA applies two techniques. First, Dog-IQA objectively scores with specific standards that utilize MLLM's behavior pattern and minimize the influence of subjective factors. Second, Dog-IQA comprehensively takes local semantic objects and the whole image as input and aggregates their scores, leveraging local and global information. Our proposed Dog-IQA achieves state-of-the-art (SOTA) performance compared with training-free methods, and competitive performance compared with training-based methods in cross-dataset scenarios. Our code will be available at https://github.com/Kai-Liu001/Dog-IQA.

cs.CV

Handling errors in four-dimensional variational data assimilation by balancing the degrees of freedom and the model constraints: A new approach

For many years, strongly and weakly constrained approaches were the only options to deal with errors in four-dimensional variational data assimilation (4DVar), with the aim of balancing the degrees of freedom and model constraints. Strong model constraints were imposed to reduce the degrees of freedom encountered when optimizing the strongly constrained 4DVar problem, and it was assumed that the models were perfect. The weakly constrained approach sought to distinguish initial errors from model errors, and to correct them separately using weak model constraints. Our proposed i4DVar* method exploits the hidden mechanism that corrects initial and model errors simultaneously in the strongly constrained 4DVar. The i4DVar* method divides the assimilation window into several sub-windows, each of which has a unique integral and flow-dependent correction term to simultaneously handle the initial and model errors over a relatively short period. To overcome the high degrees of freedom of the weakly constrained 4DVar, for the first time we use ensemble simulations not only to solve the 4DVar optimization problem, but also to formulate this method. Thus, the i4DVar* problem is solvable even if there are many degrees of freedom. We experimentally show that i4DVar* provides superior performance with much lower computational costs than existing methods, and is simple to implement.

physics.ao-ph

BayesFT: Bayesian Optimization for Fault Tolerant Neural Network Architecture

To deploy deep learning algorithms on resource-limited scenarios, an emerging device-resistive random access memory (ReRAM) has been regarded as promising via analog computing. However, the practicability of ReRAM is primarily limited due to the weight drifting of ReRAM neural networks due to multi-factor reasons, including manufacturing, thermal noises, and etc. In this paper, we propose a novel Bayesian optimization method for fault tolerant neural network architecture (BayesFT). For neural architecture search space design, instead of conducting neural architecture search on the whole feasible neural architecture search space, we first systematically explore the weight drifting tolerance of different neural network components, such as dropout, normalization, number of layers, and activation functions in which dropout is found to be able to improve the neural network robustness to weight drifting. Based on our analysis, we propose an efficient search space by only searching for dropout rates for each layer. Then, we use Bayesian optimization to search for the optimal neural architecture robust to weight drifting. Empirical experiments demonstrate that our algorithmic framework has outperformed the state-of-the-art methods by up to 10 times on various tasks, such as image classification and object detection.

cs.LG

Graphical Direct-Writing of Macroscale Domain Structures with Nanoscale Spatial Resolution in Non-Polar-Cut Lithium Niobate on Insulators

We reported on a graphical domain engineering technique with the capability to fabricate macroscale domain structures with nanoscale spatial resolution in non-polar-cut lithium niobate thin film on insulators through the biased probe tip of scanning atomic force microscopy. It was found that the domain writing process is asymmetric with respect to the spontaneous polarization Ps even though the tip-induced poling field is mirror-symmetric. Various domain structures, with a dimension larger than millimeters while consisting of nanoscale domain elements and with arbitrary domain-wall inclination angle with respect to Ps, were designed graphically and then written directly into non-polar-cut lithium niobate crystals. As a proof of principle demonstration, periodically poled x-cut lithium niobate thin film on insulators with a period of 600 nm, a depth of 460 nm and a length of ~1 mm was fabricated. This technique could be useful for device applications in integrated optics and opto-electronics and domain-wall nanoelectronics based on lithium niobate on insulator.

physics.app-ph