arXiv ScienceSearch

arXiv subjects

Cheng Zhao

Publications and source records attributed to Cheng Zhao.

At least 19 recordsLinked to original sources

Exact PID and PI Gain Regions for Uncertain Non-Affine MIMO Systems

Characterizing all PID gains that guarantee uniform exponential regulation of uncertain nonlinear systems remains a fundamental problem. Most existing results provide only sufficient conditions derived from a particular Lyapunov construction. This paper gives exact gain regions, defined directly by uniform regulation over the entire uncertainty class, for two classes of nonaffine MIMO systems with $m\ge2$ controlled coordinates and input channels. For systems of second order under PID control, we independently characterize a common quadratic inner region and an outer region based on a linear subclass. The two regions coincide and therefore characterize the PID gain region exactly at the system level. For systems of first order under PI control, a lossless reduction to a linear family on the uncertainty boundary, combined with a common quadratic certificate, yields the exact PI region. Both regions strictly enlarge existing sufficient regions, and the PID result also provides a sequential tuning rule. A compact linear family with one uncertain parameter shows that common quadratic certification is not generally lossless. These results establish endpoint coincidence and boundary family equivalence as two routes to exact gain regions and motivate the broader question of which nonlinear uncertainty classes admit such characterizations.

eess.SY

On Global Regulatability of Robot Manipulators by Classical PID

A long-standing open problem in robot manipulator control is whether global regulation can be achieved by classical PID control. This paper provides an answer to this question for classical PID controllers with triple parameters (k_p,k_i,k_d) in R^3. We find and prove that for one-degree-of-freedom manipulators, the classical PID control guarantees global stability and asymptotic regulation under standard structural assumptions, and further derive explicit quantitative design conditions for the PID gains. However, for multi-degree-of-freedom cases, we can construct a robot manipulator satisfying the same structural assumptions for which no choice of PID gains (k_p,k_i,k_d) can achieve global asymptotic regulation. These results provide a fundamental understanding of the abovementioned open problem, revealing both the fundamental capability and intrinsic limitation of the classical PID control for robot manipulator dynamics.

cs.RO

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA

Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions. We introduce EndoCA, a paired complex-atomic answer consistency benchmark for evaluating whether complex answers remain consistent with same-image atomic answers. EndoCA contains two suites: EndoCA-Core evaluates compact question-complexity patterns commonly seen in practical endoscopic VQA, and EndoCA-Diagnostic supports controlled analysis across increasing question complexity. We evaluate 11 VLMs spanning open, medical, endoscopy-adapted, and closed-source models on EndoCA. Some VLMs achieve high complex-answer accuracy, yet their atomic-answer accuracy and complex-atomic answer consistency remain substantially lower. To reduce this complex-atomic inconsistency, we introduce Atomic-Support Reconciliation (ASR), a training-free mechanism that uses model-generated atomic answers as contextual premises for answer revision and consistency-guided selective answering. On four selected publicly available models, ASR-Revise improves paired complex-atomic correctness with modest changes in complex-answer accuracy, while ASR-Selective improves accuracy on answered cases by allowing the model to abstain from less reliable cases. Together, EndoCA and ASR provide a consistency-aware benchmark and a training-free mechanism for answer reconciliation and selective answering in endoscopic VQA.

cs.CV

Optimization of Tessellation-based Statistics: Void Statistics

Tessellation methods are extensively employed in the analyses of cosmic large-scale structure (LSS). However, these techniques are highly sensitive to perturbations in both densities and positions of points, often leading to substantial rearrangements of tessellation configurations. As a result, considerable additional statistical errors are introduced in various tessellation-based statistics, thereby weakening their cosmological constraints. In this work, we identify this issue and propose an efficacious measurement scheme through subsampling and averaging to enhance the stabilities of tessellation-based statistics. As a case study, we apply the new scheme to measure multiple primary void statistics [i.e., void size function (VSF), void two-point correlation function (VTCF), and void power spectrum (VPS)] in two distinct classes of voids, based on Delaunay and Voronoi tessellations, respectively. We notice that the statistical uncertainties in void statistics can be predominantly attributed to tessellation instabilities. Through rigorous testing, we demonstrate that the proposed method can substantially eliminate these scatters to deeply mine the statistical power of void statistics. Specifically, we find that our method can dramatically boost the signal-to-noise ratios (SNRs) of void Baryon Acoustic Oscillations (BAOs) and significantly improve the constraining power of void statistics on cosmological parameters. These findings showcase enormous application potentials of our new method in maximizing extraction of cosmological information from galaxy surveys. Importantly, our method is simple yet highly potent with broad applicability, hopefully evolving into a standard framework for measuring tessellation-based statistics in the future.

astro-ph.CO

Towards Guaranteed Optimal PID Tuning for Uncertain Nonlinear Systems

Despite the widespread use of PID controllers in engineering practice, designing optimal PID parameters has long been regarded as a challenging problem in both theory and practice, particularly when faced with uncertain nonlinear dynamical systems. Based on the authors' PID control theory established recently for MIMO nonlinear uncertain systems (Zhao and Guo, 2022), which provides a concrete PID parameter set for global stability of PID controlled systems, this paper further proposes a near-optimal PID tuning method, where only input-output (zeroth-order) data on the control performance is available. The tuning method is formulated as a constrained optimization problem and solved by an iterative learning algorithm, referred to as HRS-KW algorithm, that combines a hysteretic random search with the Kiefer-Wolfowitz algorithm, aiming at utilizing the advantages of both global exploration and local gradient acceleration. This method operates without requiring precise structural knowledge of the system dynamics, yet its almost sure convergence to an epsilon-optimal solution for the PID parameters can be guaranteed in theory while ensuring closed-loop system stability. Simulation results illustrate that our HRS-KW algorithm outperforms other related optimization methods, exhibiting better convergence to the prescribed epsilon-optimal performance set.

eess.SY

From Large Telescopes to the MUltiplexed Survey Telescope (MUST)

Recent advances in astronomical observations have ushered in an era of remarkable discoveries. We now probe the Universe through multi-messenger signals, image the sky with unprecedented depth and resolution, and investigate individual sources using powerful large-aperture telescopes. Yet, a critical gap persists: the lack of wide-field, highly multiplexed spectroscopic capabilities needed to fully exploit the wealth of imaging data from current and upcoming surveys. In this review, we trace the historical development of large optical telescopes and spectroscopic surveys, assess the capabilities of ongoing and near-future facilities, and motivate the need for next-generation Stage-V spectroscopic experiments. As a representative example, we present the MUltiplexed Survey Telescope (MUST), the first Stage-V spectroscopic facility currently under construction. MUST is a 6.5-meter telescope designed to obtain optical spectra for over 20,000 targets simultaneously within a $\sim$5 deg$^2$ field, using a modular focal plane populated with 6.2-mm pitch fiber-positioning robots. Over an 8-year survey in the 2030s, MUST aims to build the most comprehensive 3D spectroscopic map of the Universe to date, measuring redshifts for over 100 million galaxies and quasars and opening new windows into cosmology, Galactic structure, and time-domain astrophysics.

astro-ph.IM

Constraining Neutrino Mass with the Void Weak Lensing Effect

Cosmic voids, the underdense regions of the Large Scale Structure (LSS), provide cosmological information highly complementary to that obtained from overdense regions. In this work, we investigate the constraining power of the void-shear cross-correlation (void lensing effect) on the total neutrino mass. Based on cosmological $N$-body simulations with varying neutrino masses, we populate BOSS LOW-Z-like galaxies at $0.2<z<0.4$ using HOD fitting, identify voids with the DIVE void finder and obtain their density profiles from the underlying dark matter and neutrino distributions. We then generate mock shear catalogues through ray-tracing and measure the corresponding void lensing signals, assuming a source number density of $10/{\rm arcmin}^{2}$ and sky area of around $8400\,{\rm deg}^2$. Under this setup, void lensing independently yields a constraint on total neutrino mass as $\sigma(M_{\nu})=0.096\,{\rm eV}$ ($M_{\nu}<0.232\,{\rm eV}$, 95% C.L.) in the absence of shape noise, and $\sigma(M_{\nu})=0.340\,{\rm eV}$ ($M_{\nu}<0.707\,{\rm eV}$, 95% C.L.) when adopting a Stage-III-like shape noise. Moreover, we find a clear linear relationship between the void lensing signal and neutrino mass. We further validate the forward modelling of the void lensing signal from the void density profiles across different cosmologies, demonstrating its accuracy and potential for future applications. These findings highlight void lensing as a promising probe of massive neutrinos and motivate its applications to galaxy survey data as well as the combination with other cosmological observables.

astro-ph.CO

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent methods either rely on large-scale human preference annotations or operate on misaligned embeddings from pre-trained vision-language models, leading to limited scalability or suboptimal supervision. We present $\texttt{PISCES}$, an annotation-free post-training algorithm that addresses these limitations via a novel Dual Optimal Transport (OT)-aligned Rewards module. To align reward signals with human judgment, $\texttt{PISCES}$ uses OT to bridge text and video embeddings at both distributional and discrete token levels, enabling reward supervision to fulfill two objectives: (i) a Distributional OT-aligned Quality Reward that captures overall visual quality and temporal coherence; and (ii) a Discrete Token-level OT-aligned Semantic Reward that enforces semantic, spatio-temporal correspondence between text and video tokens. To our knowledge, $\texttt{PISCES}$ is the first to improve annotation-free reward supervision in generative post-training through the lens of OT. Experiments on both short- and long-video generation show that $\texttt{PISCES}$ outperforms both annotation-based and annotation-free methods on VBench across Quality and Semantic scores, with human preference studies further validating its effectiveness. We show that the Dual OT-aligned Rewards module is compatible with multiple optimization paradigms, including direct backpropagation and reinforcement learning fine-tuning. Project page: https://roar-ai.github.io/pisces

cs.CV

Counting voids and filaments: Betti Curves as a Topological Probe for Cosmology

Topological analysis of galaxy distributions has gathered increasing attention in cosmology, as they are able to capture non-Gaussian features of large-scale structures (LSS) that are overlooked by conventional two-point clustering statistics. We utilize Betti curves, a summary statistic derived from persistent homology, to characterize the multiscale topological features of the LSS, including connected components, loops, and voids, as a complementary cosmological probe. Using halo catalogs from the \textsc{Quijote} suite, we construct Betti curves, assess their sensitivity to cosmological parameters, and train automated machine learning based emulators to model their dependence on cosmological parameters. Our Bayesian inference recovers unbiased estimation of cosmological parameters, notably $n_{\mathrm{s}}$, $\sigma_8$, and $\Omega_{\mathrm{m}}$, while validation on sub-box simulations confirms robustness against cosmic variance. We further investigate the impact of redshift-space distortions (RSD) on Betti curves and demonstrate that including RSD enhances sensitivity to growth-related parameters. By jointly analyzing Betti curves and the power spectrum, we achieve significantly tightened constraints than using power spectrum alone on parameters such as $n_{\mathrm{s}}$, $\sigma_8$, and $w$. These findings highlight Betti curves -- especially when combined with traditional two-point statistics -- as a promising, interpretable tool for future galaxy survey analyses.

astro-ph.CO

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-space video-diffusion losses, which lack explicitly temporally aware, feature-level tracking guidance. We present \emph{Track4DGen}, a two-stage framework that couples a multi-view video diffusion model with a foundation point tracker and a hybrid 4D Gaussian Splatting (4D-GS) reconstructor. The central idea is to explicitly inject tracker-derived motion priors into intermediate feature representations for both multi-view video generation and 4D-GS. In Stage One, we enforce dense, feature-level point correspondences inside the diffusion generator, producing temporally consistent features that curb appearance drift and enhance cross-view coherence. In Stage Two, we reconstruct a dynamic 4D-GS using a hybrid motion encoding that concatenates co-located diffusion features (carrying Stage-One tracking priors) with Hex-plane features, and augment them with 4D Spherical Harmonics for higher-fidelity dynamics modeling. \emph{Track4DGen} surpasses baselines on both multi-view video generation and 4D generation benchmarks, yielding temporally stable, text-editable 4D assets. Lastly, we curate \emph{Sketchfab28}, a high-quality dataset for benchmarking object-centric 4D generation and fostering future research.

cs.CV

Necessary and Sufficient PID Gain Regions for Global Stabilization of Uncertain Second-Order MIMO Nonlinear Systems

As is well known, classical PID control is ubiquitous in industrial processes, yet a rigorous and explicit design theory for nonlinear uncertain MIMO second-order systems remains underdeveloped. In this paper we consider a class of such systems with both uncertain dynamics and an unknown but strictly positive input gain, where the nonlinear uncertainty is characterized by bounds on the Jacobian with respect to the state variables. We explicitly construct a three-dimensional region for the PID gains that is sufficient to guarantee global stability and asymptotic tracking of constant references for all nonlinearities satisfying these Jacobian bounds. We then derive a corresponding necessary region, thereby revealing the inherent conservatism required to cope with worst-case uncertainties. Moreover, under additional structural assumptions on the nonlinearities, these sufficient and necessary regions coincide, yielding a precise necessary-and-sufficient characterization of all globally stabilizing PID gains. All these regions are given in closed form and depend only on the prescribed Jacobian bounds and the known lower bound of the input gain, in contrast to many qualitative tuning methods in the literature.

eess.SY

Theory and Design of Extended PID Control for Stochastic Systems with Structural Uncertainties

Since the classical proportional-integral-derivative (PID) controller has continued to be the most widely used feedback methods in engineering systems by far, it is crucial to investigate the working mechanism of PID in dealing with nonlinearity, uncertainty and random noises. Recently, Zhao and Guo (2022) has established the global stability of PID control for a class of uncertain nonlinear control systems with relative degree two without random perturbations. In this paper, we will consider a more general class of nonlinear stochastic systems with an arbitrary relative degree $n$, and discuss the stability and design of extended PID controller (a natural extension of PID). We demonstrate that, the closed-loop control systems will be globally stable in mean square with bounded tracking errors provided the extended PID parameters are selected from an $(n+1)$-dimensional unbounded set, even if both the system nonlinear drift and diffusion terms contain a wide range of structural uncertainties. Moreover, the steady-state tracking error is proved to be proportional to the noise intensity at the setpoint, which can also be made arbitrarily small by choosing the controller parameters suitably large.

math.OC

Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding

Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning. However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens across diverse 3D understanding tasks, remain highly challenging. We present NDTokenizer3D, a generalist 3D VLM that performs a wide range of 3D scene understanding tasks while naturally supporting human interactions, thereby bridging language-level reasoning with 3D spatial understanding. The core of our approach is a novel three-stage scene tokenization pipeline built upon a Multi-Scale Normal Distributions Transform (NDT) representation, paired with a Multi-Scale NDT Decoder (MSDec). Specifically, NDTokenizer3D first constructs a multi-scale NDT representation from raw high-resolution point clouds, preserving both global context and fine-grained geometric details. Next, the MSDec progressively fuses cross-scale NDT features, producing holistic scene tokens consumable by LLM endpoints. Beyond tokenization, MSDec is repurposed as a general interface for human-interactive prompting (points, boxes, masks) and segmentation-mask decoding, unifying diverse 3D scene understanding tasks within a single architecture. With this compact and unified design, NDTokenizer3D offers a fine-grained, general-purpose 3D VLM, achieving remarkable improvements in 3D Referring Segmentation, 3D Visual Question Answering, and 3D Dense Captioning.

cs.CV

The Impact of Spectroscopic Redshift Errors on Cosmological Measurements

Spectroscopic redshift errors, including redshift uncertainty and catastrophic failures, can bias cosmological measurements from galaxy redshift surveys at sub-percent level. In this work, we investigate their impact on the full-shape analysis using contaminated mock catalogs. We find that redshift uncertainty introduces a scale-dependent damping effect on the power spectrum, which is absorbed by counterterms in clustering model, keeping parameter biases below $5\%$. Catastrophic failures suppress the power spectrum amplitude by an approximately constant factor that scales with the catastrophic rate $f_c$. While this effect is negligible for DESI galaxy populations ($f_c=1\%$), the slitless-like errors, combining redshift uncertainty with $f_c=5\%$ catastrophics, introduce significant biases in cosmological constraints. In this case, we observe $6\%$ to $16\%$ shifts ($\sim2.2\sigma$ level) in estimating the fractional growth rate $df\equiv f/f^{\rm{fid}}$ and the log primordial amplitude $\ln(10^{10} A_{s})$. Applying the correction factor $(1-f_c)^2$ on the galaxy power spectrum mitigates the bias but weakens the parameter constraints due to new degeneracies. Alternatively, fixing $f_c$ to its expected value restores the constraining power with a modest bias of $1.0\sigma$. Our results indicate that for space-based slitless surveys such as \textit{Euclid}, at minimum accurate estimation of $f_c$ and its incorporation into the clustering model are essential to get unbiased cosmological inference. Extending to evolving dark energy and massive neutrino cosmologies, redshift errors do not bias the dark energy properties parametrized by $w_0$ and $w_a$, but can degrade constraints on the summed neutrino mass $\sum m_\nu$ by up to 80% in the worst case.

astro-ph.CO

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of processing a large number of clips in long videos. To address this issue, we introduce DeCafNet, an approach employing ``delegate-and-conquer'' strategy to achieve computation efficiency without sacrificing grounding performance. DeCafNet introduces a sidekick encoder that performs dense feature extraction over all video clips in a resource-efficient manner, while generating a saliency map to identify the most relevant clips for full processing by the expert encoder. To effectively leverage features from sidekick and expert encoders that exist at different temporal resolutions, we introduce DeCaf-Grounder, which unifies and refines them via query-aware temporal aggregation and multi-scale temporal refinement for accurate grounding. Experiments on two LTVG benchmark datasets demonstrate that DeCafNet reduces computation by up to 47\% while still outperforming existing methods, establishing a new state-of-the-art for LTVG in terms of both efficiency and performance. Our code is available at https://github.com/ZijiaLewisLu/CVPR2025-DeCafNet.

cs.CV

Cosmological Constraints with Void Lensing I: the Simulation-Based Inference Framework

We present a Simulation-Based Inference (SBI) framework for cosmological parameter estimation via void lensing analysis. Despite the absence of an analytical model of void lensing, SBI can effectively learn posterior distributions through forward modeling of mock data. We develop a forward modeling pipeline that accounts for both cosmology and the galaxy-halo connection. By training a neural density estimator on simulated data, we infer the posteriors of two cosmological parameters, $\Omega_m$ and $S_8$. Validation tests are conducted on posteriors derived from different cosmological parameters and a fiducial sample. The results demonstrate that SBI provides unbiased estimates of mean values and accurate uncertainties. These findings highlight the potential to apply void lensing analysis to observational data even without an analytical void lensing model.

astro-ph.CO

Dynamical Dark Energy in light of the DESI DR2 Baryonic Acoustic Oscillations Measurements

Understanding whether cosmic acceleration arises from a cosmological constant or a dynamical component is a central goal of cosmology, and the Dark Energy Spectroscopic Instrument (DESI) enables stringent tests with high-precision distance measurements. We analyze baryon acoustic oscillation (BAO) measurements from DESI Data Release 1 (DR1) and Data Release 2 (DR2), combined with Type Ia supernovae and a cosmic microwave background (CMB) distance prior. With the larger statistical power and wider redshift coverage of DR2, the preference for dynamical dark energy does not diminish relative to DR1. Using both a shape-function reconstruction and non-parametric approaches with a Horndeski-motivated correlation prior, we find that the dark-energy equation of state $w(z)$ varies with redshift. BAO data alone yield modest constraints, but in combination with independent supernova compilations and the CMB prior they strengthen the evidence for dynamics. Bayesian model comparison shows moderate support for departures from $\Lambda$CDM when multiple degrees of freedom in $w(z)$ are allowed, corresponding to $\approx3\sigma$ tension with $\Lambda$CDM (and higher for some data sets). Despite methodological differences, our results are consistent with companion DESI papers, underscoring the complementarity of approaches. Possible systematics remain under study; forthcoming DESI, \emph{Euclid}, and next-generation CMB data will provide decisive tests.

astro-ph.CO

Measuring the Mean Free Path of HI Ionizing Photons at $3.2\leq z\leq4.6$ with DESI Y1 Quasars

The mean free path of ionizing photons ($\lambda_\mathrm{mfp}^{912}$) in the intergalactic medium (IGM) is a crucial quantity in modelling the ionization state of IGM and the extragalactic ultraviolet background (EUVB), and is widely used in hydrodynamical simulations of galaxies and reionization. We construct the largest quasar spectrum dataset to date -- 12,595 $\mathrm{S/N}>3$ spectra -- using the Y1 observation of Dark Energy Spectroscopic Instrument (DESI) to make the most precise model-independent measurement of the mean free path at $3.2\leq z\leq 4.6$. By stacking the spectra in 17 redshift bins and modelling the Lyman continuum profile, we get a redshift evolution $\lambda_\mathrm{mfp}^{912}\propto(1+z)^{-4.27}$ at $2\leq z\leq 5$, which is much shallower than previous estimate $\lambda_\mathrm{mfp}^{912}\propto(1+z)^{-5.4}$. We then explore the sources of systematic bias, including the choice of intrinsic quasar continuum, the consideration of Lyman series opacity and Lyman limit opacity evolution and the definition of $\lambda_\mathrm{mfp}^{912}$. Combining our results with estimates of $\lambda_\mathrm{mfp}^{912}$ at higher redshifts, we conclude at high confidence that the evolution in $\lambda_\mathrm{mfp}^{912}$ steepens at $z \approx 5$. We interpret this inflection as the transition from the end of HI reionization to a fully ionized plasma which characterizes the intergalactic medium of the past $\sim10$ billion years.

astro-ph.CO