arXiv ScienceSearch

arXiv subjects

Zihao Shi

Publications and source records attributed to Zihao Shi.

4 recordsLinked to original sources

Occupancy-based Quantile Risk Control

Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack rigorous finite-sample guarantees. To address these limitations, we introduce Occupancy-based Quantile Risk Control (OQRC), a novel method that provides tight risk control bounds with finite-sample validity. Our key idea is to formulate risk control as a finite-occupancy problem by partitioning the loss space with the ordered calibration losses. Specifically, we estimate the distribution of test losses across the resulting bins and upper-bound the risk by the maximum loss attained within each bin. We then select the parameter $λ$ such that this upper bound does not exceed a predefined threshold $α$ with high probability $1-δ$. Theoretically, we establish a finite-sample guarantee showing that OQRC yields tight risk control bounds that converge to the optimal bounds at a provable rate of $\mathcal{O}_ p(n^{-1/2})$. Extensive experiments demonstrate the effectiveness of our method, reducing the risk gap by up to 78.64\% on common benchmarks.

stat.ML

Preserving Conservation Laws in the Time-Evolving Natural Gradient Method via Relaxation and Projection Techniques

Neural networks have demonstrated significant potential in solving partial differential equations (PDEs). While global approaches such as Physics-Informed Neural Networks (PINNs) offer promising capabilities, they often lack inherent temporal causality, which can limit their accuracy and stability for time-dependent problems. In contrast, local training frameworks that progressively update network parameters over time are naturally suited for evolving PDEs. However, a critical challenge remains: many physical systems possess intrinsic invariants -- such as energy or mass -- that must be preserved to ensure physically meaningful solutions. This paper addresses this challenge by enhancing the Time-Evolving Natural Gradient (TENG) method, a recently proposed local training framework. We introduce two complementary techniques: (i) a relaxation algorithm that ensures the target solution $u_{\text{target}}$ preserves both quadratic and general nonlinear invariants of the original system, providing a structure-preserving learning target; and (ii) a projection technique that maps the updated network parameters $θ(t)$ back onto the invariant manifold, ensuring the final neural network solution strictly adheres to the conservation laws. Numerical experiments on the inviscid Burgers equation, Korteweg-de Vries equation, and acoustic wave equation demonstrate that our proposed approach significantly improves conservation properties while maintaining high accuracy.

math.NA

Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Score

Conformal prediction (CP) is a powerful framework for uncertainty quantification, generating prediction sets with coverage guarantees. Split conformal prediction relies on labeled data in the calibration procedure. However, the labeled data is often limited in real-world scenarios, leading to unstable coverage performance in different runs. To address this issue, we extend CP to the semi-supervised setting and propose SemiCP, a new paradigm that leverages both labeled and unlabeled data for calibration. To achieve this, we introduce an unlabeled nonconformity score, Nearest Neighbor Matching (NNM) score. Specifically, NNM estimates the nonconformity scores of unlabeled samples using their most similar pseudo-labeled counterparts during calibration, while maintaining the original scores for labeled data. Theoretically, we demonstrate that the average coverage gap (i.e., the absolute difference between the empirical marginal coverage and the target coverage) of SemiCP can decrease significantly at a rate $\mathcal{O}(1/\sqrt{N})$ and converge to an error term, where $N$ is the number of unlabeled data. Extensive experiments validate the effectiveness of SemiCP under limited labeled data, reducing the average coverage gap by up to 77% on common benchmarks with 4000 unlabeled examples, when there are only 20 labeled examples.

cs.LG

Measuring Over-smoothing beyond Dirichlet energy

While Dirichlet energy serves as a prevalent metric for quantifying over-smoothing, it is inherently restricted to capturing first-order feature derivatives. To address this limitation, we propose a generalized family of node similarity measures based on the energy of higher-order feature derivatives. Through a rigorous theoretical analysis of the relationships among these measures, we establish the decay rates of Dirichlet energy under both continuous heat diffusion and discrete aggregation operators. Furthermore, our analysis reveals an intrinsic connection between the over-smoothing decay rate and the spectral gap of the graph Laplacian. Finally, empirical results demonstrate that attention-based Graph Neural Networks (GNNs) suffer from over-smoothing when evaluated under these proposed metrics.

cs.LG