arXiv ScienceSearch

arXiv subjects

Wei Qian

Publications and source records attributed to Wei Qian.

At least 19 recordsLinked to original sources

A Framework for Directed Acyclic Hypergraph Learning

Continuous optimization methods for learning Directed Acyclic Graphs (DAGs) operate on weighted adjacency matrices and are therefore limited to pairwise causal relationships. We propose a framework for learning Directed Acyclic Hypergraphs (DAHGs) from observational data, capturing joint parental influences that pairwise models cannot represent. Our approach rests on three components: (i) a generalized linear structural equation model (SEM) with multiplicative interaction terms whose non-zero weights correspond one-to-one with directed hyperedges; (ii) a weighted adjacency tensor representation whose acyclicity is characterized via nilpotency under the tensor t-product; and (iii) a differentiable acyclicity constraint derived through the Fourier decomposition of the t-product, which reduces tensor nilpotency to slice-wise matrix nilpotency and enables least-squares learning via the augmented Lagrangian method.

cs.LG

Beyond Convolution: Advancing Hypergraph Neural Networks with Hypergraph U-Nets

Convolutions have successfully transitioned from image processing to the complex realm of non-Euclidean higher-order domains, particularly in hypergraphs. Despite the success in convolution, the exploration of a popular architecture named U-Net remains largely unexplored for hypergraph data due to the lack of well-defined pooling and unpooling operations. This work pioneers the study of U-Net architectures for hypergraph data, addressing the critical challenge of designing effective pooling and unpooling operations that retain maximal structural information from the input hypergraph. Motivated by hierarchical clustering, we propose to construct the pooling and unpooling operators all at once by cutting the clustering dendrogram at different granularities, named the Parallel Hierarchical Pooling (PHPool) and Unpooling (PHUnpool) operators. Unlike existing pooling methods that risk local structural damage through a sequential learning procedure, our PHPool operators are designed in a global and parallel manner to ensure fidelity to the original hypergraph structure with efficient computation while the PHUnpool operators are tailored to perform inverse operations of the PHPools for hypergraph reconstruction. We validate our model through hypergraph reconstruction simulation, hypergraph classification, and node-level anomaly detection, where it demonstrates superior performance over existing state-of-the-art graph and hypergraph deep learning methods.

cs.LG

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition

Micro-gesture online recognition aims to temporally localize and classify subtle gestures in untrimmed videos. Owing to their extremely short duration, low motion amplitude, and ambiguous visual cues, capturing discriminative spatiotemporal representations remains highly challenging. Existing parameter-efficient adapters typically employ a single branch to model spatial and temporal cues jointly, which may fail to capture the fine-grained patterns of micro-gestures. To address this limitation, we propose a Spatial-Temporal Decoupled Adapter that decomposes video adaptation into independent temporal and spatial branches via lightweight depthwise convolutions. In addition, to alleviate the long-tailed class distribution inherent in the benchmark dataset, we introduce an Adaptive Soft Balanced Augmentation method, which dynamically allocates augmentation intensity based on class rarity and learning difficulty, without manual thresholds. Our method achieves an F1 score of 0.43808, ranking 1st in Track 2 of the 4th EI-MiGA-IJCAI Challenge.

cs.CV

Selective Forgetting for Large Reasoning Models

Large Reasoning Models (LRMs) generate structured chains of thought (CoTs) before producing final answers, making them especially vulnerable to knowledge leakage through intermediate reasoning steps. Yet, the memorization of sensitive information in the training data such as copyrighted and private content has led to ethical and legal concerns. To address these issues, selective forgetting (also known as machine unlearning) has emerged as a potential remedy for LRMs. However, existing unlearning methods primarily target final answers and may degrade the overall reasoning ability of LRMs after forgetting. Additionally, directly applying unlearning on the entire CoTs could degrade the general reasoning capabilities. The key challenge for LRM unlearning lies in achieving precise unlearning of targeted knowledge while preserving the integrity of general reasoning capabilities. To bridge this gap, we in this paper propose a novel LRM unlearning framework that selectively removes sensitive reasoning components while preserving general reasoning capabilities. Our approach leverages multiple LLMs with retrieval-augmented generation (RAG) to analyze CoT traces, identify forget-relevant segments, and replace them with benign placeholders that maintain logical structure. We also introduce a new feature replacement unlearning loss for LRMs, which can simultaneously suppress the probability of generating forgotten content while reinforcing structurally valid replacements. Extensive experiments on both synthetic and medical datasets verify the desired properties of our proposed method.

cs.AI

FreqPhys: Repurposing Implicit Physiological Frequency Prior for Robust Remote Photoplethysmography

Remote photoplethysmography (rPPG) enables contactless physiological monitoring by capturing subtle skin-color variations from facial videos. However, most existing methods predominantly rely on time-domain modeling, making them vulnerable to motion artifacts and illumination fluctuations, where weak physiological clues are easily overwhelmed by noise. To address these challenges, we propose FreqPhys, a frequency-guided rPPG framework that explicitly leverages physiological frequency priors for robust signal recovery. Specifically, FreqPhys first applies a Physiological Bandpass Filtering module to suppress out-of-band interference, and then performs Physiological Spectrum Modulation together with adaptive spectral selection to emphasize pulse-related frequency components while suppress residual in-band noise. A Cross-domain Representation Learning module further fuses these spectral priors with deep time-domain features to capture informative spatial--temporal dependencies. Finally, a frequency-aware conditional diffusion process progressively reconstructs high-fidelity rPPG signals. Extensive experiments on six benchmarks demonstrate that FreqPhys yields significant improvements over state-of-the-art approaches, particularly under challenging motion conditions. It highlights the importance of explicitly modeling physiological frequency priors. The source code will be released.

cs.CV

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly frame-level labeling. To further tackle the challenges of imprecise sentiment boundaries in P-WTSL, we propose the Face-guided Sentiment Boundary Enhancement Network (\textbf{FSENet}), a unified framework that leverages fine-grained facial features to guide sentiment localization. Specifically, our approach \textit{first} introduces the Face-guided Sentiment Discovery (FSD) module, which integrates facial features into multimodal interaction via dual-branch modeling for effective sentiment stimuli clues; We \textit{then} propose the Point-aware Sentiment Semantics Contrast (PSSC) strategy to discriminate sentiment semantics of candidate points (frame-level) near annotation points via contrastive learning, thereby enhancing the model's ability to recognize sentiment boundaries. At \textit{last}, we design the Boundary-aware Sentiment Pseudo-label Generation (BSPG) approach to convert sparse point annotations into temporally smooth supervisory pseudo-labels. Extensive experiments and visualizations on the benchmark demonstrate the effectiveness of our framework, achieving state-of-the-art performance under full supervision, video-level, and point-level weak supervision, thereby showcasing the strong generalization ability of our FSENet across different annotation settings.

cs.CV

Coupling Brownian loop soups and random walk loop soups at all polynomial scales

Lawler and Trujillo Ferreras constructed a well-known coupling between the Brownian loop soups on $\mathbb{R}^2$ and the (discrete-time) random walk loop soups on $\mathbb{Z}^2$ (one rescales the random walk loops by $1/N$, their time parametrizations by $1/(2N^2)$, and lets $N\to \infty$), which led to numerous applications. It nevertheless only holds for loops with time length at least $N^{\theta-2}$ for $\theta \in(2/3,2)$. In particular, there is no control on mesoscopic loops with time length less than $N^{-4/3}$ (i.e. roughly diameter less than $N^{-2/3}$). This coupling was subsequently extended by Sapozhnikov and Shiraishi to $\mathbb{Z}^d$ with $d\ge 3$, for loops with time length at least $N^{\theta-2}$, for $\theta \in(2d/(d+4),2)$. In this paper, we find a simple way to remove the restriction $\theta>2d/(d+4)$, so that such a coupling works for all $\theta\in (0,2)$, i.e. for loops at all polynomial scales. We establish couplings for both discrete-time and continuous-time random walk loop soups on $\mathbb{Z}^d$, for $d\ge 1$. As an intermediate step, we also establish a KMT coupling between the continuous-time random walk bridge on $\mathbb{Z}^d$ and the Brownian bridge on $\mathbb{R}^d$.

math.PR

Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models

The rapid advancements in artificial intelligence (AI) have primarily focused on the process of learning from data to acquire knowledgeable learning systems. As these systems are increasingly deployed in critical areas, ensuring their privacy and alignment with human values is paramount. Recently, selective forgetting (also known as machine unlearning) has shown promise for privacy and data removal tasks, and has emerged as a transformative paradigm shift in the field of AI. It refers to the ability of a model to selectively erase the influence of previously seen data, which is especially important for compliance with modern data protection regulations and for aligning models with human values. Despite its promise, selective forgetting raises significant privacy concerns, especially when the data involved come from sensitive domains. While new unlearning-induced privacy attacks are continuously proposed, each is shown to outperform its predecessors using different experimental settings, which can lead to overly optimistic and potentially unfair assessments that may disproportionately favor one particular attack over the others. In this work, we present the first comprehensive benchmark for evaluating privacy vulnerabilities in selective forgetting. We extensively investigate privacy vulnerabilities of machine unlearning techniques and benchmark privacy leakage across a wide range of victim data, state-of-the-art unlearning privacy attacks, unlearning methods, and model architectures. We systematically evaluate and identify critical factors related to unlearning-induced privacy leakage. With our novel insights, we aim to provide a standardized tool for practitioners seeking to deploy customized unlearning applications with faithful privacy assessments.

cs.LG

Personalized Pricing in Social Networks with Individual and Group Fairness Considerations

Personalized pricing assigns different prices to customers for the same product based on customer-specific features to improve retailer revenue. However, this practice often raises concerns about fairness at both the individual and group levels. At the individual level, a customer may perceive unfair treatment if he/she notices being charged a higher price than others. At the group level, pricing disparities can result in discrimination against certain protected groups, such as those defined by gender or race. Existing studies on fair pricing typically address individual and group fairness separately. This paper bridges the gap by introducing a new formulation of the personalized pricing problem that incorporates both dimensions of fairness in social network settings. To solve the problem, we propose FairPricing, a novel framework based on graph neural networks (GNNs) that learns a personalized pricing policy using customer features and network topology. In FairPricing, individual perceived unfairness is captured through a penalty on customer demand, and thus the profit objective, while group-level discrimination is mitigated using adversarial debiasing and a price regularization term. Unlike existing optimization-based personalized pricing, which requires re-optimization whenever the network updates, the pricing policy learned by FairPricing assigns personalized prices to all customers in an updated network based on their features and the new network structure, thereby generalizing to network changes. Extensive experimental results show that FairPricing achieves high profitability while improving individual fairness perceptions and satisfying group fairness requirements.

cs.CY

Arm events in critical planar loop soups

We establish up-to-constants estimates for arm events in the Brownian loop soup on the 2D metric graph associated with the square lattice. More specifically, we consider two natural geometric events: first, ``bulk'' four-arm events, corresponding to two large connected components of loops getting close to each other; and then, two-arm events in the half-plane, used to estimate the probability that a cluster of loops approaches the boundary. Our proof relies on an estimate by Lupu-Werner [Probab. Theory Related Fields 171(3):775-818, 2018], thanks to the well-known coupling between the loop soup and the Gaussian free field on the metric graph [Lecture Notes in Mathematics, volume 2026, 2011] and [Ann. Probab. 44(3):2117-2146, 2016]. As a consequence, we also obtain up-to-constant upper bounds for the corresponding arm events in the random walk loop soup on the square lattice. In this way, we verify Assumptions 5.7 and 5.11 in arXiv:2409.16273: in a box with side length $N$, this implies the existence of crossings where the Gaussian free field remains below $a\sqrt{\log \log N}$ in absolute value, for some constant $a > 0$ large enough.

math.PR

Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be Forgotten

Currently, various uncertainty quantification methods have been proposed to provide certainty and probability estimates for deep learning models' label predictions. Meanwhile, with the growing demand for the right to be forgotten, machine unlearning has been extensively studied as a means to remove the impact of requested sensitive data from a pre-trained model without retraining the model from scratch. However, the vulnerabilities of such generated predictive uncertainties with regard to dedicated malicious unlearning attacks remain unexplored. To bridge this gap, for the first time, we propose a new class of malicious unlearning attacks against predictive uncertainties, where the adversary aims to cause the desired manipulations of specific predictive uncertainty results. We also design novel optimization frameworks for our attacks and conduct extensive experiments, including black-box scenarios. Notably, our extensive experiments show that our attacks are more effective in manipulating predictive uncertainties than traditional attacks that focus on label misclassifications, and existing defenses against conventional attacks are ineffective against our attacks.

cs.LG

Membership Inference Attacks with False Discovery Rate Control

Recent studies have shown that deep learning models are vulnerable to membership inference attacks (MIAs), which aim to infer whether a data record was used to train a target model or not. To analyze and study these vulnerabilities, various MIA methods have been proposed. Despite the significance and popularity of MIAs, existing works on MIAs are limited in providing guarantees on the false discovery rate (FDR), which refers to the expected proportion of false discoveries among the identified positive discoveries. However, it is very challenging to ensure the false discovery rate guarantees, because the underlying distribution is usually unknown, and the estimated non-member probabilities often exhibit interdependence. To tackle the above challenges, in this paper, we design a novel membership inference attack method, which can provide the guarantees on the false discovery rate. Additionally, we show that our method can also provide the marginal probability guarantee on labeling true non-member data as member data. Notably, our method can work as a wrapper that can be seamlessly integrated with existing MIA methods in a post-hoc manner, while also providing the FDR control. We perform the theoretical analysis for our method. Extensive experiments in various settings (e.g., the black-box setting and the lifelong learning setting) are also conducted to verify the desirable performance of our method.

stat.ML

Non-existence of several random fractals in Brownian motion and Brownian loop soup

We develop a unified approach to establish the non-existence of three types of random fractals: (1) the pioneer triple points of the planar Brownian motion, answering an open question in [7], (2) the pioneer double cut points of the planar and three-dimensional Brownian motions, and (3) the double points on the boundaries of the clusters of the planar Brownian loop soup at the critical intensity, answering an open question in [39]. These fractals have the common feature that they are associated with an intersection or disconnection exponent which yields a Hausdorff dimension ``exactly zero''.

math.PR

A Self-training Framework for Semi-supervised Pulmonary Vessel Segmentation and Its Application in COPD

Background: It is fundamental for accurate segmentation and quantification of the pulmonary vessel, particularly smaller vessels, from computed tomography (CT) images in chronic obstructive pulmonary disease (COPD) patients. Objective: The aim of this study was to segment the pulmonary vasculature using a semi-supervised method. Methods: In this study, a self-training framework is proposed by leveraging a teacher-student model for the segmentation of pulmonary vessels. First, the high-quality annotations are acquired in the in-house data by an interactive way. Then, the model is trained in the semi-supervised way. A fully supervised model is trained on a small set of labeled CT images, yielding the teacher model. Following this, the teacher model is used to generate pseudo-labels for the unlabeled CT images, from which reliable ones are selected based on a certain strategy. The training of the student model involves these reliable pseudo-labels. This training process is iteratively repeated until an optimal performance is achieved. Results: Extensive experiments are performed on non-enhanced CT scans of 125 COPD patients. Quantitative and qualitative analyses demonstrate that the proposed method, Semi2, significantly improves the precision of vessel segmentation by 2.3%, achieving a precision of 90.3%. Further, quantitative analysis is conducted in the pulmonary vessel of COPD, providing insights into the differences in the pulmonary vessel across different severity of the disease. Conclusion: The proposed method can not only improve the performance of pulmonary vascular segmentation, but can also be applied in COPD analysis. The code will be made available at https://github.com/wuyanan513/semi-supervised-learning-for-vessel-segmentation.

eess.IV

Up-to-constants estimates on four-arm events for simple conformal loop ensemble

We prove up-to-constants estimates for a general class of four-arm events in simple conformal loop ensembles, i.e. CLE$_\kappa$ for $\kappa\in (8/3,4]$. The four-arm events that we consider can be created by either one or two loops, with no constraint on the topology of the crossings. Our result is a key input in our series of works arxiv:2409.16230 and arxiv:2409.16273 on percolation of the two-sided level sets in the discrete Gaussian free field (and level sets in the occupation field of the random walk loop soup). In order to get rid of all constraints on the topology of the crossings, we rely on the Brownian loop-soup representation of simple CLE [Ann. Math. 176 (2012) 1827-1917], and a "cluster version" of a separation lemma for the Brownian loop soup. As a corollary, we also obtain up-to-constants estimates for a general version of four-arm events for SLE$_\kappa$ for $\kappa\in (8/3,4]$. This fixes (in the case of four arms and $\kappa\in(8/3,4]$) an essential gap in [Ann. Probab. 46 (2018) 2863-2907] and improves some estimates therein.

math.PR

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporally pinpoint all events simultaneously occurring in both audio and visual streams. Typically, each video encompasses dense events of multiple classes, which may overlap on the timeline, each exhibiting varied durations. Given these challenges, effectively exploiting the audio-visual relations and the temporal features encoded at various granularities becomes crucial. To address these challenges, we introduce a novel CCNet, comprising two core modules: the Cross-Modal Consistency Collaboration (CMCC) and the Multi-Temporal Granularity Collaboration (MTGC). Specifically, the CMCC module contains two branches: a cross-modal interaction branch and a temporal consistency-gated branch. The former branch facilitates the aggregation of consistent event semantics across modalities through the encoding of audio-visual relations, while the latter branch guides one modality's focus to pivotal event-relevant temporal areas as discerned in the other modality. The MTGC module includes a coarse-to-fine collaboration block and a fine-to-coarse collaboration block, providing bidirectional support among coarse- and fine-grained temporal features. Extensive experiments on the UnAV-100 dataset validate our module design, resulting in a new state-of-the-art performance in dense audio-visual event localization. The code is available at https://github.com/zzhhfut/CCNet-AAAI2025.

cs.CV

Backbone exponent and annulus crossing probability for planar percolation

We report the recent derivation of the backbone exponent for 2D percolation. In contrast to previously known exactly solved percolation exponents, the backbone exponent is a transcendental number, which is a root of an elementary equation. We also report an exact formula for the probability that there are two disjoint paths of the same color crossing an annulus. The backbone exponent captures the leading asymptotic, while the other roots of the elementary equation capture the asymptotic of the remaining terms. This suggests that the backbone exponent is part of a conformal field theory (CFT) whose bulk spectrum contains this set of roots. Our approach is based on the coupling between SLE curves and Liouville quantum gravity (LQG), and the integrability of Liouville CFT that governs the LQG surfaces.

cond-mat.stat-mech

Percolation of discrete GFF in dimension two I. Arm events in the random walk loop soup

In this work, which is the first part of a series of two papers, we study the random walk loop soup in dimension two. More specifically, we estimate the probability that two large connected components of loops come close to each other, in the subcritical and critical regimes. The associated four-arm event can be estimated in terms of exponents computed in the Brownian loop soup, relying on the connection between this continuous process and conformal loop ensembles (with parameter $\kappa \in (8/3,4]$). Along the way, we need to develop several useful tools for the loop soup, based on separation for random walks and surgery for loops, such as a "locality" property and quasi-multiplicativity. The results established here then play a key role in a second paper, in particular to study the connectivity properties of level sets in the random walk loop soup and in the discrete Gaussian free field.

math.PR