arXiv ScienceSearch

arXiv subjects

Qing Lu

Publications and source records attributed to Qing Lu.

At least 19 recordsLinked to original sources

Multimarker genetic association tests for panel count data

The existing multimarker survival tests focus on time to event outcomes. However, recurrent events are common in real world clinical and biomedical studies, especially in the research of chronic and recurrent diseases. In this paper, we develop a suite of set based genetic association tests for panel count outcomes under a unified weighted V statistic framework. These tests can effectively account for genetic effect heterogeneity in panel count data. Additionally, we develop small sample corrections to the tests to enhance the accuracy of the tests under small samples. Simulation studies show that the proposed tests perform well in terms of size and power across various scenarios and have a bigger power than the existing set based tests for interval censored outcomes. A dental caries GWAS data set is analyzed to illustrate the utility of the tests.

stat.ME

Genetic association testing with multivariate survival phenotypes under interval censoring

Set-based genetic association tests provide a powerful framework for detecting genetic effects on complex traits by jointly analyzing multiple genetic variants. Although set-based methods have been developed for interval-censored survival outcomes, existing approaches primarily focus on a single survival phenotype and therefore do not fully use information from multiple correlated outcomes. In this paper, we develop two Weighted V Tests for Multivariate Interval-Censored Data (WV-M-IC), extending the weighted V-statistic framework (Wu et al., 2021) to the joint analysis of multiple correlated interval-censored survival outcomes. The performance of these methods is evaluated through simulation studies, showing that the proposed approaches can provide power gains compared with single-outcome analyses. We apply the proposed methods to the ZOE 2.0 study to investigate dental caries progression in children.

stat.ME

Investigating the $P$-wave $DK^{*}$ Molecular Interpretation of the $D_{s1}(2700)$ via Its Strong Decays

In this work, we investigate the possibility of interpreting the $D_{s1}(2700)$ as a $P$-wave $DK^{*}$ hadronic molecular state through a systematic study of its strong decay properties. Using the effective Lagrangian approach, we calculate the strong decay widths of the $D_{s1}(2700)$ into the two-body final states $DK$, $D^{*}K$, $D_s\eta$, and $D_s^{*}\eta$, as well as the three-body $D\pi K$ channel. The coupling of the $D_{s1}(2700)$ to its constituents, $DK^{*}$, is constrained by the available experimental measurement of $R_{D_{s1}} =\Gamma(D_{s1}(2700) \to D^{*}K)/\Gamma(D_{s1}(2700)\to DK)=0.91\pm0.13\pm0.12$. With the resulting coupling, the total decay width can be readily obtained and is found to be in good agreement with the experimental measurement, providing strong support for interpreting the $D_{s1}(2700)$ as a $P$-wave $DK^{*}$ molecular state. The experimentally unobserved $D_{s1}(2700)\to D_s^{*}\eta$ decay channel provides a further test of this interpretation, as its predicted sizable decay width differs significantly from the conventional quark model prediction.

hep-ph

Searching for the $G(3900)$ via the $K^- p \to D_s^- \Lambda_c^+ G(3900)^0$ reaction

The nature of the $G(3900)$ structure, observed in $e^{+}e^{-}\to D\bar{D}$, remains unclear and may stem either from a genuine resonance or from charmonium interference and threshold effects. We therefore propose searching for the $G(3900)$ signal in the reaction $K^- p \to D_s^- \Lambda_c^+ G(3900)^0$, where the interference effects present in $e^{+}e^{-}\to \bar{D}^{*}D$ are absent. We employ an effective Lagrangian approach, where the reaction proceeds via a central production mechanism dominated by $t$-channel $D^{0}$ and $D^{*0}$ exchanges, based on the possible interpretation of $G(3900)$ as a $P$-wave $\bar{D}^{*}D$ molecular state, whose coupling to the $\bar{D}^{*}D$ channel is fixed from our previous fit to the $e^{+}e^{-}\to \bar{D}^{*}D$ data. The $\bar{K}N$ initial-state interaction, mediated by Pomeron and Reggeon exchanges, is also included and leads to a significant enhancement of the production cross section. If measured in future experiments, the predicted total cross sections and angular distributions can provide a promising probe of the nature of the $G(3900)$, and in particular of its possible genuine resonance nature.

hep-ph

A Quasi-Regression Method for the Mediation Analysis of Zero-Inflated Single-Cell Data

Recent advances in single-cell technologies have advanced our understanding of gene regulation and cellular heterogeneity at single-cell resolution. Single-cell data contain both gene expression levels and the proportion of expressing cells, which makes them structurally different from bulk data. Currently, methodological work on causal mediation analysis for single-cell data remains limited and often requires specific distributional assumptions. To address this challenge, we present QuasiMed, a mediation framework specialized for single-cell data. Our proposed method comprises three steps, including (i) screening mediator candidates through penalized regression and marginal models (similar to sure independence screening), (ii) estimation of indirect effects through the average expression and the proportion of expressing cells, (iii) and hypothesis testing with multiplicity control. The key benefit of QuasiMed is that it specifies only the mean functions of the mediation models through a quasi-regression framework, thereby relaxing strict distributional assumptions. The method performance was evaluated through the real-data-inspired simulations, and demonstrated high power, false discovery rate control, and computational efficiency. Lastly, we applied QuasiMed to ROSMAP single-cell data to illustrate its potential to identify mediating causal pathways. R package is freely available on GitHub repository at https://github.com/sjahnn/QuasiMed.

stat.ME

Alternating geometric progressions modulo one and Sturmian words

Let $b\ge 2$ be an integer. Using Sturmian words we describe all irrational real numbers $\xi$ such that the image in $\mathbb{R}/\mathbb{Z}$ of the sequence $(\xi (-b)^n)_{n\ge 0}$ is contained in an interval of length $b^{-1}+b^{-2}-b^{-3}$. In previous work (arXiv:2603.16794) we showed that the image cannot be contained in a shorter interval.

math.NT

Giant intrinsic dichroism in \b{eta}-Ga2O3 enables filter-free, high-fidelity polarization division multiplexing

Conventional polarization detection relies on external filters, which incur significant efficiency loss and polarization crosstalk, especially in the deep ultraviolet band where subwavelength nanofabrication is challenging. Here, we report that monoclinic \b{eta}-Ga2O3 exhibits intrinsic giant polarization dichroism, allowing near-ideal polarization photodetection without external optical elements and coherent polarization-division multiplexing (PDM) capability. The giant dichroism originates from the crystallographic symmetry-driven selectivity of optical transitions, which, combined with a large valence band splitting, results in vastly distinct absorption for orthogonal polarizations. A theoretical analysis of the transition selection rules in \b{eta}-Ga2O3 reveals only the E//c-polarized vb1-to-conduction band transition is activated, within the 245-258 nm spectral window. An admirable polarization ratio surpassing 500 and a polarization crosstalk ratio below 0.2% is hence achieved. The polarization-sensitive photodetector exhibits a high responsivity of 73 A/W and fast response (20 ms). Furthermore, we showcase its practical utility in PDM free-space communication, successfully decoding encoded optical signals, and demonstrate its capability for high-fidelity Stokes vector retrieval. The intrinsic anisotropy of \b{eta}-Ga2O3, dictated by its crystal symmetry, lays the groundwork for filter-free, high-fidelity polarization polarimetry. This work further paves the way for a general design principle in next-generation optoelectronics that harness polarization transition selection rules.

physics.app-ph

Fractional parts of powers of negative rationals

We prove that for any real number $\xi\neq 0$ and any coprime integers $p>q\ge1$ such that $\xi$ is irrational or $q>1$, the image in $\mathbb{R}/\mathbb{Z}$ of the sequence $(\xi (-p/q)^n)_{n\ge 0}$ is not contained in any interval of length less than $(1+q/p-q^2/p^2)/p$.

math.NT

Interpretation of $\Upsilon(11020)$ as an $S$-Wave $B_1\bar{B}$--$B_1\bar{B}^*$ Molecular State

Although heavy-quark symmetry predicts a $B_1\bar{B}$ molecular partner of the $D_1\bar{D}$ molecule, no such state has been observed. We propose that the experimentally observed $\Upsilon(11020)$ may be a candidate for such a state, possibly containing a $B_1\bar{B}^{*}$ component. To test this, we interpret $\Upsilon(11020)$ as an $S$-wave $B_1\bar{B}$--$B_1\bar{B}^{*}$ molecule and compute its strong decay widths using the compositeness condition and effective Lagrangians. The couplings to $B_1$ and $\bar{B}^{(*)}$ are extracted by fitting $\Upsilon(11020)\to e^+ e^-$ and $\Upsilon(11020)\to \chi_{bJ} \pi\pi\pi$ data. Using these couplings, we evaluate partial widths into $B^{(*)}_{(s)}\bar{B}^{(*)}_{(s)}$, $\pi\pi \Upsilon(nS)$, $\pi\pi h_b(nP)$, and $\pi\pi\pi \chi_{b1}$ via hadronic loops, as well as three-body $B^{*}\pi \bar{B}^{(*)}$ decays via tree diagrams. The results indicate that $\Upsilon(11020)$ is predominantly a $B_1\bar{B}$ molecule, with its main decay channel being $B_s^{*}\bar{B}^{*}$. The $\pi\pi \Upsilon(nS)$ and $\pi\pi h_b(nP)$ widths are only a few eV, whereas $\pi\pi\pi \chi_{b1}$ reaches 0.167~MeV and the unobserved $\pi\pi\pi \chi_{b0}$ could be 0.754~keV. These distinctive decay patterns provide clear experimental signatures of the molecular nature of $\Upsilon(11020)$ and offer a test of heavy-quark symmetry.

hep-ph

Enabling Ultra-Fast Cardiovascular Imaging Across Heterogeneous Clinical Environments with A Generalist Foundation Model and Multimodal Database

Multimodal cardiovascular magnetic resonance (CMR) imaging provides comprehensive and non-invasive insights into cardiovascular disease (CVD) diagnosis and underlying mechanisms. Despite decades of advancements, its widespread clinical adoption remains constrained by prolonged scan times, inconsistent image quality, and heterogeneity across medical environments. This underscores the urgent need for a generalist reconstruction foundation model for ultra-fast CMR imaging, one formulated for physics-constrained inverse problems in the sensor (k-space) domain, capable of adapting across diverse imaging scenarios and serving as the essential substrate for all downstream analyses. To enable this goal, we curate MMCMR-427K, the largest and most comprehensive multimodal CMR k-space database to date, comprising 427,465 multi-coil k-space data paired with structured metadata across 13 international centers, 12 CMR modalities, 15 scanners spanning four field strengths, and 17 CVD categories in populations across three continents. Building on this unprecedented resource, we introduce CardioMM, a generalist reconstruction foundation model capable of dynamically adapting to heterogeneous fast CMR imaging scenarios. CardioMM unifies semantic contextual understanding with physics-informed data consistency to deliver robust reconstructions across varied scanners, protocols, and patient presentations. Comprehensive evaluations demonstrate that CardioMM achieves state-of-the-art performance across internal centers and exhibits strong zero-shot generalization to unseen external settings. Importantly, CardioMM supports acceleration up to 24x, providing the first evidence that such extreme acquisition speed can preserve key cardiac phenotypes, quantitative myocardial biomarkers, and diagnostic image quality without compromising clinical integrity.

eess.IV

DiffKnock: Diffusion-based Knockoff Statistics for Neural Networks Inference

We introduce DiffKnock, a diffusion-based knockoff framework for high-dimensional feature selection with finite-sample false discovery rate (FDR) control. DiffKnock addresses two key limitations of existing knockoff methods: preserving complex feature dependencies and detecting non-linear associations. Our approach trains diffusion models to generate valid knockoffs and uses neural network--based gradient and filter statistics to construct antisymmetric feature importance measures. Through simulations, we showed that DiffKnock achieved higher power than autoencoder-based knockoffs while maintaining target FDR, indicating its superior performance in scenarios involving complex non-linear architectures. Applied to murine single-cell RNA-seq data of LPS-stimulated macrophages, DiffKnock identifies canonical NF-$\kappa$B target genes (Ccl3, Hmox1) and regulators (Fosb, Pdgfb). These results highlight that, by combining the flexibility of deep generative models with rigorous statistical guarantees, DiffKnock is a powerful and reliable tool for analyzing single-cell RNA-seq data, as well as high-dimensional and structured data in other domains.

stat.ME

Neural Tangent Kernels for Complex Genetic Risk Prediction: Bridging Deep Learning and Kernel Methods in Genomics

Given the complexity of genetic risk prediction, there is a critical need for the development of novel methodologies that can effectively capture intricate genotype--phenotype relationships (e.g., nonlinear) while remaining statistically interpretable and computationally tractable. We develop a Neural Tangent Kernel (NTK) framework to integrate kernel methods into deep neural networks for genetic risk prediction analysis. We consider two approaches: NTK-LMM, which embeds the empirical NTK in a linear mixed model with variance components estimated via minimum quadratic unbiased estimator (MINQUE), and NTK-KRR, which performs kernel ridge regression with cross-validated regularization. Through simulation studies, we show that NTK-based models outperform the traditional neural network models and linear mixed models. By applying NTK to endophenotypes (e.g., hippocampal volume) and AD-related genes (e.g., APOE) from Alzheimer's Disease Neuroimaging Initiative (ADNI), we found that NTK achieved higher accuracy than existing methods for hippocampal volume and entorhinal cortex thickness. In addition to its accuracy performance, NTK has favorable optimization properties (i.e., having a closed-form or convex training) and generates interpretable results due to its connection to variance components and heritability. Overall, our results indicate that by integrating the strengths of both deep neural networks and kernel methods, NTK offers competitive performance for genetic risk prediction analysis while having the advantages of interpretability and computational efficiency.

stat.AP

Collapsing ROC approach for risk prediction research on both common and rare variants

Risk prediction that capitalizes on emerging genetic findings holds great promise for improving public health and clinical care. However, recent risk prediction research has shown that predictive tests formed on existing common genetic loci, including those from genome-wide association studies, have lacked sufficient accuracy for clinical use. Because most rare variants on the genome have not yet been studied for their role in risk prediction, future disease prediction discoveries should shift toward a more comprehensive risk prediction strategy that takes into account both common and rare variants. We are proposing a collapsing receiver operating characteristic CROC approach for risk prediction research on both common and rare variants. The new approach is an extension of a previously developed forward ROC FROC approach, with additional procedures for handling rare variants. The approach was evaluated through the use of 533 single-nucleotide polymorphisms SNPs in 37 candidate genes from the Genetic Analysis Workshop 17 mini-exome data set. We found that a prediction model built on all SNPs gained more accuracy AUC = 0.605 than one built on common variants alone AUC = 0.585. We further evaluated the performance of two approaches by gradually reducing the number of common variants in the analysis. We found that the CROC method attained more accuracy than the FROC method when the number of common variants in the data decreased. In an extreme scenario, when there are only rare variants in the data, the CROC reached an AUC value of 0.603, whereas the FROC had an AUC value of 0.524.

cs.LG

A U-Statistic-based random forest approach for genetic interaction study

Variations in complex traits are influenced by multiple genetic variants, environmental risk factors, and their interactions. Though substantial progress has been made in identifying single genetic variants associated with complex traits, detecting the gene-gene and gene-environment interactions remains a great challenge. When a large number of genetic variants and environmental risk factors are involved, searching for interactions is limited to pair-wise interactions due to the exponentially increased feature space and computational intensity. Alternatively, recursive partitioning approaches, such as random forests, have gained popularity in high-dimensional genetic association studies. In this article, we propose a U-Statistic-based random forest approach, referred to as Forest U-Test, for genetic association studies with quantitative traits. Through simulation studies, we showed that the Forest U-Test outperformed existing methods. The proposed method was also applied to study Cannabis Dependence CD, using three independent datasets from the Study of Addiction: Genetics and Environment. A significant joint association was detected with an empirical p-value less than 0.001. The finding was also replicated in two independent datasets with p-values of 5.93e-19 and 4.70e-17, respectively.

q-bio.GN

A Generalized Genetic Random Field Method for the Genetic Association Analysis of Sequencing Data

With the advance of high-throughput sequencing technologies, it has become feasible to investigate the influence of the entire spectrum of sequencing variations on complex human diseases. Although association studies utilizing the new sequencing technologies hold great promise to unravel novel genetic variants, especially rare genetic variants that contribute to human diseases, the statistical analysis of high-dimensional sequencing data remains a challenge. Advanced analytical methods are in great need to facilitate high-dimensional sequencing data analyses. In this article, we propose a generalized genetic random field (GGRF) method for association analyses of sequencing data. Like other similarity-based methods (e.g., SIMreg and SKAT), the new method has the advantages of avoiding the need to specify thresholds for rare variants and allowing for testing multiple variants acting in different directions and magnitude of effects. The method is built on the generalized estimating equation framework and thus accommodates a variety of disease phenotypes (e.g., quantitative and binary phenotypes). Moreover, it has a nice asymptotic property, and can be applied to small-scale sequencing data without need for small-sample adjustment. Through simulations, we demonstrate that the proposed GGRF attains an improved or comparable power over a commonly used method, SKAT, under various disease scenarios, especially when rare variants play a significant role in disease etiology. We further illustrate GGRF with an application to a real dataset from the Dallas Heart Study. By using GGRF, we were able to detect the association of two candidate genes, ANGPTL3 and ANGPTL4, with serum triglyceride.

stat.ME

Functional Analysis of Variance for Association Studies

While progress has been made in identifying common genetic variants associated with human diseases, for most of common complex diseases, the identified genetic variants only account for a small proportion of heritability. Challenges remain in finding additional unknown genetic variants predisposing to complex diseases. With the advance in next-generation sequencing technologies, sequencing studies have become commonplace in genetic research. The ongoing exome-sequencing and whole-genome-sequencing studies generate a massive amount of sequencing variants and allow researchers to comprehensively investigate their role in human diseases. The discovery of new disease-associated variants can be enhanced by utilizing powerful and computationally efficient statistical methods. In this paper, we propose a functional analysis of variance (FANOVA) method for testing an association of sequence variants in a genomic region with a qualitative trait. The FANOVA has a number of advantages: (1) it tests for a joint effect of gene variants, including both common and rare; (2) it fully utilizes linkage disequilibrium and genetic position information; and (3) allows for either protective or risk-increasing causal variants. Through simulations, we show that FANOVA outperform two popularly used methods - SKAT and a previously proposed method based on functional linear models (FLM), - especially if a sample size of a study is small and/or sequence variants have low to moderate effects. We conduct an empirical study by applying three methods (FANOVA, SKAT and FLM) to sequencing data from Dallas Heart Study. While SKAT and FLM respectively detected ANGPTL 4 and ANGPTL 3 associated with obesity, FANOVA was able to identify both genes associated with obesity.

stat.AP

Hodge-Riemann polynomials

We show that Schur classes of ample vector bundles on smooth projective varieties satisfy Hodge-Riemann relations on $H^{p,q}$ under the assumption that $H^{p-2,q-2}$ vanishes. More generally, we study Hodge-Riemann polynomials, which are partially symmetric polynomials that produce cohomology classes satisfying the Hodge-Riemann property when evaluated at Chern roots of ample vector bundles. In the case of line bundles and in bidegree $(1,1)$, these are precisely the nonzero dually Lorentzian polynomials. We prove various properties of Hodge-Riemann polynomials, confirming predictions and answering questions of Ross and Toma. As an application, we show that the derivative sequence of any product of Schur polynomials is Schur log-concave, confirming conjectures of Ross and Wu.

math.AG

A multi-locus predictiveness curve and its summary assessment for genetic risk prediction

With the advance of high-throughput genotyping and sequencing technologies, it becomes feasible to comprehensive evaluate the role of massive genetic predictors in disease prediction. There exists, therefore, a critical need for developing appropriate statistical measurements to access the combined effects of these genetic variants in disease prediction. Predictiveness curve is commonly used as a graphical tool to measure the predictive ability of a risk prediction model on a single continuous biomarker. Yet, for most complex diseases, risk prediciton models are formed on multiple genetic variants. We therefore propose a multi-marker predictiveness curve and provide a non-parametric method to construct the curve for case-control studies. We further introduce a global predictiveness U and a partial predictiveness U to summarize prediction curve across the whole population and sub-population of clinical interest, respectively. We also demonstrate the connections of predictiveness curve with ROC curve and Lorenz curve. Through simulation, we compared the performance of the predictiveness U to other three summary indices: R square, Total Gain, and Average Entropy, and showed that Predictiveness U outperformed the other three indexes in terms of unbiasedness and robustness. Moreover, we simulated a series of rare-variants disease model, found partial predictiveness U performed better than global predictiveness U. Finally, we conducted a real data analysis, using predictiveness curve and predictiveness U to evaluate a risk prediction model for Nicotine Dependence.

stat.ME