arXiv ScienceSearch

arXiv subjects

Chun Li

Publications and source records attributed to Chun Li.

At least 19 recordsLinked to original sources

Doubly Robust Estimators of Quantile Treatment Effects With Semiparametric Cumulative Probability Models

The causal inference literature has traditionally focused on estimating the mean of the potential outcome, whereas evaluating how a treatment affects the entire outcome distribution can provide additional information in biomedical research. Quantile treatment effect (QTE) captures such distributional differences, particularly when outcomes are skewed. However, existing approaches for estimating QTE make distributional assumptions about the outcome and are thus sensitive to model misspecification. Motivated by an HIV study with skewed outcomes, one of which is subject to detection limits, we propose a doubly robust framework for estimating QTE based on the cumulative probability model (CPM), which is a rank-based, semiparametric linear transformation model. We develop two CPM-based estimation strategies: (1) an inverse-cumulative distribution function (CDF) approach that first estimates the marginal CDF of potential outcomes using the efficient influence function (EIF) and then obtains marginal quantiles via weighted quantile interpolation by inverting the distribution, and (2) a direct approach that solves the EIF of potential marginal quantiles. The proposed estimators are doubly robust and asymptotically normal. We further extend the framework to probability treatment effects (PTEs) and their conditional counterparts. For statistical inference, we investigate several variance estimation procedures, including EIF-based estimators, sandwich estimators, and the nonparametric bootstrap. Simulation studies illustrate that the empirical sandwich estimator and the nonparametric bootstrap provide doubly robust variance estimation with stable finite-sample performance under nuisance model misspecification. The proposed methods are evaluated through extensive Monte Carlo simulations and illustrated using an HIV data application.

stat.ME

What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program

This paper presents the design and outcomes of a seven-weekend AI storytelling program developed for Black girls aged 10-12. Grounded in Afrofuturism and Black feminist thought, the program adopted AI-enabled counter-storytelling, supported the development of foundational AI literacies, and fostered future-oriented imagination. Activities included brainstorming AI-related topics, developing character and story plots, and delivering collaborative group presentations. Drawing on the analysis of learners' artifacts from the case study, findings show that participants created Afrofuturist narratives rooted in their identities and everyday experiences. At the same time, they developed core AI literacies, including prompt engineering, bias critique, and awareness of data privacy. This program demonstrates that integrating Afrofuturist storytelling with generative AI in informal learning spaces can be a powerful approach for engaging Black girls in computer science education.

cs.CY

Addressing errors in multiple variables using generalized raking and cumulative probability models

Routinely collected data, such as electronic health record (EHR) data, are frequently used for biomedical research, but these data are prone to errors, which can bias study findings. Validating data in subsamples of records can reduce bias, and the efficiency of estimates can be improved by incorporating in analyses both the error-prone data available on the entire cohort and the validated data available on the subsample. One approach to incorporate both data sources is with generalized raking, which calibrates validation sampling weights using error-prone data from the entire cohort. Motivated by an EHR study of maternal weight gain during pregnancy with a validation subsample, we develop and illustrate generalized raking techniques for cumulative probability models (CPMs). CPMs are robust, rank-based and semiparametric models for continuous, ordinal, or mixed type outcome data. We develop efficient generalized raking estimators for CPMs, evaluate their performance relative to competing methods, and demonstrate the utility and strengths of generalized raking with CPMs in a study that examines factors associated with weight gain during pregnancy.

stat.ME

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors

Existing approaches for unsupervised 3D point cloud segmentation predominantly rely on a purely visual similarity-based learning-by-clustering paradigm, which suffers from a fundamental limitation: long-tail ambiguity. In such a paradigm, features of minor classes are consistently absorbed by dominant clusters, leading to severely imbalanced predictions. To address this issue, we propose LangTail, a language-guided hierarchical learning framework that leverages the balanced world knowledge encoded in language models to mitigate long-tail ambiguity in unsupervised 3D segmentation. The key idea is to establish multi-level associations between language-derived semantic priors and visually underrepresented minor classes, thereby compensating for the biased attention of purely visual clustering toward dominant classes. Specifically, LangTail first constructs an entity-level semantic prior from language models, capturing balanced and fine-grained world knowledge across categories. These priors are injected into a hierarchical clustering framework via contrastive alignment. This guides multi-granularity semantic structure formation and prevents minor classes from being absorbed by dominant clusters, yielding more discriminative representations for underrepresented categories. Extensive experiments on ScanNet-v2, S3DIS, and nuScenes demonstrate that LangTail consistently outperforms existing methods by significant margins, \ie, +13.5, +12.9, and +8.9 mIoU, respectively. These results demonstrate the effectiveness of language priors in improving the representation of minority classes in 3D point clouds. The code will be released at: https://github.com/Whisky0129/langtail_official.

cs.CV

The Stellar Abundances and Galactic Evolution Survey (SAGES). V. The First Data Release of the DDO51 Band

We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The DDO51 filter is centered near the \ion{Mg}{1}~$b$ triplet and the adjacent MgH feature, offering sensitivity to stellar surface gravity. The data reduction pipeline incorporates an improved astrometric solution anchored to Gaia DR3 and a photometric calibration strategy tied to synthetic photometry from Gaia XP spectra. These procedures yield a point-source depth of $\sim$18.9 mag at S/N$\sim$10 and an internal photometric precision $\approx$6-7 mmag at the bright end. A preliminary color--color analysis using Gaia broadband photometry confirms the expected sensitivity of the DDO51 band to stellar surface gravity, demonstrating a clear photometric separation between dwarf and giant sequences for late-type stars. This dataset, when combined with existing SAGES photometry in other bands, provides a crucial tool for disentangling the substructures of the Milky Way. All data products from this release upon publication will be available.

astro-ph.SR

The Spectroscopic and Photometric Study of a Star Cluster Sample in Andromeda Halo

Halo star clusters serve as vital tracers for the formation and evolution of the Andromeda galaxy. In this work, we present physical parameters for 29 M31 halo star clusters, derived from a combination of spectroscopic and photometric data. Low-resolution spectra were acquired using the BFOSC spectrograph on the NAOC Xinglong 2.16-m telescope. For the photometric analysis, we utilized uSC and vSAGE bands from the SAGE survey, complemented by archival data from GALEX(NUV, FUV), PAN-STARRS(grizy) and the 2MASS(JHK). Ages and metallicities were determined via ULySS (Vazdekis et al. and pegase-hr) SSP model and the Bruzual & Charlot (2003) (BC03) stellar population synthesis models. The derived parameters show good agreement with literature values. Notably, for three of these clusters, this study represents the first combined photometric and spectroscopic analysis.

astro-ph.GA

QED corrections of orders $m\alpha^6$ and $m\alpha^6(m/M)$ for HD$^+$ rovibrational transitions beyond Born-Oppenheimer approximation

The effective Hamiltonian of $m\alpha^6$ and $m\alpha^6(m/M)$ order corrections for hydrogen molecular ions has been derived in [ Z.-X. Zhong, \emph{et al.}, Phys. Rev. A {\bf98}, 032502(2018).], in this work we express the energy correction in the form of finite-value effective operators. The cut-off regularization scheme is used to determine finite part of divergent operators of the leading-order recoil corrections. Numerical calculations of first-order contributions are performed in the Hylleraas basis set. Combining the second-order terms calculated in recent work [V. I. Korobov, \emph{et al.}, Mol. Phys. e2563023 (2025).], the $m\alpha^6$-order corrections for the fundamental rovibrational transition are obtained with an uncertainty three times smaller than in previous calculations.

physics.atom-ph

Mechanics of hierarchical twisted and coiled polymer artificial muscles: Decoupling force from kinematic limits

Thermally actuated twisted and coiled polymer (TCP) artificial muscles exhibit exceptional specific work capacities but are limited by an inherent competition between load-bearing capacity and actuation stroke. To address this limitation, we investigate a hierarchical helical structure designed to decouple force generation from kinematic limits. We propose a coupled thermo-mechanical model incorporating inter-filamentary contact mechanics and geometric nonlinearities to predict the assembly's equilibrium response. The results indicate that this hierarchical topology significantly amplifies isometric actuation stress compared to monofilament baselines, while maintaining a biological-like contraction stroke of approximately 22%. A critical topological threshold governed by the balance between cooperative load-sharing and geometric confinement is identified. Beyond an optimal bundle complexity, the geometric jamming dominates, as excessive inter-filamentary friction hinders actuation. Furthermore, we elucidate a stiffness-stroke synergy in homochiral configurations, where high helical angles amplify the thermal untwisting torque to overcome increased structural rigidity. Crucially, the volumetric energy density exhibits scale invariance regarding the hierarchical radius, implying that absolute force output can be linearly scaled through geometric upsizing without compromising efficiency. These findings provide a mechanics-based rationale for the structural programming, demonstrating that soft actuator performance limits are dictated by topological order rather than intrinsic material properties.

cond-mat.soft

ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents

Existing GUI agent models relying on coordinate-based one-step visual grounding struggle with generalizing to varying input resolutions and aspect ratios. Alternatives introduce coordinate-free strategies yet suffer from learning under severe data scarcity. To address the limitations, we propose ToolTok, a novel paradigm of multi-step pathfinding for GUI agents, where operations are modeled as a sequence of progressive tool usage. Specifically, we devise tools aligned with human interaction habits and represent each tool using learnable token embeddings. To enable efficient embedding learning under limited supervision, ToolTok introduces a semantic anchoring mechanism that grounds each tool with semantically related concepts as natural inductive bias. To further enable a pre-trained large language model to progressively acquire tool semantics, we construct an easy-to-hard curriculum consisting of three tasks: token definition question-answering, pure text-guided tool selection, and simplified visual pathfinding. Extensive experiments on multiple benchmarks show that ToolTok achieves superior performance among models of comparable scale (4B) and remains competitive with a substantially larger model (235B). Notably, these results are obtained using less than 1% of the training data required by other post-training approaches. In addition, ToolTok demonstrates strong generalization across unseen scenarios. Our training & inference code is open-source at https://github.com/ZephinueCode/ToolTok.

cs.LG

A DeepONet joint Neural Tangent Kernel Hybrid Framework for Physics-Informed Inverse Source Problems and Robust Image Reconstruction

This work presents a novel hybrid approach that integrates Deep Operator Networks (DeepONet) with the Neural Tangent Kernel (NTK) to solve complex inverse problem. The method effectively addresses tasks such as source localization governed by the Navier-Stokes equations and image reconstruction, overcoming challenges related to nonlinearity, sparsity, and noisy data. By incorporating physics-informed constraints and task-specific regularization into the loss function, the framework ensures solutions that are both physically consistent and accurate. Validation on diverse synthetic and real datasets demonstrates its robustness, scalability, and precision, showcasing its broad potential applications in computational physics and imaging sciences.

cs.CV

Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities

Multimodal Emotion Recognition in Conversations (MERC) enhances emotional understanding through the fusion of multimodal signals. However, unpredictable modality absence in real-world scenarios significantly degrades the performance of existing methods. Conventional missing-modality recovery approaches, which depend on training with complete multimodal data, often suffer from semantic distortion under extreme data distributions, such as fixed-modality absence. To address this, we propose the Federated Dialogue-guided and Semantic-Consistent Diffusion (FedDISC) framework, pioneering the integration of federated learning into missing-modality recovery. By federated aggregation of modality-specific diffusion models trained on clients and broadcasting them to clients missing corresponding modalities, FedDISC overcomes single-client reliance on modality completeness. Additionally, the DISC-Diffusion module ensures consistency in context, speaker identity, and semantics between recovered and available modalities, using a Dialogue Graph Network to capture conversational dependencies and a Semantic Conditioning Network to enforce semantic alignment. We further introduce a novel Alternating Frozen Aggregation strategy, which cyclically freezes recovery and classifier modules to facilitate collaborative optimization. Extensive experiments on the IEMOCAP, CMUMOSI, and CMUMOSEI datasets demonstrate that FedDISC achieves superior emotion classification performance across diverse missing modality patterns, outperforming existing approaches.

cs.CV

Diffusion Low Rank Hybrid Reconstruction for Sparse View Medical Imaging

This work presents TV-LoRA, a novel method for low-dose sparse-view CT reconstruction that combines a diffusion generative prior (NCSN++ with SDE modeling) and multi-regularization constraints, including anisotropic TV and nuclear norm (LoRA), within an ADMM framework. To address ill-posedness and texture loss under extremely sparse views, TV-LoRA integrates generative and physical constraints, and utilizes a 2D slice-based strategy with FFT acceleration and tensor-parallel optimization for efficient inference. Experiments on AAPM-2016, CTHD, and LIDC datasets with $N_{\mathrm{view}}=8,4,2$ show that TV-LoRA consistently surpasses benchmarks in SSIM, texture recovery, edge clarity, and artifact suppression, demonstrating strong robustness and generalizability. Ablation studies confirm the complementary effects of LoRA regularization and diffusion priors, while the FFT-PCG module provides a speedup. Overall, Diffusion + TV-LoRA achieves high-fidelity, efficient 3D CT reconstruction and broad clinical applicability in low-dose, sparse-sampling scenarios.

cs.CV

The Stellar Abundances and Galactic Evolution Survey (SAGES). IV. Surface Gravity Estimation and Giant-Dwarf Separation with the DDO51 Filter

Reliable estimation of stellar surface gravity (log $g$) for a large sample is crucial for evaluating stellar evolution models and understanding galactic structure; However, it is not easy to accomplish due to the difficulty in gathering a large spectroscopic data set. Photometric sky survey using a specific filter, on the other hand, can play a substantial role in the assessment of log $g$. The Stellar Abundances and Galactic Evolution Survey (SAGES) utilizes eight filters to provide accurate stellar parameters for $\sim10^{7}$ stars, with its DDO51 intermediate-band filter specifically designed for robust log $g$ determination. In this work, the observed SAGES $u_{\rm SC}$ and $v_{\rm SAGES}$ photometry, the synthetic photometry in $g$, $r$, $i$, and DDO51 bands derived from \textit{Gaia} XP spectra are employed to investigate the importance of the DDO51 filter in the determination of log $g$. We applied machine-learning-based extinction correction and employed XGBoost models, trained on stellar parameters from LAMOST, to predict log $g$ using photometric data. By comparing model predicted log $g$ with LAMOST values, we find that including DDO51 filter improve the accuracies of log $g$ estimates by 21.0\% (from 0.224\,dex to 0.177\,dex) overall, and by 26.5\% (from 0.302\,dex to 0.222\,dex ) for GK-type stars, as compared to those obtained without DDO51. The DDO51 filter is also validated to be particularly effective for metal-poor stars ([Fe/H]$<$-1.0), where it significantly mitigates systematic biases. Our findings highlight the diagnostic power of the SAGES DDO51 filter, providing enhanced stellar characterization vital for future in-depth studies of the Milky Way.

astro-ph.SR

FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation

Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn limits their broader adoption. To address this, we propose FedVLA, the first federated VLA learning framework, enabling distributed model training that preserves data privacy without compromising performance. Our framework integrates task-aware representation learning, adaptive expert selection, and expert-driven federated aggregation, enabling efficient and privacy-preserving training of VLA models. Specifically, we introduce an Instruction Oriented Scene-Parsing mechanism, which decomposes and enhances object-level features based on task instructions, improving contextual understanding. To effectively learn diverse task patterns, we design a Dual Gating Mixture-of-Experts (DGMoE) mechanism, where not only input tokens but also self-aware experts adaptively decide their activation. Finally, we propose an Expert-Driven Aggregation strategy at the federated server, where model aggregation is guided by activated experts, ensuring effective cross-client knowledge transfer.Extensive simulations and real-world robotic experiments demonstrate the effectiveness of our proposals. Notably, DGMoE significantly improves computational efficiency compared to its vanilla counterpart, while FedVLA achieves task success rates comparable to centralized training, effectively preserving data privacy.

cs.RO

Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures

Existing multi-view classification and clustering methods typically improve task accuracy by leveraging and fusing information from different views. However, ensuring the reliability of multi-view integration and final decisions is crucial, particularly when dealing with noisy or corrupted data. Current methods often rely on Kullback-Leibler (KL) divergence to estimate uncertainty of network predictions, ignoring domain gaps between different modalities. To address this issue, KPHD-Net, based on H\"older divergence, is proposed for multi-view classification and clustering tasks. Generally, our KPHD-Net employs a variational Dirichlet distribution to represent class probability distributions, models evidences from different views, and then integrates it with Dempster-Shafer evidence theory (DST) to improve uncertainty estimation effects. Our theoretical analysis demonstrates that Proper H\"older divergence offers a more effective measure of distribution discrepancies, ensuring enhanced performance in multi-view learning. Moreover, Dempster-Shafer evidence theory, recognized for its superior performance in multi-view fusion tasks, is introduced and combined with the Kalman filter to provide future state estimations. This integration further enhances the reliability of the final fusion results. Extensive experiments show that the proposed KPHD-Net outperforms the current state-of-the-art methods in both classification and clustering tasks regarding accuracy, robustness, and reliability, with theoretical guarantees.

cs.CV

Robust Brain Tumor Segmentation with Incomplete MRI Modalities Using H\"older Divergence and Mutual Information-Enhanced Knowledge Transfer

Multimodal MRI provides critical complementary information for accurate brain tumor segmentation. However, conventional methods struggle when certain modalities are missing due to issues such as image quality, protocol inconsistencies, patient allergies, or financial constraints. To address this, we propose a robust single-modality parallel processing framework that achieves high segmentation accuracy even with incomplete modalities. Leveraging Holder divergence and mutual information, our model maintains modality-specific features while dynamically adjusting network parameters based on the available inputs. By using these divergence- and information-based loss functions, the framework effectively quantifies discrepancies between predictions and ground-truth labels, resulting in consistently accurate segmentation. Extensive evaluations on the BraTS 2018 and BraTS 2020 datasets demonstrate superior performance over existing methods in handling missing modalities.

cs.CV

Estimating treatment effects with a unified semi-parametric difference-in-differences approach

Difference-in-differences (DID) approaches are widely used for estimating causal effects with observational data before and after an intervention. DID traditionally estimates the average treatment effect among the treated after making a parallel trends assumption on the means of the outcome. With skewed outcomes, a transformation is often needed; however, the transformation may be difficult to choose, results may be sensitive to the choice, and parallel trends assumptions are made on the transformed scale. Recent DID methods estimate alternative treatment effects that may be preferable with skewed outcomes. However, each alternative DID estimator requires a different parallel trends assumption. We introduce a new DID method capable of estimating average, quantile, probability, and novel Mann-Whitney treatment effects among the treated with a single unifying parallel trends assumption. The proposed method uses a semi-parametric cumulative probability model (CPM). The CPM is a linear model for a latent variable on covariates, where the latent variable results from an unspecified transformation of the outcome. Our DID approach makes a universal parallel trends assumption on the expectation of the latent variable conditional on covariates. Hence, our method avoids specifying outcome transformations and does not require separate assumptions for each estimand. We introduce the method; describe identification, estimation, and inference; conduct simulations evaluating its performance; and apply it to assess the impact of Medicaid expansion on CD4 count among people with HIV.

stat.ME

DeepRAG: Integrating Hierarchical Reasoning and Process Supervision for Biomedical Multi-Hop QA

We propose DeepRAG, a novel framework that integrates DeepSeek hierarchical question decomposition capabilities with RAG Gym unified retrieval-augmented generation optimization using process level supervision. Targeting the challenging MedHopQA biomedical question answering task, DeepRAG systematically decomposes complex queries into precise sub-queries and employs concept level reward signals informed by the UMLS ontology to enhance biomedical accuracy. Preliminary evaluations on the MedHopQA dataset indicate that DeepRAG significantly outperforms baseline models, including standalone DeepSeek and RAG Gym, achieving notable improvements in both Exact Match and concept level accuracy.

cs.CL