arXiv ScienceSearch

arXiv subjects

Shujie Li

Publications and source records attributed to Shujie Li.

At least 19 recordsLinked to original sources

When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation

\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ in data distributions, adaptation dynamics, and representation learning progress, making a single split point insufficient to accommodate client-specific training states. In this paper, we propose \textsc{FedSGA}, a \textbf{S}ufficiency-\textbf{G}uided \textbf{A}daptive split \textbf{Fed}erated learning framework that addresses this question through client-specific shallow sufficiency estimation. First, we introduce a client-specific adaptation channel based on private prompt tokens, which tracks local adaptation dynamics separately from the shared backbone and provides a lightweight signal for detecting whether client adaptation remains active. To further avoid repeated online probing over multiple candidate depths, we design a shallow sufficiency estimator that combines cross-client semantic alignment, temporal interface stability, and prompt-state variation to estimate whether the shallowest split is already sufficient. Finally, we introduce a split-compatible interface harmonization module that projects activations from different split depths into a shared semantic space, improving the comparability of heterogeneous client interfaces before server-side prediction. Extensive experiments on multiple heterogeneous benchmarks demonstrate the effectiveness of \textsc{FedSGA} in improving model performance compared with state-of-the-art methods while reducing unnecessary client-side computation.

cs.DC

One Graph, Multiple Gains: Single High-Quality Item-Item Graph for Multimodal Recommendation

Multimodal recommendation leverages item multimodal features alongside collaborative signals to capture user preferences. While item-item graphs have become a key building block in advanced models, existing methods typically construct them with noisy similarity edges and limit their role to a single function of item-item representation propagation, leaving substantial potential untapped. In this paper, we propose IIMRec, a framework that constructs a single high-quality item-item graph during preprocessing and systematically reuses it across three stages of the recommendation pipeline: representation enhancement, interaction graph enhancement, and optimization enhancement. The graph is built by fusing semantic and co-occurrence signals, then refined via Neighborhood Consistency Edge Reweighting (NCER), which applies the triadic closure principle to amplify structurally reliable edges and suppress spurious ones. Once constructed, the graph is leveraged in three complementary ways: (1) Item-item propagation with a Residual II Gate (RIG) that adaptively controls per-item absorption of semantic neighborhood signals for representation enhancement; (2) A content-guided UI graph expansion that introduces virtual user-item edges through high-confidence semantic neighbors for interaction graph enhancement; (3) II-Neighbor BPR Augmentation (INA) that treats top neighbors of positive items as discounted soft positives for optimization enhancement. We provide theoretical analysis showing that NCER reduces the spectral noise-to-signal ratio, RIG converges to a non-degenerate gating regime, and INA yields a tighter generalization bound. Extensive experiments on four datasets demonstrate that IIMRec consistently outperforms state-of-the-art baselines while running faster and consuming less GPU memory, with particularly strong gains under cold-start and sparse-interaction conditions.

cs.IR

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.

cs.CV

CJ26 Global QCD Analysis with Large-$x$ Jefferson Lab 6 and 12 GeV Data

We present CJ26, the new CTEQ-JLab global QCD analysis that incorporates for the first time the complete suite of JLab 6 GeV DIS measurements and the first published JLab 12 GeV measurements. Focused on the large-$x$ region, the analysis utilizes the increased $Q^2$ leverage of the 12 GeV data to uniquely disentangle higher-twist effects from off-shell nucleon corrections. This leads to a highly accurate determination of the $n/p$ structure function ratio and the $d/u$ valence quark ratio, with uncertainties reduced by 30-50% and 5-10%, respectively. We highlight the critical role of experimental correlated systematic uncertainties in achieving this precision and provide the resulting NLO PDFs and structure functions in LHAPDF format for general use.

hep-ph

Probing Star-Forming Properties via ALMA Observations of Massive Protocluster IRAS 15596-5301

To deepen our understanding of star-forming properties, we studied a massive protocluster IRAS 15596-5301 using ALMA 870 um and 3 mm data. High-resolution 870 um data reveal 34 dense cores, including 3 hot molecular cores, with subsequent line surveys detecting 22 molecular species toward them. Two velocity components (I15596-red/I15596-blue) were found in the averaged H13CO+(1-0) spectrum, and two filaments were identified from velocity-resolved integrated intensity maps. A spatial overlap between the two filaments was observed, and this overlapping region exhibits a distinct bridge-shaped feature in the position-velocity diagram constructed along the entire filamentary structures. Combined with the reduced H13CO+/HCO+ ratio in the overlapping region and the three-dimensional position-position-velocity cube data, we conclude that a non-head-on collision occurs between the edges of the two filamentary structures in IRAS 15596-5301. Cluster analysis demonstrates that clusters located in the collision region host more evolved chemical rich dense cores than their counterparts in other regions. Our results thus indicate that star formation in I15596 is triggered or accelerated by a mild non-head-on collision between two filaments.

astro-ph.GA

A proof for the conjecture on superlinear problems with Ambrosetti-Rabinowitz condition

This paper is devoted to exploring a new minimax approach by introducing a characteristic mapping family which is invariant under the smooth descending flow for initial value. The minimax approach is self-contained, and its features are markedly different from standard ones, as it identifies the existence of critical points and intrinsically presents a lower-bound estimate for the generalized Morse index at the corresponding critical point. This quantity can be effectively viewed as an alternative to the group action. As applications, under the Ambrosetti-Rabinowitz condition we offer a positive answer to the long-standing open problem on the existence of infinitely many distinct solutions for superlinear elliptic equations without symmetric hypothesis.

math.FA

The BigBite Calorimeter for the Super Bigbite Spectrometer Program at Jefferson Lab

We report features of the design, construction, installation, and performance of the BigBite Calorimeter (BBCal), a lead-glass electromagnetic calorimeter constructed as part of the BigBite Spectrometer (BBS), which served as the electron arm for the Super Bigbite Spectrometer (SBS) program of high-precision neutron electromagnetic form factor measurements in Hall A at Jefferson Lab. As a total-absorption calorimeter, BBCal provided the primary electron trigger for BBS, detecting (quasi-) elastically scattered electrons in the 1-4 GeV energy range with an energy resolution of approximately 6.2%, position resolution of 1.2 cm, and timing resolution of 0.5 ns.

physics.ins-det

Long Range Outlook for Short-Range Correlations

Short range correlated (SRC) N N pairs are pairs of nucleons with high relative momentum (prel > kF where kF ~ 250 MeV/c is the Fermi momentum in medium to heavy nuclei) and lower center of mass momentum. The motivation for studying SRC pairs ranges from a desire to achieve a more comprehensive understanding of the many-body nuclear wave-function at high-resolution to searching for explicit QCD-dynamics effects within the nuclear medium, not to mention connections to many other open problems in nuclear physics. Exploring short-range correlations was one of the physics motivations for building CEBAF (now Jefferson Lab). Scientists used the high luminosity and high energy of this cutting-edge machine to find kinematics that cleanly showed the signals of short-range correlations. This paved the way in the last two decades for tremendous progress understanding these correlations. This paper reviews recent progress and highlights outstanding questions and areas that need further study.

nucl-ex

FedSM: Robust Semantics-Guided Feature Mixup for Bias Reduction in Federated Learning with Long-Tail Data

Federated Learning (FL) enables collaborative model training across decentralized clients without sharing private data. However, FL suffers from biased global models due to non-IID and long-tail data distributions. We propose \textbf{FedSM}, a novel client-centric framework that mitigates this bias through semantics-guided feature mixup and lightweight classifier retraining. FedSM uses a pretrained image-text-aligned model to compute category-level semantic relevance, guiding the category selection of local features to mix-up with global prototypes to generate class-consistent pseudo-features. These features correct classifier bias, especially when data are heavily skewed. To address the concern of potential domain shift between the pretrained model and the data, we propose probabilistic category selection, enhancing feature diversity to effectively mitigate biases. All computations are performed locally, requiring minimal server overhead. Extensive experiments on long-tail datasets with various imbalanced levels demonstrate that FedSM consistently outperforms state-of-the-art methods in accuracy, with high robustness to domain shift and computational efficiency.

cs.LG

GraphLAMA: Enabling Efficient Adaptation of Graph Language Models with Limited Annotations

Large language models (LLMs) have demonstrated their strong capabilities in various domains, and have been recently integrated for graph analysis as graph language models (GLMs). With LLMs as the predictor, some GLMs can interpret unseen tasks described by natural language, and learn from a few examples in the prompts without parameter tuning, known as in-context learning (ICL). Another subset of GLMs utilizes abundant training labels to enhance model performance, known as instruction tuning. However, we argue that ICL on graphs has effectiveness issues due to fixed parameters and efficiency issues due to long context. Meanwhile, the large amount of labeled data required for instruction tuning can be difficult to obtain in real-world scenarios. To this end, we aim to introduce an extra parameter adaptation stage that can efficiently tailor GLMs to an unseen graph and task with only a few labeled examples, in exchange for better prediction accuracy and faster inference speed. For implementation, in this paper we propose GraphLAMA method, with its model backbone and learning schemes specialized for efficient tuning and inference. Specifically, for model backbone, we use a graph neural network (GNN) with several well-designed components to transform nodes into the representation space of LLM tokens. Task instructions can then be represented as a mixture of node and language tokens. In the pre-training stage, model parameters except the LLM will be trained with different tasks to capture general knowledge. In the adaptation stage, only a few pre-trained parameters will be updated based on few-shot examples. Extensive experiments on few/zero-shot node classification and summary generation show that our proposed GraphLAMA achieves state-of-the-art performance with 4.91% absolution improvement in accuracy. Compared with ICL, our inference speed can be 10 times faster under 5-shot setting.

cs.CL

Learning Speaker-Invariant Visual Features for Lipreading

Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape, color, texture), which introduce spurious correlations between vision and text. These correlations lead to suboptimal lipreading accuracy and restrict model generalization. To address this challenge, we introduce SIFLip, a speaker-invariant visual feature learning framework that disentangles speaker-specific attributes using two complementary disentanglement modules (Implicit Disentanglement and Explicit Disentanglement) to improve generalization. Specifically, since different speakers exhibit semantic consistency between lip movements and phonetic text when pronouncing the same words, our implicit disentanglement module leverages stable text embeddings as supervisory signals to learn common visual representations across speakers, implicitly decoupling speaker-specific features. Additionally, we design a speaker recognition sub-task within the main lipreading pipeline to filter speaker-specific features, then further explicitly disentangle these personalized visual features from the backbone network via gradient reversal. Experimental results demonstrate that SIFLip significantly enhances generalization performance across multiple public datasets. Experimental results demonstrate that SIFLip significantly improves generalization performance across multiple public datasets, outperforming state-of-the-art methods.

cs.CV

Blend the Separated: Mixture of Synergistic Experts for Data-Scarcity Drug-Target Interaction Prediction

Drug-target interaction prediction (DTI) is essential in various applications including drug discovery and clinical application. There are two perspectives of input data widely used in DTI prediction: Intrinsic data represents how drugs or targets are constructed, and extrinsic data represents how drugs or targets are related to other biological entities. However, any of the two perspectives of input data can be scarce for some drugs or targets, especially for those unpopular or newly discovered. Furthermore, ground-truth labels for specific interaction types can also be scarce. Therefore, we propose the first method to tackle DTI prediction under input data and/or label scarcity. To make our model functional when only one perspective of input data is available, we design two separate experts to process intrinsic and extrinsic data respectively and fuse them adaptively according to different samples. Furthermore, to make the two perspectives complement each other and remedy label scarcity, two experts synergize with each other in a mutually supervised way to exploit the enormous unlabeled data. Extensive experiments on 3 real-world datasets under different extents of input data scarcity and/or label scarcity demonstrate our model outperforms states of the art significantly and steadily, with a maximum improvement of 53.53%. We also test our model without any data scarcity and it still outperforms current methods.

cs.LG

HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs

Graph neural networks (GNNs) have demonstrated success in modeling relational data primarily under the assumption of homophily. However, many real-world graphs exhibit heterophily, where linked nodes belong to different categories or possess diverse attributes, such as webpages, Wikipedia articles, social networks, and e-commerce platforms. Additionally, nodes in many domains are associated with textual descriptions, forming heterophilic text-attributed graphs (TAGs). Despite their significance, heterophilic TAGs remain underexplored due to the lack of dedicated benchmarks that jointly capture heterophilic structures and rich textual attributes. To address this gap, we introduce the \textbf{He}terophilic \textbf{T}ext-attributed \textbf{G}raph \textbf{B}enchmark (HeTGB), a novel benchmark comprising five real-world heterophilic graph datasets from diverse domains, with nodes enriched by extensive textual descriptions. HeTGB enables systematic evaluation of GNNs, pre-trained language models (PLMs) and co-training methods on the node classification task. Through extensive benchmarking experiments, we showcase the utility of text attributes in heterophilic graphs, analyze the challenges posed by heterophilic TAGs and the limitations of existing models, and provide insights into the interplay between graph structures and textual attributes.

cs.CL

Systematic uncertainties from higher-twist corrections in DIS at large x

We investigate the systematic uncertainties and potential biases arising from the inclusion of large-$x$ corrections to proton and deuteron deep inelastic scattering (DIS) data in global quantum chromodynamics (QCD) analyses. Using the CTEQ-JLab framework, we examine various approaches to implementing higher-twist corrections in nucleon structure functions and off-shell PDF modifications in deuteron targets. We analyze how these components interact and influence the determination of the $d$-quark PDF and the neutron structure function at large $x$. We find that it is very important to consider isospin-dependent higher-twist corrections in order to minimize implementation biases in the extracted quantities.

hep-ph

New Measurements of the Deuteron to Proton F2 Structure Function Ratio

Nucleon structure functions, as measured in lepton-nucleon scattering, have historically provided a critical observable in the study of partonic dynamics within the nucleon. However, at very large parton momenta it is both experimentally and theoretically challenging to extract parton distributions due to the probable onset of non-perturbative contributions and the unavailability of high precision data at critical kinematics. Extraction of the neutron structure and the d-quark distribution have been further challenging due to the necessity of applying nuclear corrections when utilizing scattering data from a deuteron target to extract free neutron structure. However, a program of experiments has been carried out recently at the energy-upgraded Jefferson Lab electron accelerator aimed at significantly reducing the nuclear correction uncertainties on the d-quark distribution function at large partonic momentum. This allows leveraging the vast body of deuterium data covering a large kinematic range to be utilized for d-quark parton distribution function extraction. We present new data from experiment E12-10-002 carried out in Jefferson Lab Hall C on the deuteron to proton cross-section ratio at large BJorken-x. These results significantly improve the precision of existing data, and provide a first look at the expected impact on quark distributions extracted from global parton distribution function fits.

hep-ex

Exploring the Potential of Large Language Models for Heterophilic Graphs

Large language models (LLMs) have presented significant opportunities to enhance various machine learning applications, including graph neural networks (GNNs). By leveraging the vast open-world knowledge within LLMs, we can more effectively interpret and utilize textual data to better characterize heterophilic graphs, where neighboring nodes often have different labels. However, existing approaches for heterophilic graphs overlook the rich textual data associated with nodes, which could unlock deeper insights into their heterophilic contexts. In this work, we explore the potential of LLMs for modeling heterophilic graphs and propose a novel two-stage framework: LLM-enhanced edge discriminator and LLM-guided edge reweighting. In the first stage, we fine-tune the LLM to better identify homophilic and heterophilic edges based on the textual content of their nodes. In the second stage, we adaptively manage message propagation in GNNs for different edge types based on node features, structures, and heterophilic or homophilic characteristics. To cope with the computational demands when deploying LLMs in practical scenarios, we further explore model distillation techniques to fine-tune smaller, more efficient models that maintain competitive performance. Extensive experiments validate the effectiveness of our framework, demonstrating the feasibility of using LLMs to enhance node classification on heterophilic graphs.

cs.LG

Systematic uncertainty of offshell corrections and higher-twist contribution in DIS at large x

We study the systematic uncertainty and biases introduced by theoretical assumptions needed to include large-$x$ DIS data in a global QCD analysis. Working in the CTEQ-JLab framework, we focus on different implementations of higher-twist corrections to the nucleon structure functions and of offshell PDF deformations in deuteron targets and discuss how their interplay impacts the extraction of the $d$-quark PDF and the calculation of the neutron structure function at large $x$.

hep-ph

Unifying Structured Data as Graph for Data-to-Text Pre-Training

Data-to-text (D2T) generation aims to transform structured data into natural language text. Data-to-text pre-training has proved to be powerful in enhancing D2T generation and yields impressive performances. However, previous pre-training methods either oversimplified structured data into a sequence without considering input structures or designed training objectives tailored for a specific data structure (e.g., table or knowledge graph). In this paper, we unify different types of structured data (i.e., table, key-value data, knowledge graph) into the graph format and cast different data-to-text generation tasks as graph-to-text generation. To effectively exploit the structural information of the input graph, we propose a structure-enhanced pre-training method for D2T generation by designing a structure-enhanced Transformer. Concretely, we devise a position matrix for the Transformer, encoding relative positional information of connected nodes in the input graph. In addition, we propose a new attention matrix to incorporate graph structures into the original Transformer by taking the available explicit connectivity structure into account. Extensive experiments on six benchmark datasets show the effectiveness of our model. Our source codes are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/unid2t.

cs.CL