arXiv ScienceSearch

arXiv subjects

Hong Shen

Publications and source records attributed to Hong Shen.

At least 19 recordsLinked to original sources

Fine-grained Distributed Backdoor Attacks in Federated Learning

Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. Compared to centralized attacks, distributed backdoor attacks are more harmful but require more poisoned samples to compensate for the loss of trigger strength due to decomposition. Fixed trigger patterns are also easily detected by robust aggregation algorithms, increasing the risk of attack exposure. To address these challenges, we propose a fine-grained distributed backdoor attack framework (FDBA). This framework uses dynamic trigger generation and embedding vector optimization to perform attacks with fewer poisoned samples. First, we design a dynamic trigger generation method based on image edge structures using the Canny algorithm to extract edge features, which are then injected with Laplacian noise. RGB channel decomposition is applied for covert adaptation of the distributed trigger, reducing detection chances. Second, we introduce an embedding vector contrastive learning strategy that forces poisoned samples to approach the target class center in the feature space, enhancing attack effectiveness. On CIFAR-10, piecewise-linear estimates for target ASRs between 70\% and 90\% show that FDBA reduces the required poisoning ratio by 37.4\%--48.4\% compared with DBA. In non-independent and identically distributed (Non-IID) scenarios, FDBA retains 84.7\% of its IID attack performance under extreme heterogeneity, whereas DBA drops to 73.5\%, and the framework successfully bypasses mainstream defense mechanisms. This study offers new insights into federated learning security and emphasizes the potential threats and defense challenges posed by fine-grained distributed attacks.

cs.LG

Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration

This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. In the proposed method, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. It then filters unreliable clients, assigns reliability-based weights to the retained clients, and constructs a teacher for each type of knowledge. Finally, the corresponding distillation gradients are validated using a supervised gradient computed from private data. Conflicting prediction and boundary gradients are removed, and conflicting relation gradients are suppressed before the final model update. We prove the convergence of the proposed method by showing stable local optimization for honest clients under Byzantine distillation. Particularly, we show that our method ensures a bounded Byzantine influence on both distillation gradients and individual client private gradients after cross-modality fusion, thereby enabling stable local optimization for honest clienunder Byzantine distillation. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that the proposed method improves the prediction accuracy of heterogeneous models of clients under non-IID data and Byzantine attacks. As the booming demands of federated learning in decentralized environments such as edge computing and mission-oriented UAV collaborations, our method has a great potential for adoption of DFL in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions.

cs.LG

Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration

This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation and history-based update trend prediction, rather than purely aggregating client models as in the existing work. In R-DPFL, each client first computes the current-round model update by aggregating the received neighborhood update vectors. It then predicts what this update should be based on its historical values and local model changes. Finally, R-DPFL computes the difference between these two quantities, adaptively clips this difference, and adds it to the local update. We prove convergence of the learning process through rigorous analysis and show that honest clients maintain stable personalized descent dynamics under Byzantine neighbor perturbations without requiring consensus among neighboring models. Extensive experiments on CIFAR-10 demonstrate that RDPFL consistently outperforms state-of-the-art decentralized and personalized federated learning baselines under heterogeneous and adversarial settings.

cs.LG

Quantum information in neutron-proton scattering from the $M$ matrix

We study quantum-information aspects of neutron--proton scattering in the spin-space $M$-matrix framework. Four representative classes of input states are considered, namely diagonal mixed states, separable pure states, general two-qubit pure states, and a special Schmidt-like entangled subclass. For each class, ensemble-averaged output mutual information, reduced-state linear entropy, negativity, and geometric quantum discord are calculated in the relative momentum-- scattering angle plane. The results show that the outgoing spin correlations are governed jointly by scattering kinematics and by the structure of the incoming quantum ensemble. Input states with stronger intrinsic coherence or entanglement give larger maxima and higher minima in the mutual information, negativity, and geometric quantum discord. The enhanced regions of the mutual information and geometric discord depend on the input states, while the negativity maximum remains concentrated in the high-momentum backward-scattering region. These results extend earlier studies based on product-state entanglement power and provide an ensemble-based description of how spin correlations in neutron--proton scattering arise from the interplay between input-state structure and scattering dynamics.

nucl-th

AEGIS: Attention-Embedding Gradient Isolation Shield - Triple-Channel Gradient Masking for Privacy-Preserving Federated LLM Fine-Tuning

Gradient inversion attacks recover private training text from gradients shared in federated learning, posing a serious threat to collaborative model training. Through our analysis of transformer gradient structure, we identify three channels through which private token information leaks: the attention output projection gradient exposes a low-rank subspace that encodes input embeddings (Channel 1), the embedding gradient's row-norm sparsity directly reveals which tokens are present (Channel 2), and the MLP expansion gradient carries a recoverable subspace signal analogous to Channel 1 (Channel 3). State-of-the-art attacks exploit these channels analytically to achieve near-exact token recovery in seconds. Existing defences address at most one channel and either degrade model utility or leave the remaining structural signals intact. We introduce AEGIS (Attention-Embedding Gradient Isolation Shield), a lightweight defence that closes all three analytical channels with three backward-path operations requiring no architectural changes: freezing attention projection parameters eliminates Channel 1 by construction, calibrated noise injection into the embedding gradient destroys Channel 2's token-presence signal, and analogous per-block noise injection into the MLP expansion gradient masks Channel 3. The same masked gradient drives both the local optimiser step and the server export, so no clean signal is retained on either side. Evaluated across 11 models and six datasets, AEGIS reduces token recovery rates to near zero against a range of gradient inversion attacks, both analytical and optimisation-based, while preserving or improving model utility. We provide formal guarantees for Channels 1 and 2 and validate the full defence empirically against adaptive adversaries with complete knowledge of the mechanism.

cs.CR

Role of the $\delta$ Meson in Softening the Symmetry Energy within the DDRHF Model

We investigate the effects of the isovector-scalar $\delta$ meson on the density dependence of the symmetry energy within the density-dependent relativistic Hartree--Fock (DDRHF) framework. As a baseline, we generate $1006$ accepted DDRHF parametrizations including the $\sigma$, $\omega$, $\rho$, and $\pi$ mesons by imposing empirical constraints on the saturation properties of nuclear matter. The resulting symmetry-energy slope parameters are confined to relatively large values, $L\simeq65$--$110~\mathrm{MeV}$. Two representative parametrizations, denoted RHF-NK1 and RHF-NK2, are randomly selected from this ensemble. Starting from these two parametrizations, we introduce the $\delta$ meson and readjust the meson--nucleon couplings under the same saturation-property constraints. The numerical optimization shows that small values of $L$ are obtained most efficiently when the $\delta$ coupling is taken to be constant. In this case, $L$ is reduced from approximately $73$ to $32~\mathrm{MeV}$, while the binding energy per nucleon, saturation density, symmetry energy, and incompressibility coefficient remain nearly unchanged. A channel-by-channel decomposition shows that the softening is not caused by the direct $\delta$-meson contribution alone, but by a redistribution among the $\delta$, $\rho$, and $\pi$ mesons together with the isoscalar Fock contributions. The resulting neutron-star mass--radius relations shift toward smaller radii, indicating that the $\delta$ meson provides an efficient additional degree of freedom for controlling the isovector properties of DDRHF models.

nucl-th

Influence of effective mass of the relativistic mean field theory on core collapse supernovae and compact objects

We study the influence of the effective mass in the relativistic mean field (RMF) theory on the properties of the central core of collapse-driven supernovae and the formation of compact objects. Influence of the effective mass has been so far studied within the non-relativistic frameworks. In order to clarify the role of the effective mass in the relativistic frameworks, which is different from non-relativistic ones, we adopt the set of equation of state (EOS) tables using the parameterizations TM1e and TM1m, which have different effective masses but with the same saturation properties, in the RMF theory. We show that choices of the effective mass in supernova matter affect both the stiffness of the EOS through pressure and the thermodynamical behavior through temperature under the RMF frameworks. We explore differences in matter evolution with neutrino emissions by performing a set of numerical simulations of the gravitational collapse and bounce of massive stars and the cooling of the proto-neutron stars. The EOS with large effective mass leads to compact proto-neutron stars and early collapse to black holes with high densities and temperatures due to the softness. It leads to high energy neutrinos in long emission from the proto-neutron star cooling and in short burst from the black hole formation.

astro-ph.HE

Vortex Nucleons as Partial-Wave Filters in Nucleon--Nucleon Scattering

We propose vortex nucleon scattering as an angular-momentum-resolved probe of nucleon--nucleon partial waves. Using the standard $LSJ$ partial-wave $S$ matrix as input, we show that an on-axis vortex incident state with a fixed orbital angular-momentum projection $m_L=\ell$ imposes the direct selection rule $L\geq |\ell|$ on the initial nucleon--nucleon partial waves. As a result, the initial $S$ wave is excluded for $\ell=1$, while both the initial $S$ and $P$ waves are excluded for $\ell=2$. The underlying phase shifts are not modified. Instead, the vortex external state changes how the ordinary partial waves are projected into the scattering amplitude. We further analyze off-axis scattering, where the displacement of the target from the vortex axis introduces Bessel-function weights and partially relaxes the on-axis selection rule. These results suggest that vortex nucleons can provide a new experimental handle on the partial-wave content of the strong nucleon--nucleon interaction.

nucl-th

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existing representation alignment metrics are symmetric, collapsing both modalities into a single score and hiding which modality drives cross-modal degradation. We introduce the Spectral Alignment Score (SAS), an asymmetric metric that projects both modalities onto the principal eigenbasis of an anchor modality and computes eigenvalue-weighted per-eigenmode correlations, resulting in directional scores whose difference quantifies modality information imbalance. We embed SAS within a benchmarking framework evaluating 15 VLMs across natural and medical image-text datasets alongside 6 alignment metrics and bidirectional retrieval. Our experiments show that medical images retain richer structural information than their paired clinical reports, a directional asymmetry invisible to all competing metrics, and that SAS achieves the strongest zero-label correlation with retrieval performance in the medical domain, positioning it as a practical diagnostic tool for clinical deployment. Code is available at this URL: https://github.com/iamalegambetti/medical-vlms-assessment.

cs.CV

Design Principles for Human-Agent Interaction

AI agents are rapidly evolving into autonomous systems capable of sustained interaction, tool use, and long-term collaboration. Yet their real-world adoption remains limited, suggesting that the key barrier lies not only in technical capability but also in a lack of design knowledge for successful human-agent interaction. This position paper argues that AI agents should not be solely evaluated or deployed based on autonomous task capability alone; because agents interact with, adapt to, influence, and sometimes fail humans, human-agent interaction must be treated as a core design and evaluation target for agentic AI. We present 14 design principles that articulate the ideal human-agent relationship across four interaction stages: initially, during interaction, over time, and when things go wrong. We use these principles to evaluate nine agent systems to illustrate that these design principles can provide actionable guidance for AI design teams to systematically design and evaluate agents that are usable, trustworthy, and effective in real-world interactive settings.

cs.HC

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios

AI measurement science has a wide variety of methodologies and measurements for comparing AI systems, resulting in what often appear to be "apples-to-oranges" comparisons across AI evaluations. To move toward "apples-to-apples" comparisons in real-world AI evaluations, this work advocates for methodological transparency in evaluation scenarios, operational grounding, and human-centered design (HCD) principles. We propose a repeatable process for transforming high-level use cases to detailed scenarios by eliciting use cases from subject matter experts (SMEs) via a structured AI Use Case Worksheet with six key elements: use case, sector, user (direct and indirect), intended outcomes, expected impacts (positive and negative), and KPIs and metrics. We demonstrate utility of the worksheet and process in the U.S. financial services sector. This paper reports on example high-level AI use cases identified by financial services sector SMEs: cyber defense enablement, developer productivity, financial crime aggregation, suspicious activity report (SAR) filing, credit memo generation, and internal call center support. These AI use cases provided are illustrative of the process and not exhaustive. Central to our work is a three-stage expansion pipeline combining LLM prompting with human reviews to generate 107 scenarios from those use cases elicited from SMEs. This process integrates iterative human reviews at every juncture to ensure operational grounding: for scenario titles and descriptions; for core scenario elements like users, benefits and risks, and metrics; and for scenario narratives and evaluation objectives. Human checkpoints ensure scenarios remain reflective of real-world usage and human needs. We describe a validation rubric to assess scenario quality. By defining key scenario components, this work supports a more consistent and meaningful paradigm for human-centered AI evaluations.

cs.HC

Relativistic mean-field study of the neutron star inner crust using the asymmetric finite difference method

The ground-state properties of neutron-rich nuclear clusters in the inner crust of neutron stars are investigated within the Wigner-Seitz approximation using a relativistic mean-field framework. The radial Dirac equations are solved with an asymmetric finite-difference scheme, by which the hermiticity is preserved and spurious states are eliminated. Calculations are performed for representative Wigner-Seitz cells employing TM1-based interactions with different symmetry-energy slope parameters $L$, as well as a parametrization with a larger nucleon effective mass. It is found that the binding energy per nucleon decreases systematically with increasing $L$, while a larger effective mass leads to further reduction, particularly at higher densities. Quantum shell effects, which are absent in the Thomas-Fermi approximation, give rise to oscillatory density distributions and modify neutron properties. Within the Wigner-Seitz cell, the resulting neutron root-mean-square radius and chemical potential are shown to be sensitive to both $L$ and the effective nucleon mass, underscoring their important roles in determining the microscopic structure of the neutron-star inner crust.

nucl-th

An O(K)-Approximation Coflow Scheduling in K-Core Optical Circuit Switching Networks

Coflow has emerged as a fundamental application-layer abstraction in distributed systems, enabling collaborative management of related flows to enhance job completion efficiency. To meet the increasing bandwidth demands of modern data center networks (DCNs), optical circuit switches are widely deployed due to their high capacity and energy efficiency. Simultaneously, DCN deployments are evolving towards heterogeneous parallel architectures, where multiple independent optical circuit switching (OCS) cores operate concurrently to facilitate bandwidth expansion and incremental upgrades. However, existing research on coflow scheduling in multi-core switching fabrics primarily focuses on electrical packet switching (EPS) networks, with a few known results on OCS networks without or with a poor performance guarantee. This paper studies the coflow scheduling problem in multi-core OCS networks under the not-all-stop reconfiguration model, focusing on two major challenges of overcoming cross-core coupling for inter-core traffic allocation and satisfying the constraints of port exclusivity and reconfiguration overhead for intra-core circuit scheduling. To minimize total weighted coflow completion time (CCT), we propose an efficient algorithm by integrating LP-guided global coflow ordering, inter-core flow allocation and intra-core circuit scheduling that achieves approximation ratios of $8K$ and $\left(8K+1\right)$ for zero and arbitrary release times of coflows, respectively, where $K$ is the number of OCS cores. This framework is also applicable to $H$-core EPS networks, providing approximation guarantees of $4H$ and $\left(4H+1\right)$ for zero-time and arbitrary-time release, respectively.

cs.DC

What People See (and Miss) About Generative AI Risks: Perceptions of Failures, Risks, and Who Should Address Them

Despite growing concerns about the risks of Generative AI (GenAI), there is limited understanding of public perceptions of these risks and their associated failure modes -- defined as recurring patterns of sociotechnical breakdown across the GenAI lifecycle that contribute to risks of real-world harm. To address this gap, we present a survey instrument, validated with eight subject matter experts and deployed on a sample of 960 U.S.-based participants, to assess awareness and perceptions of GenAI's failure modes, their associated risks, and stakeholder responsibilities to address them. To support realism and content validity, our instrument is structured around scenarios grounded in publicly reported incidents and a taxonomy of GenAI's failure modes. Findings suggest that our instrument is (1) effective for assessing risk awareness and perceptions in a way that is grounded in people's current contexts of use, yet is extensible to new contexts that will inevitably arise; and (2) potentially useful for informing the design of AI literacy tools and interventions. We argue for AI literacy and governance approaches that align with how people encounter and reason about GenAI in everyday life.

cs.HC

Bayesian Inference of Dense-Matter Equations of State from Small-Radius Compact Stars with Twin-Star Scenarios

We investigate dense-matter equations of state (EOSs) within a Bayesian framework, with particular emphasis on whether recent small-radius compact-star candidates can be accommodated in a twin-star scenario. For the hadronic sector, we adopt a meta-modeling EOS constrained by the NICER mass--radius measurements of PSR J0030$+$0451, PSR J0437$-$4715, PSR J0614$-$3329, and the massive pulsar PSR J0740$+$6620. The hadronic inference indicates that PSR J0614$-$3329 favors a somewhat softer EOS than the other two \(\sim1.4\,M_\odot\) pulsars, while the \(\sim2\,M_\odot\) constraint prevents the EOS from becoming too soft. We then introduce a strong first-order phase transition through a constant-speed-of-sound quark-matter segment. Using HESS J1731$-$347 and XTE J1814$-$338 to constrain the phase-transition parameters, we find a preferred transition density of \(n_\mathrm{t}\sim2.7\text{--}2.8\,n_0\), a sizable energy-density jump of \(600\text{--}700\) MeV, and a relatively large post-transition sound speed of \(c_s^2/c^2\sim0.85\). Such a phase transition generates a disconnected hybrid branch with radii of about \(6\text{--}7\) km at masses around \(1.2\text{--}1.4\,M_\odot\), and strongly suppresses the dimensionless tidal deformability relative to the purely hadronic branch. This pronounced change in tidal deformability is a characteristic signature of the twin-star mechanism and may provide an important observational tool for identifying phase transitions in neutron-star matter in future multimessenger measurements. These results show that small-radius compact stars can provide direct constraints on both the strength of a first-order phase transition and the stiffness of the post-transition phase in dense matter.

astro-ph.HE

Impact of Effective Nucleon Mass and Multineutron States on the Equation of State for Core-Collapse Supernovae

In this study, we investigate the impact of effective nucleon mass and the existence of the dineutron $(\mathrm{^{2}n})$ and the tetraneutron $(\mathrm{^{4}n})$ on the thermodynamic properties and nuclear compositions by constructing new equations of state. Our results indicate that the model with a larger effective nucleon mass slightly alters the nuclear composition in neutron-rich environments primarily due to differences in the symmetry energy: the mass fractions of unbound neutrons, protons, and heavy nuclei increase. The impact on the thermodynamic properties is negligible, except for the chemical potentials. On the other hand, multineutron states become prominent at high densities in neutron-rich environments, leading to a substantial reduction in the unbound neutron fraction. This depletion lowers the chemical potential of unbound neutrons, which in turn reduces the abundance of neutron-rich nuclei. Consequently, the number of unbound protons increases, leading to a corresponding rise in proton chemical potential. These shifts in chemical potentials promote the formation of heavy nuclei with larger mass and atomic numbers. Ultimately, this compositional shift results in a lower free energy, primarily driven by the emergence of these heavy nuclei.

nucl-th

Crossover Equation of State Constrained by Astronomical Observations and pQCD

The hadron--quark crossover equation of state (EOS) of neutron star (NS) matter is investigated by combining relativistic mean-field (RMF) hadronic models with the Nambu--Jona-Lasinio (NJL) model for quark matter. The vector and diquark coupling constants of the NJL model are constrained using perturbative QCD (pQCD) calculations at high density through a scale-averaging likelihood approach, together with constraints from NS observations and the causality condition on the speed of sound. It is found that the diquark coupling is tightly constrained to $H \simeq 1.5G_s$, while the vector coupling is restricted to $G_v \lesssim 1.1G_s$ by the combined pQCD and astrophysical constraints. Crossover EOSs are constructed based on three hadronic RMF parameter sets, and their thermodynamic properties, sound speed behaviour, and trace anomaly are analysed. The resulting EOSs are applied to calculate NS global and dynamical properties, including mass--radius relations, tidal deformabilities, and fundamental radial oscillation frequencies. Compared with pure hadronic EOSs, the hadron--quark crossover is shown to significantly enhance the maximum NS mass, particularly for softer hadronic EOSs, while remaining consistent with observational bounds. It is further shown that the fundamental radial oscillation frequencies predicted by different EOSs exhibit pronounced differences, especially for intermediate-mass NSs, indicating that radial modes may provide a sensitive probe of the internal composition of NSs. These results indicate that quantitative NS observables may provide potential signatures of quark matter in NS interiors.

nucl-th

Scheduling Coflows in Multi-Core OCS Networks with Performance Guarantee

Coflow provides a key application-layer abstraction for capturing communication patterns, enabling the efficient coordination of parallel data flows to reduce job completion times in distributed systems. Modern data center networks (DCNs) are employing multiple independent optical circuit switching (OCS) cores operating concurrently to meet the massive bandwidth demands of application jobs. However, existing coflow scheduling research primarily focuses on the single-core setting, with multi-core fabrics only for EPS (electrical packet switching) networks. To address this gap, this paper studies the coflow scheduling problem in multi-core OCS networks under the \textit{not-all-stop} reconfiguration model in which one circuit's reconfiguration does not interrupt other circuits. The challenges stem from two aspects: (i) cross-core coupling induced by traffic assignment across heterogeneous cores; and (ii) per-core OCS scheduling constraints, namely \textit{port exclusivity} and \textit{reconfiguration delay}. We propose an approximation algorithm that jointly integrates cross-core flow assignment and per-core circuit scheduling to minimize the total weighted coflow completion time (CCT) and establish a provable worst-case performance guarantee. Trace-driven simulations using real Facebook workloads demonstrate that our algorithm effectively reduces weighted CCT and tail CCT.

cs.DC